AI Factory for Service Providers

Turn Your AI Infrastructure Into a Billable Service

You have already invested in GPUs, models and inference endpoints. What is usually missing is everything that turns them into a service you can sell: per-tenant isolation, access control, usage metering, spend attribution and an invoice at the end of the month.

That is what Hybr® adds. The same engine that has run Microsoft CSP billing and multi-tenant cloud operations for service providers since 2016 now covers AI and ML workloads — without replacing the infrastructure you already chose.

Watch: turning racks into revenue

Two minutes on where the margin actually sits in an AI factory, and what has to exist between the rack and the customer’s invoice.

Read the transcript

This is the part everyone photographs. Steel, power, silicon. It is also the part that is already solved — if you have the capital, you can buy it. What you cannot buy off a shelf is the layer above it. That service layer is what connects your infrastructure to your customers’ invoices.

Every operator I talk to describes the same month end. Usage exported from one system. A price book in a spreadsheet. Someone reconciling it by hand, three days before invoices go out. It works, right up until an auditor or an investor asks you to prove it.

Start with what you already own. Hybr® registers the clusters and the fleet you already run — nodes, racks, GPUs, across sites. It does not provision your data centre and it does not replace your Kubernetes. It makes the estate legible, because you cannot sell or bill what you cannot see per tenant.

Hybr® is that service layer. It connects to the LLM gateway you already operate. What it adds is everything between the rack and the customer: a catalogue they can order from, metering per tenant, rating against their price list, and an invoice that is correct the first time.

Your customers already know what a cloud console feels like. Multi-tenant, role-scoped, self-service. Order a service, watch spend accrue, pull the invoice — without opening a ticket. And every unit consumed is metered against that tenant: completions, image units, GPU-hours.

And in most markets, capacity does not reach the mid-market directly. It reaches it through a channel. Multi-level resellers, each with their own price list, their own margin, their own branding. Colocation operators, enterprises charging back to business units, and public-sector programmes.

If you are building AI capacity, the useful test is small. Take one service you intend to sell. Model the offer, the meter, and the invoice.

Where Hybr sits in your AI stackTenants call OpenAI SDK compatible endpoints issued by the Hybr LLM Gateway, which meters usage and attributes spend, then routes to the gateway deployment, models and GPU capacity you already run.Where Hybr® sits in your AI stackYour infrastructure stays yours. Hybr® adds the tenancy, metering and billing layer in front of it.YOUR CUSTOMERSTenant AOpenAI SDK · base URL + keyTenant BOpenAI SDK · base URL + keyTenant COpenAI SDK · base URL + keyHYBR® LLM GATEWAYPer-tenant inference endpointsModel access controlUsage meteringSpend attributionHealth monitoringYOUR INFRASTRUCTUREGateway deploymentadmin endpoint + master keyModelsopen or commercialGPU capacityon-prem or cloudrequestsroutedConnect direct, or via Hybr Edge for private networks.Hybr® platformService catalog · offers & SKUs · multi-currency pricing · RBAC · one consolidated invoiceusage + spend

The gap between running AI and selling AI

Most teams reach the same wall. Inference works. The models are served. Then the questions start: which customer consumed what? Which model did the spend go to? How do we stop one tenant exhausting capacity? How do we put a margin on it and bill it?

Answering those questions with spreadsheets and exported logs does not scale past a handful of tenants — and it is where margin quietly disappears.

LLM Gateway — metered, multi-tenant model access

Point the Hybr® LLM Gateway at the gateway deployment you already run and it becomes a multi-tenant service. Each tenant gets an inference endpoint that works with the standard OpenAI SDK, so your customers and their developers change a base URL and a key — nothing else.

  • Per-tenant inference endpoints — OpenAI SDK compatible, issued and managed centrally.
  • Spend attribution per tenant — every registered gateway reports total spend alongside the tenants and models using it.
  • Model-level visibility — see which models are exposed and consumed, per gateway instance.
  • Health monitoring — gateway state is tracked continuously, not discovered when a customer complains.
  • Direct or via Hybr Edge — connect straight to the gateway, or route through Hybr Edge where the gateway sits on a private network.

Hybr® does not replace your gateway or host your models. It sits in front of what you already run and adds the commercial layer. More on the Hybr® LLM Gateway →

AI Applications — package once, publish, bill

An AI application is only a product once someone can order it, get it provisioned, and be billed for it. Hybr® treats AI apps the way it already treats every other service in your catalog.

  • Define a package — name, description, icon, category and sub-category, version, and menu classification.
  • Publish to the service catalog — with an offer and SKU attached, so the app is orderable and billable on the same commerce engine as the rest of your portfolio.
  • Deploy or import instances — stand up a new instance, or bring an already-running application under management.
  • Track health across the estate — every instance reports healthy, degraded or offline, broken down by subscription and by category, with version and location.
  • Control tenant visibility — hide packages from tenant infrastructure views when they should stay internal.

More on AI application and token billing →

It runs on the billing engine you already trust

This is the part competitors bolt on afterwards. AI usage in Hybr® lands in the same place as your Microsoft CSP subscriptions, your VMware and Azure Stack consumption, and your Kubernetes workloads — one catalog, one set of offers and price lists, one consolidated invoice per customer.

Multi-currency pricing, role-based access, reseller tiers and white-labelled self-service portals all apply to AI services exactly as they do to everything else you sell.

A console your customers already know how to use

Buyers arrive with an expectation set by the large public clouds: sign in, see what they are entitled to, order it, watch spend accrue, pull the invoice. Hybr® gives you that console over the infrastructure you already run — multi-tenant from the first customer rather than bolted on at the tenth.

  • Tenant isolation and role-based access — scoped so each customer sees only their own estate, users and spend.
  • A catalogue they order from — offers, SKUs and entitlements, so ordering and provisioning are the same action rather than a ticket.
  • Live consumption per tenant — GPU-hours, tokens, storage and egress attributed continuously, not reconstructed from logs at month end.
  • Self-service that stays self-service — change a subscription, add a user, download an invoice, without opening a support case.
  • Your brand, not ours — white-labelled portals so the experience belongs to you.

Sell through your channel, not only direct

In most markets, capacity does not reach the mid-market directly. It reaches it through partners. Hybr® treats the channel as part of the commerce model rather than something to reconcile afterwards — multi-level resellers, each with their own customers, price lists and margin, billing under one platform.

  • Multi-level reseller hierarchies — distributors, resellers and their end customers, each scoped to what they are allowed to see and do.
  • Price lists and markup per tier — pricing profiles with markups and discounts applied per reseller, per provider or per service.
  • Partner branding — white-labelled portals so your reseller sells under their own name, not yours.
  • Margin you can see — cost, sell price and margin visible per tier rather than derived from a spreadsheet at quarter end.
  • One engine for everything — GPU-hours, tokens, Microsoft CSP seats and managed services on the same invoice, under the same reseller structure.

Who this is for

  • Service providers and telcos building a sovereign or regional AI offering and needing per-customer metering from day one.
  • MSPs and Microsoft CSPs adding AI services to an existing portfolio without running a second billing stack.
  • Enterprises operating a shared internal AI platform that has to charge back to business units.

Frequently asked questions

Does Hybr® host the models?

No. Hybr® connects to the LLM gateway and infrastructure you already operate, and adds multi-tenancy, governance, metering and billing on top. You keep control of where models run and where data goes.

How does the Hybr® LLM Gateway connect to our infrastructure?

You register the gateway deployment you already run using its admin endpoint and master key, then set the tenant endpoint your customers use for inference. Hybr® issues each tenant an OpenAI SDK–compatible endpoint against it.

Do our customers have to change their code?

The tenant endpoint is used with the standard OpenAI SDK, so in most cases it is a base URL and an API key change.

What if the gateway is on a private network?

Register it in Via Hybr Edge connectivity mode rather than direct, so it does not need to be publicly reachable.

Can AI usage appear on the same invoice as Microsoft CSP?

Yes. That is the point of running it on the same commerce engine — one consolidated invoice per customer across every service you provide.

The AI Factory in detail: LLM Gateway · AI Application & Token Billing · Kubernetes & GPU-as-a-Service
Related: CSP Billing Ultimate · Cloud FinOps · VMware Alternative · What is an AI Factory?

The research behind this

We publish the analysis this positioning is built on, sources and all. Three recent pieces on where AI capacity is being built, what it rents for, and what has to be true before it can be sold.

See it with your own workloads

The fastest way to judge this is against your own gateway and your own tenant model. Book a working session and we will walk through it.