
How Stixor AI Hub Works: Architecture, Features and How It Compares

Written by Muhammad Rizwan
October 9, 2026
4 min Read
Stixor AI Hub is Stixor's sovereign AI infrastructure platform: one layer between the people who use AI models and the hardware that runs them. This guide goes deeper than the product page. It covers how the platform is built, what each part does, how it compares with other ways of running AI, and how a rollout works.
In short, Stixor AI Hub gives every team or customer one OpenAI-compatible API, its own keys and budget, and a meter on every call. It also lets teams fine tune models and publish them to that same API. It is neither a model nor hardware. It works with open models on your own servers and with outside providers, on NVIDIA GPUs, Huawei Ascend NPUs and seven other accelerator families.
Why Do Enterprises Need an AI Platform Layer?
Because AI use now grows faster than anyone can track it, and most of the cost comes from running models every day, not from building them. Gartner forecasts that worldwide spending on AI-optimised infrastructure as a service will grow 96% in 2026 to $42 billion, and that inference spending (23.3billion)will surpass training(19 billion) for the first time this year.
Inference is the part that spreads across a company. Inside most organisations it looks like this:
- Each team opens its own vendor account and creates its own API keys.
- Nobody can say what each team is spending until the invoice arrives.
- Nothing stops a looping script or a leaked key before it runs up the bill.
- Compliance cannot show which model processed which data.
A platform layer fixes this by putting one front door in front of every model. Access, limits, metering and billing happen in one place, instead of in every team's own setup. That front door is what Stixor AI Hub provides.
How Is Stixor AI Hub Built?
Stixor AI Hub combines two products on one platform: Model as a Service (MaaS) for using models, and Platform as a Service (PaaS) for building and adapting them. Both share one scheduler and one pool of accelerators underneath.

The two layers are usually bought and run as separate systems, which means a fine-tuned model has to be handed from one team's tools to another's before anyone can use it. In Stixor AI Hub, a model built in the PaaS layer is registered and deployed to the MaaS API in one click, and from then on it is governed and metered exactly like every other model. The sections below take each layer in turn.
What Does the MaaS Layer Do?
The Model as a Service layer is a catalogue of AI models behind one API, with tenants, API keys, rate limits, budgets, metering and billing. It is where applications call models, and where every call is checked and counted.
One catalogue, one API
Every model you serve sits in one catalogue: open-source models on your own hardware, models fine tuned in-house and, where your policy allows, models from outside providers. Applications reach all of them through one OpenAI-compatible API, so code written for that format moves over by changing the base URL. Adding a model to the catalogue changes nothing in the applications.
Tenants, keys and rate limits
Each department, application or customer becomes a tenant. Tenants create their own API keys and see their own usage and remaining budget from a web console. Administrators decide which models each tenant may use and set rate limits per tenant and per key, in requests and tokens per minute. When a tenant needs more, it asks for an increase and an administrator approves it.
Budgets checked before the call
There are two limits. Each tenant sets its own budget cap, and an administrator sets a hard cap that the tenant cannot exceed. Both are checked at the API before the request reaches the model. A call that would break the budget is refused, so it generates no tokens and costs nothing.
Metering and billing
Every call is recorded per key: tokens in, tokens out, cost and time. Billing runs on prepaid credits or on-demand, with clear rules for switching between them, and per-tenant usage records come out ready to invoice. For an enterprise, that is internal chargeback. For a data centre, it is customer billing.
Because metering and billing run on your own platform, not a foreign provider's, you set the prices and the currency. A Pakistani enterprise can charge AI costs back to departments in rupees, and a Pakistani data centre can sell tokens to its customers in PKR. That removes the exchange rate risk that comes with dollar-billed AI services, where a weaker rupee raises the bill even when usage stays flat.
Serving from a shared pool
Models run from one shared pool of accelerators. Popular models stay always on, models with moderate demand share accelerators, and rarely used models load when called. Many fine-tuned variants can run on top of a single base model, so one cluster serves a large catalogue.
What Does the PaaS Layer Do?
The Platform as a Service layer is where teams build and adapt models, then publish them to the same API everyone else uses. It covers the path from first experiment to a model in production.
- Notebooks and raw compute. Hosted Jupyter notebooks with accelerator access for exploring data and prototyping.
- Training jobs. Run training and fine tuning jobs, bringing your own container when you need a custom setup.
- Fine-tune recipes. Ready recipes for LoRA, QLoRA and full fine tuning, and no-code tuning for teams without ML engineers. See LoRA vs QLoRA vs Full Fine Tuning [internal link: Blog 20].
- Experiment tracking. Compare runs and keep a record of what was tried.
- Model registry. Store approved versions and deploy them to the MaaS API in one click.
Once a model is deployed, it appears in the catalogue under the same keys, limits and metering as every other model. Compute used in this layer is metered per accelerator-second, so training and notebook time can be charged back just like inference.
Which Hardware Does Stixor AI Hub Run On?
Stixor AI Hub runs on nine accelerator families: NVIDIA, AMD, Huawei Ascend, T-head, Hygon, MetaX, Moore Threads, Cambricon and Iluvatar. Different vendors can run in the same deployment, because the accelerator is treated as a property of each job rather than a choice made for the whole platform.
Underneath both layers sits a Kubernetes-based scheduler. It pools many servers into one cluster and places each workload on suitable hardware. It also does three things that raise utilisation:
- Fractional slicing. An accelerator can be split into halves or quarters, so small models and notebooks share one card.
- Priority tiers and pre-emption. Production inference runs first, and lower-priority work gives way when needed.
- Spot backfill. Idle capacity is filled with cheaper, interruptible jobs.
For a data centre, that means more billable hours from the same racks. For an enterprise, it means fewer idle GPUs. Applications never need to know which chip answered a request.
Where Can Stixor AI Hub Be Deployed, and How Is Data Protected?
Stixor AI Hub runs entirely inside the environment you deploy it in: on-premise servers, a private cloud, a Kubernetes cluster you already run, public cloud instances, or a fully air-gapped site. Prompts, outputs, logs and billing records stay in that environment.
The controls that protect data are built into the platform rather than added later:
- Tenant isolation. Each tenant is separated at the namespace level, so one team's workloads and usage never affect another's.
- Scoped access. Model entitlements decide which models each tenant can reach, and every API key belongs to one tenant.
- Identity. Role-based access, with single sign-on through your existing identity provider.
- A record of every call. Each request is logged with its key, tenant, model, time and token count, so compliance can answer "which model saw this data" from the platform itself.
For sovereign and regulated deployments, such as banks or government in Pakistan, the air-gapped option runs with zero external dependencies. Stixor's Data Governance Compliance and Cloud Strategy Management teams help design these environments.
How Does Stixor AI Hub Compare With Other Ways of Running AI?

When each option makes sense
Vendor APIs are fine for prototypes and data that can leave the country. Self-hosting with vLLM or Ollama suits one application run by one team. A basic gateway helps when you only need one front door to outside providers. Building in-house can work for companies with a dedicated platform engineering team and time to spend.
Stixor AI Hub fits when several of these needs arrive at once: many teams or customers, a budget that must hold, data that must stay local, models you want to fine tune, and hardware you want to choose yourself
Who Uses Stixor AI Hub?
Two kinds of organisation: enterprises that want control over how their own teams use AI, and data centres or cloud providers that want to sell AI as a service. The product is the same for both. What each one gets out of it differs.

Sky47 is a customer of the platform. Stixor Technologies builds and owns Stixor AI Hub and offers it to other organisations too.
Operators who want help planning the hardware side can work with Stixor's Data Center Consulting & GPU Services team.
How Does a Stixor AI Hub Rollout Work?
It starts with mapping the models and teams you have today, and the platform deploys onto infrastructure you already run, including existing Kubernetes clusters. The steps differ slightly for the two kinds of user.
For an enterprise
- IT deploys Stixor AI Hub in front of the models the company uses: on its own servers, through outside providers, or both.
- Each team or application is set up as a tenant, with a model list, rate limits and a budget.
- Teams create their own API keys from the console.
- Developers point their applications at the new API. No code rewrite is needed.
- Every call is metered. IT sees spend per team in real time, teams see their remaining budget, and hard caps stop runaway spend automatically.
For a data centre or cloud provider
- The provider installs Stixor AI Hub on its accelerator servers and publishes a model catalogue.
- Customers sign up as tenants, choose prepaid credits or on-demand billing, and create keys.
- Customers call models through the API, or open a notebook, fine tune a model and deploy it to the same API.
- The scheduler spreads work across the pooled hardware to keep accelerators busy.
- Per-tenant usage records feed billing, while caps and rate limits protect the provider.
Stixor's MLOps Implementation Services team supports rollout, from the first tenant to production traffic.
The Bottom Line
As AI use moves from experiments to daily operations, the hard part becomes knowing who is using which model, keeping spend inside budget, and being able to show where data went. Stixor AI Hub handles that in one platform, for using models and for building them, on hardware you choose and in an environment you control.
Frequently Asked Questions
FAQs
Stixor AI Hub is an enterprise AI platform from Stixor Technologies. It sits between the people who use AI models and the hardware that runs them, giving each team one OpenAI-compatible API, its own keys and budget, per-call metering and billing, and tools to fine tune and deploy models.
Discuss Your Enterprise Use Case
From small to large scale enterprises, we deliver next-gen AI, data engineering, and actionable insights.