01Deployment

Run AI without your data leaving the building.

Business AI runs in our cloud, in a hybrid split, or entirely inside your network on your own hardware. The same identity, metering and spend controls apply in all three. Where the model runs is a configuration choice, not a different product.

Self-hosted models are a first-class mode, not an escape hatch

The platform's model configuration has a self-hosted mode that accepts an arbitrary base URL. Anything exposing an OpenAI-compatible API can be the target, which in practice covers Ollama, LM Studio and vLLM.

That model is then selected centrally for the workspace and inherited by every connected product, exactly like a frontier model would be. Your products do not know or care whether the endpoint they are reaching is in a datacentre or in your server room.

The three deployment shapes

Cloud
We run it. Fastest to value, and it works on your provider keys or ours depending on which billing mode you pick.
Hybrid
Frontier models for everyday work, local models for anything you would not paste into a public API. Because model choice is central configuration, moving a workload from one to the other is a setting rather than a rebuild.
On-premise
The whole layer runs inside your network on your hardware. No prompt, no document and no customer record leaves the building.

Why hybrid is usually the honest answer

Local models are the right choice for the fraction of work that carries real confidentiality risk. They are usually the wrong choice for everything else, because the frontier models are better and the difference shows up in output quality.

The failure mode we see is companies choosing one model for the whole organisation to avoid the routing question, then either sending sensitive material to a public API or accepting weaker output everywhere. Splitting the decision per workload avoids both, and it only stays practical if changing the split does not require engineering work.

What stays the same on-premise

Single sign-on, role-based access, per-call metering, spend caps and the company report are all part of the platform layer rather than the hosting arrangement. An on-premise deployment keeps them.

This matters more than it sounds. Most self-hosted AI setups solve the data question and lose the governance question, because the local endpoint has no idea who is calling it or what it has cost this month. Metering sits in the platform, so it survives the move.

FAQQuestions

Straight answers.

Any endpoint exposing an OpenAI-compatible API, which covers Ollama, LM Studio and vLLM among others. The platform stores a base URL for the self-hosted mode rather than a fixed list of vendors, so a runtime we have never heard of works if it speaks the same protocol.

No. Metering, limits, roles and reporting live in the platform layer, not in the hosting arrangement, so they behave the same whether the model is a frontier API or a box in your building.

Yes, that is the hybrid mode and it is the most common choice. Because the model is selected centrally and inherited by connected products, changing which workload goes where is a configuration change rather than a code change.

For inference against a local model, nothing: the prompt goes to your endpoint. The whole layer runs inside your network on your hardware, which is the point of the deployment mode.

Live in production

Build your company’s AI System.

Bring one real workflow to the call. We run it through the gateway while you watch, metered and attributed, before you decide anything.

Book a demo

30 minutes · no deck · reply within 24h