Zero cloud. Zero risk. — on-prem models · tenant isolation · audit-chained

Your Data. Your Premises. Your AI.

Build agents grounded in your own manuals, contracts and data — answers with citations, rules your team authored, and a paper trail your auditors can verify. Deployed on infrastructure you own.

60–80%

lower model cost

on-prem inference vs metered cloud APIs

100%

data sovereignty

row-level isolation, enforced by the database

Days

to deploy

three services, one Postgres, your hardware

Zero

cloud dependency

egress refused by default, allowlist by ADR

Stop bolting AI on. Start building it in.

What the platform does

Six pillars, one request path — every one of them enforced in code, not promised in a slide.

Knowledge that answers with citations

Upload manuals, contracts and drawings — up to thousands of pages each. RAG-powered hybrid retrieval (dense, lexical, rerank, graph) finds the passage, and every answer cites it or the agent refuses.

Agents you author, not just prompt

Each agent has a SOUL: a mission, a verification gate, a failure protocol and security rules, all versioned. Compose them into workflows, councils that deliberate, and durable multi-step runs.

Memory with provenance

Agents remember across sessions — and every belief traces to the document chunks that taught it. Erase a document and the memory it grounded is quarantined, with a receipt.

Governance built into the request path

PII redaction before any model call, prompt-injection screening, a tamper-evident audit chain, and per-tenant row-level isolation enforced by the database itself.

Your infrastructure, tier by tier

Bring your own Postgres, your own Qdrant, your own model-provider keys. White-label endpoint aliases for your integrations. Your data plane, your bill, your domain.

Every surface is an API

REST with scoped API keys, webhooks for run and invoice events, an MCP server for your tools, streaming runs over SSE. The web UI uses the same endpoints you do.

Pricing

Every plan includes on-prem inference tokens. Overage is metered at the posted rate — the plan page in your workspace shows the same numbers you see here.

Loading plans…

Questions engineers ask first

Where does my data actually live?

In your workspace's tenant rows, isolated by database-enforced row-level security — and on Business and Enterprise plans you can bind your own Postgres and Qdrant, so new data lands on infrastructure you operate. Models run on the on-prem fleet; nothing reaches a cloud provider unless you bring your own key for one.

How is this different from calling a model API?

A model is a component, not a product. This is the layer that turns one into a deployable system: retrieval grounded in your documents, agents with authored rules, memory with provenance, and an audit chain — engineered for answers your auditors can verify.

What happens when an answer can't be grounded?

The agent refuses and says why. Every answer cites the passage it stands on or does not ship — a refusal with a reason beats a confident guess.

How does billing work?

Each plan includes a bundle of on-prem inference tokens; usage beyond it is metered at the posted rate and invoiced. Bring your own provider key and that provider's usage bills to your account instead of ours.

Can we white-label it?

Enterprise plans get named endpoint aliases — your integrations call /a/your-bot/run instead of a raw agent id — plus data-residency binding and SSO/SAML. Custom domains ride on the same alias layer.

AI that never leaves your building.

Start on the free plan today, or talk to the AI architects behind the platform about an on-premise deployment.