Zero cloud. Zero risk. — on-prem models · tenant isolation · audit-chained
Build agents grounded in your own manuals, contracts and data — answers with citations, rules your team authored, and a paper trail your auditors can verify. Deployed on infrastructure you own.
60–80%
lower model cost
on-prem inference vs metered cloud APIs
100%
data sovereignty
row-level isolation, enforced by the database
Days
to deploy
three services, one Postgres, your hardware
Zero
cloud dependency
egress refused by default, allowlist by ADR
Stop bolting AI on. Start building it in.
Six pillars, one request path — every one of them enforced in code, not promised in a slide.
Upload manuals, contracts and drawings — up to thousands of pages each. RAG-powered hybrid retrieval (dense, lexical, rerank, graph) finds the passage, and every answer cites it or the agent refuses.
Each agent has a SOUL: a mission, a verification gate, a failure protocol and security rules, all versioned. Compose them into workflows, councils that deliberate, and durable multi-step runs.
Agents remember across sessions — and every belief traces to the document chunks that taught it. Erase a document and the memory it grounded is quarantined, with a receipt.
PII redaction before any model call, prompt-injection screening, a tamper-evident audit chain, and per-tenant row-level isolation enforced by the database itself.
Bring your own Postgres, your own Qdrant, your own model-provider keys. White-label endpoint aliases for your integrations. Your data plane, your bill, your domain.
REST with scoped API keys, webhooks for run and invoice events, an MCP server for your tools, streaming runs over SSE. The web UI uses the same endpoints you do.
Every plan includes on-prem inference tokens. Overage is metered at the posted rate — the plan page in your workspace shows the same numbers you see here.
In your workspace's tenant rows, isolated by database-enforced row-level security — and on Business and Enterprise plans you can bind your own Postgres and Qdrant, so new data lands on infrastructure you operate. Models run on the on-prem fleet; nothing reaches a cloud provider unless you bring your own key for one.
A model is a component, not a product. This is the layer that turns one into a deployable system: retrieval grounded in your documents, agents with authored rules, memory with provenance, and an audit chain — engineered for answers your auditors can verify.
The agent refuses and says why. Every answer cites the passage it stands on or does not ship — a refusal with a reason beats a confident guess.
Each plan includes a bundle of on-prem inference tokens; usage beyond it is metered at the posted rate and invoiced. Bring your own provider key and that provider's usage bills to your account instead of ours.
Enterprise plans get named endpoint aliases — your integrations call /a/your-bot/run instead of a raw agent id — plus data-residency binding and SSO/SAML. Custom domains ride on the same alias layer.
Start on the free plan today, or talk to the AI architects behind the platform about an on-premise deployment.