Your databases
Postgres, MySQL, BigQuery, Snowflake, and 20+ sources queried in place.
Run the complete analytics agent on your infrastructure with local models. Prompts, schema, and answers stay private while inference costs stay predictable.
Reference architecture / self-hosted GenBI
The context layer, the GenBI agent, and the model all deploy on your infrastructure — your security review has nothing external to chase.
Postgres, MySQL, BigQuery, Snowflake, and 20+ sources queried in place.
Grounds every question in your schema, metrics, and permissions.
Validated open-weight models served by Ollama, vLLM, or any OpenAI-compatible endpoint.
Data can't cross a border it never approaches — residency is met by the architecture, not a vendor contract.
Runs on networks with no internet route at all. Included with Enterprise Plus.
Permission-aware SQL, audit logs, and a constrained query surface keep every answer inside policy.
The LLM is a pluggable endpoint. Serve a validated model with Ollama or vLLM, point Wren AI at your base URL, and keep the same workflow you'd get in the cloud.
We benchmark open-weight models on real GenBI workloads and recommend only the ones that pass.
Schema, joins, and business definitions travel with every request, so local models write governed SQL instead of guessing.
A newer validated model is a config change — modeling, permissions, and saved questions carry over.
# Point Wren AI at the endpoint you already run.
type: llm
provider: litellm_llm
models:
- alias: default
model: openai/validated-model # from the tested list
api_base: "http://localhost:11434/v1"
timeout: 600Every question, report, and agent run is an LLM call. On metered APIs that bill compounds with adoption; on your own hardware it's a capacity plan.
Self-hosted licensing counts simultaneous active sessions — never named users.
The bigger the rollout, the stronger the case — every dashboard, report, and agent runs on inference you already own.
Run standalone or in a Kubernetes cluster, on infrastructure you size yourself — already proving out in production.
Wren AI enables natural language data interaction across 20+ databases without ETL or migration. Integrated with Phison's aiDAPTIV+ architecture, it delivers up to an 80% query hit rate and secure, on-prem AI operations—accelerating adoption with lower costs and faster deployment.
Both run standalone or in a Kubernetes cluster, with local models and unlimited users — the difference is how much of the rollout we deliver with you.
The full self-hosted platform, for teams who run their own infrastructure.
For regulated industries and organization-wide rollouts that need a formal security review and vendor-led delivery.
A concise, four-page field guide to the network boundary, governed context layer, local inference, security controls, and scale-out economics behind production GenBI.
Free PDF · 4 pages · Built for platform, data, and security teams



Wren AI runs against a curated set of open-weight models we test for SQL accuracy on GenBI workloads, served through any OpenAI-compatible endpoint such as Ollama or vLLM. We share the current validated list during your evaluation, and moving to a newer validated model is a configuration change, not a migration.
Yes. The context layer, the GenBI agent, and the model all run inside your network, so the deployment works with no internet route at all. On-prem and air-gapped deployment support is included in the Enterprise Plus plan.
Wren AI never sends a raw question to a bare model. Every request is grounded in your context layer (schema, joins, metrics, and permissions), which narrows the model's job. That grounding is what made up to an 80% query hit rate possible at Phison on fully on-prem infrastructure.
Two fixed parts: inference runs on hardware you provision, so the cost per question falls as usage grows, and the Wren AI license is priced by concurrent sessions rather than seats, with unlimited users. Business starts at 5 concurrent sessions and Enterprise Plus at 10; estimate your sessions or compare plans on the pricing page.
A concurrent session is a single LLM-driven request to Wren AI (a question, chart, report, or GenBI App), counted from request to output, whether it comes from the UI, API, Teams, Slack, or MCP. Licensing counts simultaneous active sessions, not named users; as a guide, one concurrent session supports around five active users.
Schedule a demo. We scope your deployment, share the validated local-model list, and set up a pilot in your environment on the Business or Enterprise Plus plan, so what you evaluate is exactly what you roll out.
We'll walk your data and security teams through a self-hosted rollout with local models — or start with the reference architecture your infrastructure team can review today.