00Deployment / On-premise + local models

GenBI that stays inside your network.

Run the complete analytics agent on your infrastructure with local models. Prompts, schema, and answers stay private while inference costs stay predictable.

  • Standalone or Kubernetes
  • Fully air-gapped option
  • No data migration
Private deployment / online
Your network
01 / Request
Ask a business question
UI · API · MCP
02 / Ground
Wren context layer
Schema · metrics · permissions
03 / Infer
Local model endpoint
Ollama · vLLM · compatible API
04 / Query
Your data sources
Queried in place — no migration
Egress
0
external AI calls
Data path
IN-PLACE
schema and rows stay yours
Air-gap ready
No metered APIs
Unlimited users

Reference architecture / self-hosted GenBI

Runs in your network
100%
Per-token API fees
$0
Data sources, one layer
20+
Users, session-based license
Unlimited
01Compliance

Everything runs inside your boundary.

The context layer, the GenBI agent, and the model all deploy on your infrastructure — your security review has nothing external to chase.

Private deployment boundary
Your network
Standalone or KubernetesAir-gap ready
Source

Your databases

Postgres, MySQL, BigQuery, Snowflake, and 20+ sources queried in place.

No ETL, no migration
GenBI

Wren AI

Grounds every question in your schema, metrics, and permissions.

Standalone or Kubernetes
Inference

Local LLM

Validated open-weight models served by Ollama, vLLM, or any OpenAI-compatible endpoint.

Your GPUs
External AI APIs
Egress blocked by design
Blocked
0outbound calls
0prompts shared
0tokens metered
01

Data residency by architecture

Data can't cross a border it never approaches — residency is met by the architecture, not a vendor contract.

02

Air-gapped ready

Runs on networks with no internet route at all. Included with Enterprise Plus.

03

Governed by default

Permission-aware SQL, audit logs, and a constrained query surface keep every answer inside policy.

02Local models

Local models, validated for GenBI.

The LLM is a pluggable endpoint. Serve a validated model with Ollama or vLLM, point Wren AI at your base URL, and keep the same workflow you'd get in the cloud.

  • 01

    A tested, validated model list

    We benchmark open-weight models on real GenBI workloads and recommend only the ones that pass.

  • 02

    The context layer does the heavy lifting

    Schema, joins, and business definitions travel with every request, so local models write governed SQL instead of guessing.

  • 03

    Swap models without re-platforming

    A newer validated model is a config change — modeling, permissions, and saved questions carry over.

config.yaml
# Point Wren AI at the endpoint you already run.
type: llm
provider: litellm_llm
models:
  - alias: default
    model: openai/validated-model   # from the tested list
    api_base: "http://localhost:11434/v1"
    timeout: 600
Serves viaOllamavLLMOpenAI-compatible
Ask for the validated model list
03Economics

Pay for machines, not tokens.

Every question, report, and agent run is an LLM call. On metered APIs that bill compounds with adoption; on your own hardware it's a capacity plan.

GenBI on metered APIs

Cost driver
Per token — every question, retry, and agent step
As adoption grows
Spend climbs with success
Budgeting
Variable, hard to forecast
Data path
Prompts and schema leave your network

GenBI on local models

Fixed cost
Cost driver
Machines you size once, run at capacity
As adoption grows
Cost per question falls
Budgeting
A fixed line item
Data path
Nothing leaves your network
Licensing

Concurrent sessions, not seats.

Self-hosted licensing counts simultaneous active sessions — never named users.

Users
Unlimited
One session
≈ 5 active users
API & agents
Count same as UI

The bigger the rollout, the stronger the case — every dashboard, report, and agent runs on inference you already own.

04 / Scale

Built for large-scale GenBI rollouts.

Run standalone or in a Kubernetes cluster, on infrastructure you size yourself — already proving out in production.

80%
Query hit rate at Phison, fully on-prem
20+
Databases connected without ETL
100+
Hours saved monthly by one data team
Wren AI enables natural language data interaction across 20+ databases without ETL or migration. Integrated with Phison's aiDAPTIV+ architecture, it delivers up to an 80% query hit rate and secure, on-prem AI operations—accelerating adoption with lower costs and faster deployment.
Wei
CTO, Phison (A Public Company)
05Getting there

Choose your self-hosted plan.

Both run standalone or in a Kubernetes cluster, with local models and unlimited users — the difference is how much of the rollout we deliver with you.

Plan 01 / Business

Wren AI Business

The full self-hosted platform, for teams who run their own infrastructure.

  • Full platform: UI, dashboards, governance
  • Standalone or Kubernetes, in your cloud or data center
  • Unlimited users, starts at 5 concurrent sessions
  • Standard support
Recommended
Plan 02 / Enterprise Plus

Wren AI Enterprise Plus

For regulated industries and organization-wide rollouts that need a formal security review and vendor-led delivery.

  • Standalone or Kubernetes, including air-gapped environments
  • SSO and SCIM 2.0
  • Unlimited users, starts at 10 concurrent sessions
  • Vendor-led onboarding, procurement and security review
Reference architecture · Free PDF

Take the on-premise blueprint into your next architecture review.

A concise, four-page field guide to the network boundary, governed context layer, local inference, security controls, and scale-out economics behind production GenBI.

  • 01One-boundary reference architecture
  • 02Local models and governed context
  • 03Row and column security
  • 04Scale-out infrastructure and TCO

Free PDF · 4 pages · Built for platform, data, and security teams

Cover of The On-Premise GenBI Reference Architecture
On-Premise GenBI whitepaper architecture page
On-Premise GenBI whitepaper security and model page
FAQ / On-premise

On-premise questions, answered.

Wren AI runs against a curated set of open-weight models we test for SQL accuracy on GenBI workloads, served through any OpenAI-compatible endpoint such as Ollama or vLLM. We share the current validated list during your evaluation, and moving to a newer validated model is a configuration change, not a migration.

Yes. The context layer, the GenBI agent, and the model all run inside your network, so the deployment works with no internet route at all. On-prem and air-gapped deployment support is included in the Enterprise Plus plan.

Wren AI never sends a raw question to a bare model. Every request is grounded in your context layer (schema, joins, metrics, and permissions), which narrows the model's job. That grounding is what made up to an 80% query hit rate possible at Phison on fully on-prem infrastructure.

Two fixed parts: inference runs on hardware you provision, so the cost per question falls as usage grows, and the Wren AI license is priced by concurrent sessions rather than seats, with unlimited users. Business starts at 5 concurrent sessions and Enterprise Plus at 10; estimate your sessions or compare plans on the pricing page.

A concurrent session is a single LLM-driven request to Wren AI (a question, chart, report, or GenBI App), counted from request to output, whether it comes from the UI, API, Teams, Slack, or MCP. Licensing counts simultaneous active sessions, not named users; as a guide, one concurrent session supports around five active users.

Schedule a demo. We scope your deployment, share the validated local-model list, and set up a pilot in your environment on the Business or Enterprise Plus plan, so what you evaluate is exactly what you roll out.

Next step / architecture review

Bring GenBI inside your firewall.

We'll walk your data and security teams through a self-hosted rollout with local models — or start with the reference architecture your infrastructure team can review today.