The AI router — cheapest capable model, every time

Frontier answers.
A fraction of the cost.

Vyom reads every prompt and routes it to the cheapest model that can still nail it — escalating to frontier only when it matters. Same quality, up to 90% less spend.

Try the chat
97%AIME 2025
92%HumanEval
~$0.001per routine answer
Measured, not marketed

Frontier-grade accuracy.
A tenth of the cost.

Every score below is a real Vyom-router run on a public benchmark — exact-match graded. The router escalated only the hard items, so quality stayed high while cost stayed tiny.

Vyom router · task completion
AIME 2025 · Competition math97% n=30
MMLU · College CS · CS93% n=15
HumanEval · Coding · executed92% n=12
MMLU · Law · Legal87% n=15
MMLU · Medicine · Health73% n=15
Humanity's Last Exam · Frontier40% n=20

Vyom router run live through our harness on the stated public sets, exact-match graded. Subset sizes shown. Comparable to published per-model scores on the same sets — not a claim of an official leaderboard run.

The efficiency frontier
Vyom sits where you want to be: near-frontier quality, budget-tier cost.
60%70%80%90%100%cheapcostlyquality →cost per answer →VyomAlways-frontierSingle-model scalingAlways-budget

Cost/quality is a transparent projection from list prices + a representative workload. Single-model scaling ≈ inference-time approaches (e.g. Sakana Fugu); Vyom instead routes across many models.

Savings calculator

See what routing saves you. Screenshot it.

100,000
2,000
600
70%

Cost is projected from published list prices and your inputs — a transparent model, not a benchmark. Assumes ~35% input-token reduction from Headroom.

Projected monthly spend
Vyom router
Always frontier
You save
/mo
≈ ₹/mo
What Vyom does

The routing layer your model bill has been missing.

Cheapest capable model

A cheap triage pass classifies difficulty, reasoning, tools, and language — then routes to the right tier. Budget for chat, mid for real work, frontier only on escalation.

Context compression, on by default

Headroom compresses tool output, PDFs, web results, and history. Fewer tokens, same answer — lighter on budget models where lossy context hurts more.

Metering that protects revenue

Every call — chat, harness, or API — passes one server-side ledger enforcing per-user, per-tier limits before the model is ever called. No client-side trust.

Web search with citations

When a request needs fresh facts, Vyom searches, compresses the results, feeds them in, and cites sources [n] in the answer.

Three surfaces, one brain

A Claude-style chat, a coding agent forked from OpenCode, and an OpenAI-compatible API — all share the same orchestrator and the same usage ledger.

Voice, coming soon

Speak your prompt, hear the answer back. Speech-to-text → router → the cheapest capable model → text-to-speech. Endpoints are scaffolded today; the mic UI ships in v1.1.

The pipeline

One request, five deliberate steps.

01

Auth + limit check

Resolve the user and enforce their plan limits server-side — before anything runs.

02

Search + compress

Pull fresh facts if needed, then Headroom-compress everything down to the essentials.

03

Route to a tier

Triage classifies the task; the policy picks the cheapest capable tier and model.

04

Verify + respond

Destructive actions pass a verification gate. The answer streams back with its tier shown.

05

Log usage

Every decision — features, tier, tokens, cost — lands in one ledger across all surfaces.

🇮🇳 India-sovereign mode

Indian-language traffic stays in-region.

Turn on sovereignty mode, or just send Indian-language content, and Vyom routes it to an in-region model directly — bypassing the global aggregator entirely, so no data transits a third party. Data residency by construction, not by policy. This is the one thing the frontier labs can't sell you.

  • In-region reasoning models, INR-priced
  • Never routed through the global aggregator
  • Optional self-hosted India endpoint for full residency
# same API, sovereign route
POST /v1/chat/completions
{
"sovereignty": true,
"messages": [...]
}
Sovereign→ in-region · direct · zero data egress
Open-core lineage

Built on giants. Stated proudly.

OpenCode

MIT

Our coding harness is a fork of OpenCode, rebranded to the vyom CLI/TUI. Every model call routes through our orchestrator on the shared ledger — MCP connectors intact.

Headroom

Apache 2.0

The context-compression engine behind our token savings. Per-tier aggressiveness keeps budget models sharp while squeezing frontier calls hard.

The framework is MIT & public and runs fully on your own keys. Only our tuned routing weights and hosted infra stay private.

Read the open-core split →
Pricing

Start free. Scale when it pays for itself.

One ledger across chat, harness, and API. Server-side limits, always.

Free

₹0forever
  • ~50 messages / day
  • Budget tier models
  • Web search + citations
  • Community support
Start free
Most popular

Pro

₹2,999/ month
  • High daily chat cap
  • Budget + mid tier
  • Limited frontier escalations
  • Coding harness access
  • Priority routing
Go Pro

Business

₹24,999/ month
  • Everything in Pro
  • Frontier escalation
  • OpenAI-compatible API keys
  • India-sovereign mode
  • Usage analytics
Choose Business

Enterprise

Customcontact sales
  • Custom caps & SLAs
  • Full India data residency
  • Self-hosted endpoints
  • Dedicated support
  • Security review
Contact sales

Enterprise pricing is a placeholder pending real deals — no fabricated numbers anywhere on this page.