Frontier answers.
A fraction of the cost.
Vyom reads every prompt and routes it to the cheapest model that can still nail it — escalating to frontier only when it matters. Same quality, up to 90% less spend.
Frontier-grade accuracy.
A tenth of the cost.
Every score below is a real Vyom-router run on a public benchmark — exact-match graded. The router escalated only the hard items, so quality stayed high while cost stayed tiny.
Vyom router run live through our harness on the stated public sets, exact-match graded. Subset sizes shown. Comparable to published per-model scores on the same sets — not a claim of an official leaderboard run.
Cost/quality is a transparent projection from list prices + a representative workload. Single-model scaling ≈ inference-time approaches (e.g. Sakana Fugu); Vyom instead routes across many models.
See what routing saves you. Screenshot it.
Cost is projected from published list prices and your inputs — a transparent model, not a benchmark. Assumes ~35% input-token reduction from Headroom.
The routing layer your model bill has been missing.
Cheapest capable model
A cheap triage pass classifies difficulty, reasoning, tools, and language — then routes to the right tier. Budget for chat, mid for real work, frontier only on escalation.
Context compression, on by default
Headroom compresses tool output, PDFs, web results, and history. Fewer tokens, same answer — lighter on budget models where lossy context hurts more.
Metering that protects revenue
Every call — chat, harness, or API — passes one server-side ledger enforcing per-user, per-tier limits before the model is ever called. No client-side trust.
Web search with citations
When a request needs fresh facts, Vyom searches, compresses the results, feeds them in, and cites sources [n] in the answer.
Three surfaces, one brain
A Claude-style chat, a coding agent forked from OpenCode, and an OpenAI-compatible API — all share the same orchestrator and the same usage ledger.
Voice, coming soon
Speak your prompt, hear the answer back. Speech-to-text → router → the cheapest capable model → text-to-speech. Endpoints are scaffolded today; the mic UI ships in v1.1.
One request, five deliberate steps.
Auth + limit check
Resolve the user and enforce their plan limits server-side — before anything runs.
Search + compress
Pull fresh facts if needed, then Headroom-compress everything down to the essentials.
Route to a tier
Triage classifies the task; the policy picks the cheapest capable tier and model.
Verify + respond
Destructive actions pass a verification gate. The answer streams back with its tier shown.
Log usage
Every decision — features, tier, tokens, cost — lands in one ledger across all surfaces.
Indian-language traffic stays in-region.
Turn on sovereignty mode, or just send Indian-language content, and Vyom routes it to an in-region model directly — bypassing the global aggregator entirely, so no data transits a third party. Data residency by construction, not by policy. This is the one thing the frontier labs can't sell you.
- In-region reasoning models, INR-priced
- Never routed through the global aggregator
- Optional self-hosted India endpoint for full residency
Built on giants. Stated proudly.
OpenCode
MITOur coding harness is a fork of OpenCode, rebranded to the vyom CLI/TUI. Every model call routes through our orchestrator on the shared ledger — MCP connectors intact.
Headroom
Apache 2.0The context-compression engine behind our token savings. Per-tier aggressiveness keeps budget models sharp while squeezing frontier calls hard.
The framework is MIT & public and runs fully on your own keys. Only our tuned routing weights and hosted infra stay private.
Read the open-core split →Start free. Scale when it pays for itself.
One ledger across chat, harness, and API. Server-side limits, always.
Free
- ~50 messages / day
- Budget tier models
- Web search + citations
- Community support
Pro
- High daily chat cap
- Budget + mid tier
- Limited frontier escalations
- Coding harness access
- Priority routing
Business
- Everything in Pro
- Frontier escalation
- OpenAI-compatible API keys
- India-sovereign mode
- Usage analytics
Enterprise
- Custom caps & SLAs
- Full India data residency
- Self-hosted endpoints
- Dedicated support
- Security review
Enterprise pricing is a placeholder pending real deals — no fabricated numbers anywhere on this page.