New Vantage 2.0 — semantic evals now in general availability

The control plane for production AI

Vantage traces every prompt, scores every response, and routes each request to the model that gets it right. Your team ships on evidence instead of vibes — and finds out about regressions before your customers do.

Free forever for 50,000 traces a month · No credit card · SDK drops in with two lines

vantage.app / checkout-copilot / production Live

Recent traces

last 60s
retrieval → rerank → answer 412 ms
tool: lookup_order(#40213) 88 ms
guardrail: pii_redaction 21 ms
retry: schema_mismatch ×1 1.9 s
summarise → verify → emit 306 ms

Eval score

v42 → v43
94.2composite

▲ 3.8 pts vs. last release

Frontier model 18%

Mid-tier 61%

Cached / small 21%

Trusted by AI teams shipping to millions

  • Northwind
  • Helix Labs
  • Meridian
  • Kestrel
  • Orbital

By the numbers

1.4B
Model calls traced and scored every month across customer workloads.
41 ms
Median overhead added by the Vantage gateway, measured at p50.
37%
Median inference spend removed in the first quarter after smart routing.
99.99%
Gateway uptime over the trailing twelve months, publicly reported.
The platform

Everything between your prompt and production

Three systems that work as one: see what happened, prove it was good, and send the next request somewhere smarter. No sampling, no guesswork, no rewrites.

Traces with nothing left out

Every prompt, tool call, retry, guardrail and token — captured end to end and stitched into a single timeline you can replay. Open any conversation from six weeks ago and see exactly which context the model was holding when it went wrong.

  • Full capture at 100% of traffic, not a 1% sample
  • OpenTelemetry-native, so it lands in the tools you already run
  • PII redacted in-process, before anything leaves your VPC

Evals that block bad releases

Grade responses against your own rubrics, golden datasets and semantic checks — then wire the score into CI. A prompt change that quietly drops answer quality by four points fails the pull request instead of reaching your users.

  • Model-graded, rule-based and human review in one suite
  • Regression gates for GitHub Actions, GitLab and Buildkite
  • Datasets built from real failures with two clicks

Routing that pays for itself

One gateway in front of every provider. Vantage sends each request to the cheapest model that still clears your quality bar, caches what repeats, and fails over in milliseconds when an upstream provider starts throwing errors.

  • Quality-aware routing driven by your live eval scores
  • Automatic failover and retry budgets across providers
  • Per-team spend caps, enforced at the gateway

Customers

We went from arguing about whether the assistant had gotten worse to knowing by lunchtime. Vantage caught a regression in a retrieval change that our own test suite waved straight through — and routing paid for the whole contract in the first six weeks.

Rae OkonkwoVP Engineering, Northwind

Pricing

Priced per project, not per seat

Start free and stay free while you are building. Every plan includes full-fidelity tracing — the difference is scale, retention and control.

Starter

For the first prototype and the weekend that turns it into a product.

$0 / month

Free forever — no card required

Create an account
  • 50,000 traces per month
  • 3 seats, 1 project
  • 7-day trace retention
  • 2 eval suites, run on demand
  • Community Slack support

Upgrade later without touching a line of code.

Most popular

Scale

For teams with real traffic, real customers and a real on-call rota.

$79 / month

Per project, billed monthly

Start 14-day trial
  • 2 million traces per month
  • Unlimited seats and projects
  • 90-day retention and replay
  • Smart routing and semantic cache
  • Unlimited eval suites and CI gates
  • SSO, audit log, priority support

Most teams clear the subscription in routing savings alone.

Enterprise

For regulated workloads that have to prove every answer, on your own infrastructure.

Custom pricing

Volume pricing, billed annually

Talk to sales
  • Everything in Scale, without limits
  • Self-hosted in your VPC or on-prem
  • Custom retention and data residency
  • SOC 2 Type II, HIPAA, ISO 27001
  • 99.99% SLA with named architect
  • Migration and onboarding included

Security review pack available before the first call.

All plans include unlimited trace search, the full SDK and every model provider. Overage is billed at $9 per additional million traces — never by surprise.

Find out what your AI is actually doing

Install the SDK, ship one request, and watch the first trace land in under five minutes.