Braintrust
AI evaluation and observability platform with self-serve pricing, MCP support, human eval loops, dataset management, and Series B traction. Infrastructure for eval-driven AI development teams.
Braintrust is a platform for building and maintaining quality AI products, combining observability, evaluation, and experimentation in one workflow. Its premise is that AI systems behave unlike traditional software: the same input can produce different outputs, there is rarely one correct answer, and a change that improves one thing can quietly degrade another. Braintrust is built to measure that quality systematically, catch regressions before they reach users, and give teams confidence that their system is actually improving.
The heart of the platform is evaluation. Teams define what good looks like, run experiments against real datasets, compare prompts and models side by side, and score outputs using language-model judges, code-based scorers, or human review. Offline evals run against known datasets before deployment, while online scoring runs the same judges on live production traffic to surface regressions and edge cases as they happen, automatically turning real failures into new test cases. A browser-based playground supports fast iteration, and experiments capture immutable, comparable snapshots of each run.
Observability sits alongside the evals rather than separate from them. Braintrust captures full traces of multi-step LLM and agent workflows, recording every model call and tool invocation with its inputs, outputs, cost, and latency, so engineers can see exactly what an agent did and where it went wrong. Because AI traces are large and deeply nested, the company built its own data store, Brainstore, to query millions of them quickly. Eval scores are native to the trace view, so quality signals and execution detail live in one place.
The platform closes the loop into shipping. A GitHub integration runs evals on every pull request and can block merges that degrade quality, while an AI assistant, Loop, takes natural-language instructions to analyze traces, generate eval datasets, and suggest better prompts and scorers. Braintrust also offers an AI gateway for logging, caching, and provider fallbacks, and an MCP server that lets a coding agent query logs and run evals from the IDE. It is framework-agnostic with SDKs across several languages, and carries enterprise security including SOC 2, GDPR, SSO, and HIPAA compliance.
Vendor details
Canonical URL
https://www.braintrust.dev
Category
Agent infrastructure
Company status
independent
Use cases & customers
Target customers
Deployment options
In practice
Your AI feature works in testing but quietly regresses in production and you find out from users. Braintrust scores live traffic with LLM-based judges, surfacing regressions as they happen and turning real failures into new test cases.
A prompt change looks better on one example but you can't tell if it's better overall. Braintrust runs your prompts and models against a dataset, scores them side by side, and captures an immutable experiment you can compare over time.
A bad change keeps slipping into your agent before anyone catches it. Braintrust runs your evals on every pull request and blocks the merge when a change degrades quality, gating releases on measured output quality.
Sources & related URLs
Agentic Index coverage score
8.0 / 14 capabilities · 57%
| Integrations & Tool CallingOfficial docs 2026-06-08 | Full |
|---|---|
| Workflow OrchestrationEval-focused infrastructure | Unable to verify |
| Knowledge Grounding & RAGDataset management docs 2026-06-08 | Partial |
| Human Oversight & GuardrailsHuman eval loops docs 2026-06-08 | Full |
| Security, Identity & GovernanceOfficial docs 2026-06-08 | Partial |
| Observability & AuditabilityCore platform offering 2026-06-08 | Full |
| Memory & State PersistenceEval-focused infrastructure | Unable to verify |
| Deployment & Data ResidencySelf-hosted and SaaS docs 2026-06-08 | Full |
| Prebuilt Agents, Templates & PacksEval-focused infrastructure | Unable to verify |
| Triggers & Channel CoverageEval-focused infrastructure | Unable to verify |
| Model Flexibility & RoutingMulti-model eval docs 2026-06-08 | Full |
| APIs, SDKs & MCP ExtensibilityAPI and MCP docs 2026-06-08 | Full |
| Testing, Debugging & OptimizationCore platform offering 2026-06-08 | Full |
| Browser & Computer UseOfficial docs 2026-06-08 | Unable to verify |
The Agentic Index coverage score grades every vendor Full, Partial or Unable to verify against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Braintrust introduced support for tracing Cloudflare Agents. Users can ingest traces by exporting them via OpenTelemetry or by instrumenting their agents directly in JavaScript on the Cloudflare Workers runtime.
Bears on: Integrations
View sourceBraintrust released terraform-aws-braintrust-data-plane v5.5.0, updating Braintrust Services/API ECS to v2.2.1, changing the default Redis instance type from cache.t4g.medium to cache.r7g.large, and enabling Brainstore fast readers by default.
Bears on: Deployment / data residency
View sourceBraintrust released bt CLI v0.12.0, adding dataset snapshot create/list/restore/delete support plus bt topics btmap download.
Bears on: Memory / state
View sourcePricing
Free Starter (1 GB data, 10K scores) · Pro $249/mo
usage
Included quota
Starter: 1 GB processed data and 10K scores a month, 14 day retention, unlimited seats. Pro: 5 GB and 50K scores, 30 day retention, Google SSO. No per seat fees on either plan.
What is public
Braintrust publishes full self serve pricing: free Starter with 1 GB processed data and 10K scores a month, Pro at $249/mo with 5 GB and 50K scores plus published overage rates, and custom Enterprise with BYOC and self hosted deployment. Qualifying early stage startups can get 6 to 12 months of free Pro.
Billing mechanics
Flat monthly platform fee per plan plus metered on demand usage beyond included processed data and scores. Pro overage rates are lower than Starter's. Topics and built in model usage draw from a separate monthly credit with uniform overage rates.
Cost watchouts
Overage billing has no hard spending cap: extra processed data runs $3/GB and extra scores $1.50 per 1K, so a heavy month can push the Pro bill well past $249. Starter's 14 day retention usually forces the upgrade before the data cap does. SAML/OIDC SSO, custom RBAC, BAA, and self hosting are Enterprise only.
Variable cost rationale
Platform fee is flat but processed data and score volume are metered with uncapped overage, so cost tracks trace volume rather than seats.
Additional watchouts
Model the overage math before committing: a Pro team logging 10 GB and 100K scores a month lands near $339, not $249.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free Starter plan, no card (1 GB data, 10K scores/mo, 14 day retention)
Lowest paid plan
Pro $249/mo (5 GB processed data, 50K scores, 30 day retention)
Commercial notes
Series A of $36M led by Andreessen Horowitz (Oct 2024). Enterprise customers include Dropbox. Plans restructured March 2026 into Starter/Pro/Enterprise.
Key ambiguities
Enterprise pricing is unpublished and reportedly requires a formal sales conversation. Plans were restructured in March 2026; legacy accounts may retain grandfathered features.
Missing data
Enterprise pricing unpublished.
Related vendors
- Acrab — Singapore compute infrastructure company building a full stack…
- AgentOps — Agent observability and reliability platform with broad model and…
- Agno — High-performance agent runtime and framework (formerly Phidata) with…
- AIsa — Unified resource and payment gateway for AI agents that lets them…
- AlphaBitCore — AI control plane that governs how models, agents, tools, and…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
Alternatives to Braintrust
The closest documented capability profiles to Braintrust among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Langfuse7.5 / 14A lighter documented profile than BraintrustBraintrust vs Langfuse →
- Portkey8.0 / 14Fuller documented coverage on Security, Identity & Governance
- Fiddler AI7.5 / 14Fuller documented coverage on Security, Identity & Governance
- LangWatch6.5 / 14A lighter documented profile than Braintrust
- AgentOps6.0 / 14A lighter documented profile than BraintrustBraintrust vs AgentOps →
- LiteLLM7.0 / 14Fuller documented coverage on Security, Identity & Governance
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded