Braintrust
AI observability and evaluation platform: traces, datasets, experiments, online scoring and human review, with hosted prompts and tools, a model gateway, the Loop agent and an MCP server.
Braintrust is an AI observability and evaluation platform for instrumenting, understanding and improving agents. It connects the pieces of an AI development loop: production traces, human feedback, test datasets, experiments, quality scores, prompt iteration, model routing and deployment.
Traces record each request's nested steps, including model calls, tool calls, retrieved context, errors, cost and latency, and teams filter, examine and debug them, with dashboards, alerts to Slack or webhooks, and Topics, which classifies production logs by task, sentiment and issue. Evals run an AI system over versioned datasets and score it with LLM-as-a-judge, autoevals, custom code or human reviewers; experiments are compared side by side, evals run in CI, and online scoring applies the same scorers to live traffic so failures become new test cases.
Braintrust also hosts functions: prompts, tools that deployed prompts can call, scorers and workflows that chain prompts, versioned and tagged by environment, plus a Gateway that routes model requests to providers on the customer's own keys. Loop, its built-in agent, investigates project data, writes scorers and test cases, iterates on prompts and runs on a schedule, pausing for approval before it changes anything in an interactive thread. SDKs cover Python, TypeScript, Go, Java, Ruby and C#, alongside a REST API, the bt CLI and an MCP server for coding agents.
Braintrust lists SOC 2 Type II compliance on every plan, supports SSO and SAML, OIDC, role-based access with custom groups and audit logs, and signs BAAs on Enterprise. It runs as SaaS, as BYOC in the customer's cloud operated by Braintrust, or self-hosted on Enterprise. Starter is free with unlimited users, Pro is $249 a month, both with usage-based overage, and Enterprise is custom.
Vendor details
Canonical URL
https://www.braintrust.dev
Category
Agent infrastructure
Company status
independent
Use cases & customers
Target customers
Deployment options
In practice
Your AI feature works in testing but quietly regresses in production and you find out from users. Braintrust scores live traffic with LLM-based judges, surfacing regressions as they happen and turning real failures into new test cases.
A prompt change looks better on one example but you can't tell if it's better overall. Braintrust runs your prompts and models against a dataset, scores them side by side, and captures an immutable experiment you can compare over time.
A bad change keeps slipping into your agent before anyone catches it. Braintrust runs your evals on every pull request and blocks the merge when a change degrades quality, gating releases on measured output quality.
Sources & related URLs
Agentic Index coverage score
10.0 / 14 capabilities · 71%
| Integrations & Tool Calling | Full |
|---|---|
|
Tools are hosted as functions, general-purpose code that deployed prompts can call to perform operations or access external data, running serverless in an isolated environment, and any function can be used as a tool; prompts can also connect to public MCP servers after OAuth authentication to reach external tools and data. That gives agents custom, actionable tools. SourceBraintrust, braintrust.dev/docs deploy/functions and evaluate/write-prompts (MCP servers)read 2026-09-21 |
|
| Workflow Orchestration | Partial |
|
Workflows are hosted functions that chain two or more prompts and pass outputs between them, alongside prompts, tools and scorers that compose into applications, and functions are versioned and can be tagged for production, staging or development environments. That sequencing is versioned and traceable. Branching, loops, retries and fallback paths, and deterministic nodes mixed with agent steps, are not documented. SourceBraintrust, braintrust.dev/docs deploy/functions, deploy/prompts and deploy/environmentsread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
Datasets hold test examples, traces record the context a customer's retrieval returned, and the documentation shows a RAG agent built from a customer-written vector search tool and a prompt. No document ingestion, index or retrieval layer that Braintrust maintains over company data is documented. SourceBraintrust, braintrust.dev/docs deploy/functions and annotate/datasetsread 2026-09-21 |
|
| Human Oversight & Guardrails | Full |
|
Loop, Braintrust's agent, runs read-only work without interrupting and pauses for approval before it creates or edits anything in an interactive thread; a scheduled Loop automation has nobody to ask, so each one carries write tool permissions that set what it may change, from none (read and report only) to named object types. Human review scores and review queues let people rate the customer's own traces. That puts an approval step before an agent's changes and sets different autonomy levels for scheduled and supervised runs. SourceBraintrust, braintrust.dev/docs loop, loop/automations and annotate/human-reviewread 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
Braintrust supports SSO and SAML with Okta Workforce, Microsoft Entra ID and Google Workspace and OIDC for custom providers, role-based access control with built-in and custom permission groups set at organization and project level, API keys stored as one-way hashes and scoped to projects, AES-256 encryption for provider secrets and encryption of all data at rest and in transit, configurable retention policies, and an audit log of administrative actions such as permission changes and API key creation. The pricing page lists SOC 2 Type II compliance on every plan and a BAA for HIPAA on Enterprise, with a trust center at trust.braintrust.dev. SourceBraintrust, braintrust.dev/docs security and admin/audit-logs, and braintrust.dev/pricingread 2026-09-21 |
|
| Observability & Auditability | Full |
|
Each AI request's nested steps, including model calls, tool calls, retrieved context, errors, cost and latency, are recorded in Braintrust traces, and teams examine, filter and debug them in the trace view; logs can be exported, including automatic export to the customer's own S3 bucket on Enterprise, retention is 14 days on Starter, 30 days on Pro with paid extension, and custom on Enterprise, and administrative actions go to a separate audit log. SourceBraintrust, braintrust.dev/llms.txt, braintrust.dev/docs observe, admin/data-management/export and admin/audit-logs, and braintrust.dev/pricingread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Loop's threads persist so a person can return to an investigation, and traces, datasets and experiments persist as records. No session, workflow or long-term memory that an agent reads and writes across runs, or controls to review and scope it, is documented. SourceBraintrust, braintrust.dev/docs loop and deploy/functionsread 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
Braintrust runs as SaaS, where it operates both the control plane and the data plane; as BYOC, where Braintrust operates the data plane inside the customer's cloud account; or self-hosted on Enterprise, where the customer deploys and controls the infrastructure that stores its AI data while Braintrust provides the managed UI and platform, with the data plane in the customer's own VPC. BYOC and self-hosted keep logs, traces, datasets and prompts in the customer's account and region. SourceBraintrust, braintrust.dev/docs admin/self-hosting, admin/deployment/byoc and securityread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
As its ready-made agent, Braintrust ships Loop, which investigates project data, bootstraps scorers, generates test cases, iterates on prompts and builds charts, and the same agent powers Patterns, whose preconfigured automation finds recurring failures, and the trace Debugger. That is one packaged agent with preset uses rather than a set of ready-made agents or templates a buyer selects among. SourceBraintrust, braintrust.dev/docs loop, loop/automations and observe/patterns, and braintrust.dev/pricingread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Loop automations run Braintrust's agent on a daily, weekly, interval or cron schedule with an instruction, a model and set write permissions, sending results to Slack or a webhook; online scoring runs scorers automatically on incoming production logs, and alerts fire on log, time-window and environment conditions. SourceBraintrust, braintrust.dev/docs loop/automations, evaluate/score-online and observe/alertsread 2026-09-21 |
|
| Model Flexibility & Routing | Full |
|
The Braintrust Gateway gives one API for routing model requests to AI providers, with the customer's own provider keys configured at organization level, workload identity federation for supported providers, custom providers, and Braintrust's built-in models billed against monthly credits; deployed prompts carry their own model configuration, and Loop runs on the customer's chosen model and provider. The customer controls model choice with its own providers and keys. SourceBraintrust, braintrust.dev/llms.txt and braintrust.dev/docs deploy/gateway, admin/ai-providers and loopread 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Braintrust publishes a REST API with an OpenAPI description, SDKs in Python, TypeScript, Go, Java, Ruby and C#, the bt CLI for running evals, querying logs, syncing data and managing functions, and an MCP server authenticated with OAuth 2.0 and PKCE for reasoning over project data from coding agents; custom tools, scorers and workflows are pushed from code as versioned functions, and evals run in CI. SourceBraintrust, braintrust.dev/llms.txt and braintrust.dev/docs security, deploy/functions and evaluate/run-in-ciread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
Evals run an AI system over versioned datasets and score outputs with LLM-as-a-judge, autoevals, custom code or human review scorers, experiments are compared side by side, evals run in CI, and online scoring applies the same scorers to production logs so failures become new dataset examples. That tests on datasets before production, gates quality with configurable checks, scores quality over time and closes the loop after deployment. SourceBraintrust, braintrust.dev/docs evaluate, run-in-ci, score-online and annotate/datasetsread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
Custom code functions run in isolated serverless environments, and the documentation covers tracing, evaluation, deployment and administration, but no browser, desktop or computer control by an agent is documented. SourceBraintrust, braintrust.dev/docs security and llms-full.txtread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Braintrust announced Nitro, a new asynchronous query engine for Brainstore that separates storage reads from compute work to accelerate trace investigations. Nitro is automatically enabled for SaaS and BYOC customers, with support starting at dataplane 2.15 for deployments customers host themselves. Braintrust reports that full text searches ran more than twice as fast in its production comparison.
Bears on: Observability / auditability
View sourceBraintrust introduced support for tracing Cloudflare Agents. Users can ingest traces by exporting them via OpenTelemetry or by instrumenting their agents directly in JavaScript on the Cloudflare Workers runtime.
Bears on: Integrations
View sourceBraintrust released terraform-aws-braintrust-data-plane v5.5.0, updating Braintrust Services/API ECS to v2.2.1, changing the default Redis instance type from cache.t4g.medium to cache.r7g.large, and enabling Brainstore fast readers by default.
Bears on: Deployment / data residency
View sourcePricing
Free Starter (1 GB data, 10K scores) · Pro $249/mo
usage
Included quota
Starter: 1 GB processed data and 10K scores a month, 14 day retention, unlimited seats. Pro: 5 GB and 50K scores, 30 day retention, Google SSO. No per seat fees on either plan.
What is public
Braintrust publishes full self serve pricing: free Starter with 1 GB processed data and 10K scores a month, Pro at $249/mo with 5 GB and 50K scores plus published overage rates, and custom Enterprise with BYOC and self hosted deployment. Qualifying early stage startups can get 6 to 12 months of free Pro.
Billing mechanics
Flat monthly platform fee per plan plus metered on demand usage beyond included processed data and scores. Pro overage rates are lower than Starter's. Topics and built in model usage draw from a separate monthly credit with uniform overage rates.
Cost watchouts
Usage past the included amounts is billed with no stated cap: $4/GB and $2.50 per 1K scores on Starter, $3/GB and $1.50 per 1K scores on Pro, model use past the monthly credits at token rates, and retention beyond 30 days on Pro at $0.50 per GB a month. SAML SSO, custom roles, audit logs, the BAA and self-hosting are Enterprise only.
Variable cost rationale
Platform fee is flat but processed data and score volume are metered with uncapped overage, so cost tracks trace volume rather than seats.
Additional watchouts
Model the overage math before committing: a Pro team logging 10 GB and 100K scores a month lands near $339, not $249.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free Starter plan, no card: 1 GB processed data, 10K scores a month, 14-day retention, unlimited users, $10 model credits a month. Qualifying startups can get 6 to 12 months of Pro free.
Lowest paid plan
Pro $249/mo (5 GB processed data, 50K scores, 30 day retention)
Commercial notes
Unlimited users on every plan, so cost follows processed data, scores and model use rather than seats. Qualifying startups can get 6 to 12 months of Pro free.
Key ambiguities
Enterprise pricing is not published, and model use past the included credits is billed at token rates listed on a separate detailed pricing page.
Missing data
Enterprise pricing unpublished.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Braintrust
The closest documented capability profiles to Braintrust among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- LangWatch10.0 / 14Matches Braintrust across all 14 documented capabilities
- F5 AI Guardrails9.0 / 14A lighter documented profile than Braintrust
- Galileo9.0 / 14A lighter documented profile than BraintrustBraintrust vs Galileo →
- Opik9.0 / 14A lighter documented profile than BraintrustBraintrust vs Opik →
- W&B Weave9.0 / 14A lighter documented profile than BraintrustBraintrust vs W&B Weave →
- Bernstein10.5 / 14Adds documented Memory & State Persistence
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded