LangWatch
Also known as: Langwatch
Open source platform for agent observability, evaluation, and scenario based simulation, with OpenTelemetry tracing and a built in AI gateway.
LangWatch is an open source platform, licensed under Apache 2.0, for observing, evaluating, and testing large language model powered agents. It combines production observability, continuous evaluation, and scenario based simulation in one loop, trace to dataset to evaluate to optimize to re test, so teams can understand agent behavior across real workflows and systematically improve reliability, performance, and cost. Tracing is OpenTelemetry native and framework and provider agnostic, with integrations for LangChain, LangGraph, CrewAI, the Vercel AI SDK, Mastra, and Google ADK, and support for using LangWatch through MCP clients such as Claude Desktop.
A distinguishing feature is agent simulation. LangWatch can run multi turn conversations against an agent in parallel, score them with a judge rubric on measures like faithfulness and policy adherence, and turn failures into pull requests, letting product managers own the specification while developers stay in flow. Domain experts can review runs, annotate failures, and label edge cases through annotation queues, and prompts live in Git through a GitHub integration with prompt versions linked to traces.
LangWatch also ships an AI Gateway, an OpenAI and Anthropic compatible proxy with virtual keys, hierarchical budgets, inline guardrails, and automatic fallback across providers, adding governance and cost control with roughly seven hundred nanoseconds of hot path overhead. The full platform is open source and self hosts through Docker or a Helm chart on Kubernetes, with cloud specific setups for AWS, Google Cloud, and Azure and hybrid options for teams with data residency requirements. Pricing keeps a free tier with two hundred thousand events a month, charges around twenty nine euro per seat with unlimited lite seats for sharing results, and bills additional usage at one dollar per one hundred thousand events, with the open source build free to self host.
Vendor details
Canonical URL
https://langwatch.ai
Category
Agent infrastructure
Subcategory
Observability and evaluation
Funding status
Independent. Open source under Apache 2.0 with an active project used by thousands of AI developers. Offers a managed cloud alongside self hosted, on premise, and hybrid deployment.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
OpenTelemetry native and framework and provider agnostic, with integrations for LangChain, LangGraph, CrewAI, the Vercel AI SDK, Mastra, and Google ADK. Usable through MCP clients like Claude Desktop, with a GitHub integration for prompt versioning and an OpenAI and Anthropic compatible gateway reaching hundreds of models.
In practice
Your agent passes eyeball checks but breaks in production. LangWatch runs multi turn simulations against it in parallel, scores them with a judge rubric, and turns failures into pull requests before launch.
A product manager needs to test agent behavior without code. They write the goal in plain English, LangWatch generates the scenario plan and rubric, and developers stay in flow while nothing slips through.
You want tracing plus cost control in one place. LangWatch gives OpenTelemetry native traces and an OpenAI and Anthropic compatible gateway with virtual keys, budgets, and provider fallback, self hosted if needed.
Sources & related URLs
Agentic Index coverage score
6.5 / 14 capabilities · 46%
| Integrations & Tool CallingOpenTelemetry native, framework agnostic, integrations for LangChain, LangGraph, CrewAI, Mastra, Google ADK, docs 2026-07-06 | Full |
|---|---|
| Workflow OrchestrationObserves and tests agents but is not an orchestrator | Unable to verify |
| Knowledge Grounding & RAGEvaluates RAG apps but provides no knowledge grounding | Unable to verify |
| Human Oversight & GuardrailsInline gateway guardrails plus human annotation and review queues, docs 2026-07-06 | Partial |
| Security, Identity & GovernanceGateway virtual keys and budgets plus self host and on prem for residency; enterprise identity less documented, docs 2026-07-06 | Partial |
| Observability & AuditabilityCore product: OpenTelemetry native tracing of prompts, tool calls, and agent behavior, docs 2026-07-06 | Full |
| Memory & State PersistenceNo agent memory layer | Unable to verify |
| Deployment & Data ResidencyCloud, self host via Docker or Helm, on prem for AWS, GCP, Azure, and hybrid for residency, docs 2026-07-06 | Full |
| Prebuilt Agents, Templates & PacksPrebuilt evaluators and scenario scaffolding exist but no prebuilt agents | Unable to verify |
| Triggers & Channel CoverageNo production event triggers or channel coverage | Unable to verify |
| Model Flexibility & RoutingProvider agnostic with a bundled gateway offering fallback across providers, secondary to the core product | Partial |
| APIs, SDKs & MCP ExtensibilityOpenTelemetry native, SDK, API, MCP support for Claude Desktop, webhooks, docs 2026-07-06 | Full |
| Testing, Debugging & OptimizationCore capability: evaluations plus multi turn agent simulation and scenario testing, docs 2026-07-06 | Full |
| Browser & Computer UseNo browser or computer use | Unable to verify |
The Agentic Index coverage score grades every vendor Full, Partial or Unable to verify against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
LangWatch introduced Langy, an automated AI engineering agent embedded directly inside the platform. Langy analyzes production traces to identify and cluster recurring behavior patterns, automatically writes scenario tests and evaluations, and opens pull requests on connected GitHub repositories to propose code fixes.
Bears on: Observability / auditability
View sourcePricing
Open source self host free; free cloud (200k events/mo); about $31 (29 euro) per seat plus usage
seats plus events
Included quota
Open source self host is free with the full feature set. Free cloud tier includes 200,000 events a month. Paid cloud is about 29 euro per seat, with unlimited lite seats for sharing results, plus usage at one dollar per 100,000 additional events.
What is public
The free event allowance, per seat euro price, and event overage rate are published.
Billing mechanics
Free open source self host, or cloud with a free 200,000 event tier and paid seats at about 29 euro each (unlimited lite seats for viewers), plus event based usage at one dollar per 100,000 events.
Cost watchouts
Events accrue across traces, evaluations, and simulations, so heavy agent testing and production volume can move you past the 200,000 event free allowance faster than expected, billed at one dollar per 100,000 events. Full seats (beyond unlimited lite seats) add per seat cost.
Variable cost rationale
A per seat floor is predictable, but event based usage across traces, evaluations, and simulations scales with testing and production volume, and self hosting shifts cost to your own infrastructure.
Overage / add-ons
Additional events above the included allowance bill at one dollar per 100,000 events.
Sales call required
Mixed (some tiers require a call)
Free / trial
Apache 2.0 open source free to self host, plus a free cloud tier of 200,000 events a month
Lowest paid plan
About 29 euro per seat plus usage, or free via self hosted open source
Commercial notes
Independent, open source under Apache 2.0, used by thousands of AI developers. Distinctive for agent simulation and a bundled AI gateway.
Key ambiguities
Seat pricing is quoted in euro (about 29 euro), so the dollar figure varies with exchange rate; enterprise terms may differ.
Related vendors
- Acrab — Singapore compute infrastructure company building a full stack…
- AgentOps — Agent observability and reliability platform with broad model and…
- Agno — High-performance agent runtime and framework (formerly Phidata) with…
- AIsa — Unified resource and payment gateway for AI agents that lets them…
- AlphaBitCore — AI control plane that governs how models, agents, tools, and…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
Alternatives to LangWatch
The closest documented capability profiles to LangWatch among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Opik7.0 / 14Fuller documented coverage on Security, Identity & Governance
- W&B Weave7.0 / 14Fuller documented coverage on Security, Identity & Governance
- Fiddler AI7.5 / 14Fuller documented coverage on Human Oversight & Guardrails and Security, Identity & Governance
- Langfuse7.5 / 14Fuller documented coverage on Human Oversight & Guardrails and Model Flexibility & RoutingLangWatch vs Langfuse →
- AgentOps6.0 / 14Fuller documented coverage on Model Flexibility & Routing
- Braintrust8.0 / 14Adds documented Knowledge Grounding & RAG
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded