LangWatch
Also known as: Langwatch
Open source platform for agent tracing, simulation-based agent testing and evaluation, with an AI gateway for routing and governance and Langy, an automated AI engineer that opens pull requests.
LangWatch is an open source platform for tracing, testing, routing and governing the LLM calls a company makes, from its own agents to the coding assistants its engineers use. Tracing is OpenTelemetry native across Python, TypeScript, Go and Java SDKs and integrations for major frameworks, providers and coding agents.
Its distinguishing feature is agent testing: scenarios in which a simulated user acts out a situation against the agent while a judge scores the conversation, grouped into test suites, run from the platform or in code, compared across agents, extended to voice agents and red teaming, and used as the quality gate when optimizing prompts and tool contracts.
Around that sit experiments over datasets, online monitors that score production traffic as it arrives, inline guardrails that block or modify responses, annotation queues, alerts and automations, and visual workflows that combine prompts, Python code, HTTP calls and evaluators and can be published as evaluators, agents or APIs. Langy, LangWatch's automated AI engineer, reads traces and evals, writes Scenario tests and opens pull requests through a GitHub App for a person to review.
The LangWatch AI Gateway gives one OpenAI- and Anthropic-compatible endpoint across a dozen providers, with virtual keys, routing policies and fallback order, budgets and rate limits, and governance features such as personal keys for coding assistants and OCSF export to a SIEM. LangWatch supports SSO, SCIM provisioning, custom roles, an audit log and per-scope retention, and states ISO 27001 certification and GDPR compliance. It runs as managed cloud in the EU, US, UK or APAC, self-hosted or in a hybrid setup. The Developer plan is free, Growth is €29 per core seat a month with events billed beyond the included amount, and Enterprise is custom.
Vendor details
Canonical URL
https://langwatch.ai
Category
Agent infrastructure
Subcategory
Observability and evaluation
Funding status
Independent. Open source under Apache 2.0 with an active project used by thousands of AI developers. Offers a managed cloud alongside self hosted, on premise, and hybrid deployment.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
OpenTelemetry native and framework and provider agnostic, with integrations for LangChain, LangGraph, CrewAI, the Vercel AI SDK, Mastra, and Google ADK. Usable through MCP clients like Claude Desktop, with a GitHub integration for prompt versioning and an OpenAI and Anthropic compatible gateway reaching hundreds of models.
In practice
Your agent passes eyeball checks but breaks in production. LangWatch runs multi turn simulations against it in parallel, scores them with a judge rubric, and turns failures into pull requests before launch.
A product manager needs to test agent behavior without code. They write the goal in plain English, LangWatch generates the scenario plan and rubric, and developers stay in flow while nothing slips through.
You want tracing plus cost control in one place. LangWatch gives OpenTelemetry native traces and an OpenAI and Anthropic compatible gateway with virtual keys, budgets, and provider fallback, self hosted if needed.
Sources & related URLs
Agentic Index coverage score
10.0 / 14 capabilities · 71%
| Integrations & Tool Calling | Full |
|---|---|
|
Langy, LangWatch's agent, changes code through bot-authored pull requests opened by a GitHub App, and LangWatch workflows, which can be published as agents or API endpoints, call external systems through HTTP nodes and code blocks that reference encrypted project secrets; the AI Gateway governs which tools and servers a virtual key may reach. SourceLangWatch, langwatch.ai/docs/llms.txt (Langy pull requests, workflows, secrets, routing policies)read 2026-09-21 |
|
| Workflow Orchestration | Partial |
|
A LangWatch workflow is a graph of nodes that LangWatch runs from an Entry point to an End node, mixing LLM prompt nodes with Python code, HTTP and evaluator nodes; versions are saved and restorable, and a workflow can be published as an evaluator, as an agent under test, or as an API the customer's code calls. Loops, branching, retries and fallback paths are not documented. SourceLangWatch, langwatch.ai/docs/workflows/overview and llms.txt (building a workflow, workflow as agent); langwatch.ai/docs/workflows/overview.mdread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
Built-in RAG evaluators score retrieval, and cookbooks on embedding tuning and vector versus hybrid search cover the customer's own pipeline. No document ingestion, index, retrieval layer or knowledge API that grounds an agent in company data is documented as part of LangWatch. SourceLangWatch, langwatch.ai/docs/llms.txt (built-in evaluators, cookbooks)read 2026-09-21 |
|
| Human Oversight & Guardrails | Full |
|
Guardrails run evaluators inline to block or modify harmful responses in real time, the AI Gateway enforces budgets, rate limits and routing policies, including which tools and servers a key may reach, before a request goes through, and Langy's code changes arrive as pull requests that a person reviews before merge. SourceLangWatch, langwatch.ai/docs/llms.txt (guardrails, AI Gateway budgets and routing policies, Langy pull requests)read 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
Documented controls include single sign-on with directory provisioning and sign-in policies, SCIM 2.0 provisioning from Okta, Entra ID and others with group-to-role mapping, role-based access with custom roles and role bindings at organization, team and project level, an audit log of every change made through settings, the API and the AI Gateway with who, where and what changed, OCSF export to a SIEM, retention set per organization, team or project, and encrypted write-only secrets; the homepage states ISO 27001 certification and GDPR compliance, with a Vanta-monitored trust center. SourceLangWatch, langwatch.ai/docs/llms.txt (SSO, SCIM, RBAC, audit log, data retention, secrets, OCSF export), langwatch.ai homepage and pricingread 2026-09-21 |
|
| Observability & Auditability | Full |
|
LangWatch traces LLM calls, tool executions and retrieval steps through OpenTelemetry-native SDKs and integrations, including coding assistants such as Claude Code and Codex; an audit log records every change through settings, the API and the gateway, governance events export to a SIEM as OCSF records, analytics export into the customer's dashboards, and retention defaults are published per plan and adjustable per organization, team or project. SourceLangWatch, langwatch.ai/docs/llms.txt (observability, audit log, OCSF export, data retention) and langwatch.ai/docs/pricingread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Traces, threads and datasets are stored as records of the customer's agent, and each Langy conversation is kept in its own sandboxed worker. No session, workflow or long term memory that an agent reads and writes across runs is documented. SourceLangWatch, langwatch.ai/docs/llms.txt (concepts, datasets, how Langy works)read 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
Customers can run LangWatch as managed multi-tenant SaaS in the EU, US, UK or APAC, self-hosted with Docker, Kubernetes and Helm or in the customer's VPC, or hybrid with the data plane on the customer's infrastructure and the control plane on LangWatch's; the open source repository carries every feature for local use, and an Enterprise license adds longer retention to self-hosted deployments. SourceLangWatch, langwatch.ai homepage and langwatch.ai/docs/llms.txt (self-hosting overview, editions and licensing, hybrid setup)read 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
LangWatch ships Langy, an automated AI engineer that reads traces and evals, writes Scenario tests, optimizes prompts and opens pull requests, and a Skills Directory of installable skills, including skills for product managers and domain experts, that set a coding assistant up to work with LangWatch. That is one packaged agent and a set of setup skills, not a set of ready-made agents or workflow templates a buyer selects among for their own work. SourceLangWatch, langwatch.ai/docs/llms.txt (Langy, skills directory) and langwatch.ai homepageread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Online evaluation monitors run an evaluator on every matching trace or thread as it arrives, alerts and automations detect regressions and notify teams, automations add every new trace matching a filter to a dataset, and Langy can set up alerts and automations from what it finds across traces. So scoring and actions start on incoming traffic without a person asking. SourceLangWatch, langwatch.ai/docs/llms.txt (online evaluation, alerts and automations, datasets from traces, Langy insights)read 2026-09-21 |
|
| Model Flexibility & Routing | Full |
|
The LangWatch AI Gateway gives one OpenAI- and Anthropic-compatible endpoint across providers including OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, Vertex AI, Gemini, xAI, Groq, Cerebras and DeepSeek, with provider credentials held in LangWatch, a virtual key per application, routing policies that set which providers a key routes through, the fallback order and model tiers, and budgets and rate limits enforced before the request; model providers are stored once per organization, team or project for evaluators, workflows and Langy. SourceLangWatch, langwatch.ai/docs/llms.txt (AI Gateway, routing policies, budgets, providers, model providers)read 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
LangWatch publishes a REST API (including the Workflows and Instant Evals APIs), SDKs across Python, TypeScript, Go and Java, the langwatch CLI that drives the platform from the terminal for people and coding assistants, an MCP server that gives coding assistants access to traces, tests and evaluations, installable skills, personal and service API keys for development and CI/CD, and CI integration that gates merges on eval results. SourceLangWatch, langwatch.ai/docs/llms.txt (CLI, MCP server, API keys, API reference) and langwatch.ai/pricingread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
Experiments run over datasets with built-in, saved and workflow-built evaluators. LangWatch also tests agents with multi-turn scenarios in which a simulated user acts against the agent and a judge scores the conversation, groups scenarios into test suites, compares agents side by side on pass rate, cost and latency, gates merges on eval results in CI, scores production traffic with online monitors, and optimizes prompts and tool contracts with the scenario suite as the quality gate. SourceLangWatch, langwatch.ai/docs/llms.txt (agent testing, evaluations, improve your agent) and langwatch.ai/pricingread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
Tracing, agent testing, evaluation, the AI Gateway, governance, Langy and self-hosting make up the documented product, and no browser, desktop or computer control by an agent is documented; Langy works in a sandbox through the CLI and code, not by controlling a browser or desktop. SourceLangWatch, langwatch.ai/docs/llms.txtread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Instant Evals now live in LangWatch's Trace Explorer search bar. A question no filter can express, such as whether the user sounds frustrated, becomes an LLM judged filter chip that combines with the others. Runs under $0.50 start on their own, and bigger ones ask first.
Bears on: Observability / auditability
View sourceLangWatch introduced Langy, an automated AI engineering agent embedded directly inside the platform. Langy analyzes production traces to identify and cluster recurring behavior patterns, automatically writes scenario tests and evaluations, and opens pull requests on connected GitHub repositories to propose code fixes.
Bears on: Observability / auditability
View sourcePricing
Open source free · Developer free · Growth €29/core seat/mo + usage · Enterprise custom
Per core seat plus events beyond the included amount, plus storage beyond the included retention
Included quota
Open source self host is free with the full feature set. Free cloud tier includes 200,000 events a month. Paid cloud is about €29 per seat, with unlimited lite seats for sharing results, plus usage at $1 per 100,000 additional events.
What is public
The free event allowance, per seat euro price, and event overage rate are published.
Billing mechanics
Free open source self host, or cloud with a free 200,000 event tier and paid seats at about €29 each (unlimited lite seats for viewers), plus event based usage at $1 per 100,000 events.
Cost watchouts
Evaluation results and every experiment row and simulation turn count as billable events alongside spans, so heavy testing moves past the included 200k events faster than traces alone; events beyond it cost €5 per 100k and storage beyond 30 days €3 per GB.
Variable cost rationale
Core seats are predictable, but events accrue across traces, evaluation results, experiment rows and simulation turns, so testing and production volume both drive the usage bill.
Overage / add-ons
Additional events above the included allowance bill at $1 per 100,000 events.
Sales call required
Mixed (some tiers require a call)
Free / trial
Open source, free to self host; Developer plan free forever (50k events a month, 14-day data access, 2 users, 3 scenarios, simulations and custom evals), no card; Growth has a free trial
Lowest paid plan
Growth at €29 per core seat a month plus usage
Commercial notes
Open source, with every feature in the self-hostable repository; the cloud plans charge per core seat with unlimited lite users for reviewers and stakeholders, and events cover traces, evaluations, experiments and simulations.
Key ambiguities
The pricing page includes 30-day retention on Growth, while the documentation's pricing page gives every plan a 49-day default; seat prices are quoted in euro with a dollar toggle.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to LangWatch
The closest documented capability profiles to LangWatch among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Braintrust10.0 / 14Matches LangWatch across all 14 documented capabilities
- F5 AI Guardrails9.0 / 14A lighter documented profile than LangWatch
- Galileo9.0 / 14A lighter documented profile than LangWatch
- Opik9.0 / 14A lighter documented profile than LangWatch
- W&B Weave9.0 / 14A lighter documented profile than LangWatch
- Bernstein10.5 / 14Adds documented Memory & State Persistence
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded