Langfuse
Open-source AI engineering platform for tracing, observability, prompt management, and evaluation, self-hostable or SaaS, with an OSS community, free plan, and model-agnostic coverage.
Langfuse is an open-source AI engineering platform for tracing, evaluating and improving LLM applications and agents. Hierarchical traces capture every LLM call, tool invocation and retrieval step as trace trees and agent graphs, with session and user tracking, cost and latency, dashboards and alerts. On top of tracing it adds evaluation (datasets, experiments compared against baselines, LLM-as-a-judge and code evaluators run online or offline, human annotation queues and CI checks that block regressions) and prompt management with versioning, protected deployment labels and a playground that compares models on the customer's own LLM connections.
Langfuse works with any stack through native Python and TypeScript SDKs, OpenTelemetry and more than 100 integrations, and exposes a public API, a CLI and a platform MCP server for IDE agents. Langfuse Cloud runs in US, EU or JP data regions with a HIPAA-ready region and AWS PrivateLink, and the MIT-licensed core self-hosts on Docker, Kubernetes or Terraform. It states SOC 2 Type II and ISO 27001, with RBAC, Enterprise SSO, SCIM and audit logs on higher plans. Langfuse has joined ClickHouse and continues to sell under its own name, from a free Hobby plan through Core at $29 a month, Pro at $199 and Enterprise at $2,499.
Vendor details
Canonical URL
https://langfuse.com
Category
Agent infrastructure
Company status
acquired
Use cases & customers
Target customers
Deployment options
In practice
Your agent works in demos but fails unpredictably in production and you can't see why. Langfuse traces every prompt, response, tool call, cost, and latency, so you can reconstruct exactly what the agent did and where it went wrong.
You need to run LLM observability on your own infrastructure for reasons. Langfuse is open source and self-hostable at production scale, with broad SDK and OpenTelemetry support so it fits your stack rather than locking you in.
You can't tell whether a prompt change actually improved quality. Langfuse lets you build datasets from real traffic, score outputs with model-based or human evaluation, and track quality over time to catch regressions before users do.
Sources & related URLs
Agentic Index coverage score
8.5 / 14 capabilities · 61%
| Integrations & Tool Calling | Partial |
|---|---|
|
Over 100 integrations with agent frameworks, model providers and gateways send traces into Langfuse, data exports to PostHog, Mixpanel and blob storage, and prompt changes notify through webhooks and Slack. These move telemetry in and data out; no connectors that let an agent take authenticated actions in outside systems are documented. SourceLangfuse, langfuse.com homepage and pricingread 2026-09-21 |
|
| Workflow Orchestration | Not documented |
|
For agents that run elsewhere, Langfuse provides tracing, evaluation and prompt management, and prompt composability links prompts together. No workflows that sequence, branch or retry an agent's steps, or mix deterministic nodes with agent steps, are documented. SourceLangfuse, langfuse.com homepage and pricingread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
Langfuse stores traces, prompts and evaluation datasets, and no retrieval structure over the customer's documents or knowledge that grounds an agent's answers is documented; it evaluates the grounding of RAG systems built elsewhere rather than providing one. SourceLangfuse, langfuse.com pricing and llms.txtread 2026-09-21 |
|
| Human Oversight & Guardrails | Partial |
|
Protected deployment labels restrict who can promote a prompt version to production, and CI checks can require an approved baseline before a change ships, so changes to an agent's behavior pass a human-controlled release gate at the policy layer. No approval step, escalation rule or pause before an agent acts at run time is documented. SourceLangfuse, langfuse.com pricing and llms.txtread 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
Reports for Langfuse's SOC 2 Type II and ISO 27001 attestations are available from the Pro plan, and there is a HIPAA-ready region and a GDPR DPA. Access control covers organization- and project-level RBAC, sign-in with Google, Azure AD or GitHub, Enterprise SSO with Okta or Entra ID and SSO enforcement on the Teams add-on, SCIM provisioning, audit logs, client-side data masking and data retention management. SourceLangfuse, langfuse.com pricing and homepageread 2026-09-21 |
|
| Observability & Auditability | Full |
|
Hierarchical traces capture every LLM call, tool invocation and retrieval step of an agent run, shown as trace trees and agent graphs with session and user tracking, token cost and latency, and filterable by user, session, cost, latency or metadata; audit logs are separate on Enterprise, data can be exported by batch or scheduled to blob storage and to PostHog or Mixpanel, and historical data access runs 30 days on Hobby, 90 days on Core and 3 years on Pro. That gives run level visibility of steps and tool calls. SourceLangfuse, langfuse.com homepage and pricingread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Sessions, users and traces are recorded for inspection, and no session, conversation or long-term memory that an agent reads back as context is documented. SourceLangfuse, langfuse.com pricing and homepageread 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
Langfuse Cloud offers a choice of US, EU or JP data regions, a HIPAA-ready region and AWS PrivateLink, and the MIT-licensed platform self-hosts with Docker Compose, Kubernetes via Helm, or Terraform on AWS, GCP and Azure, with paid Enterprise additions for self-hosted deployments. SourceLangfuse, langfuse.com pricing, homepage and llms.txtread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
The in-app Langfuse Assistant is the one ready-made agent shipped, and a Langfuse skill packages best-practice workflows for instrumentation, prompt management and API access into a customer's coding agent. Beyond those two, there is no set of prebuilt agents or templates for a buyer to select from. SourceLangfuse, langfuse.com llms.txt and pricingread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Online evaluation runs LLM-as-a-judge and code evaluators on production observations as they arrive, and monitors and alerts fire on metric and evaluator thresholds, so Langfuse's evaluation work starts from incoming telemetry with no person initiating each run; prompt changes also emit webhooks. SourceLangfuse, langfuse.com homepage and pricingread 2026-09-21 |
|
| Model Flexibility & Routing | Full |
|
Customers configure their own LLM connections, which the Playground uses to test prompts on real production inputs and compare models side by side, and which power the LLM-as-a-judge evaluators the customer sets up. The customer brings its own providers and chooses the models Langfuse's own LLM features run on. SourceLangfuse, langfuse.com homepage, pricing and llms.txt (LLM Connections)read 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Langfuse documents an extensive public API with published rate limits by plan, native Python and TypeScript SDKs, OpenTelemetry ingestion for Java, Go and other languages, a CLI, a platform MCP server that lets IDE agents manage prompts and query traces, stable ID-based APIs for managing evaluators, and an agent skill. SourceLangfuse, langfuse.com homepage, pricing and llms.txtread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
Evaluation of the customer's application runs online and offline, covering datasets, experiments from the SDK or UI compared side by side against baselines, LLM-as-a-judge and code evaluators managed by API, custom scores, user feedback and human annotation queues, alerts on evaluator results, and CI checks that block regressions against thresholds and approved baselines. That tests the customer's agent with datasets before production and scores output quality over time. SourceLangfuse, langfuse.com llms.txt, homepage and pricingread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
Langfuse observes and evaluates agents, and no browser, desktop or computer control by an agent is documented. SourceLangfuse, langfuse.com homepageread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Langfuse introduced stable, ID-based public APIs for managing LLM-as-a-Judge and code evaluators. Users can now programmatically create, version, and manage both evaluators and their associated evaluation rules.
Bears on: Observability / auditability
View sourceLangfuse introduced and defaulted to background execution for in-app agents, including client and web support. This changes the default operational path for these agents to run asynchronously.
Bears on: Agent capability
View sourceLangfuse added the ability to perform bulk exports and imports of prompts.
Bears on: Observability / auditability
View sourcePricing
Free Hobby (50K units/mo) · Core $29/mo · open source self host
usage
Included quota
A billable unit is one trace, observation, or score. Hobby: 50K units a month, 2 users. Core: 100K units, unlimited users, 90 day data access, 4K req/min ingestion. Pro: 3 year retention, 20K req/min, compliance certifications. All paid plans share graduated $8 per 100K unit overage.
What is public
Langfuse publishes complete self serve pricing: free Hobby at 50K units a month, Core $29/mo, Pro $199/mo, Enterprise $2,499/mo, uniform $8 per 100K overage, and a free MIT licensed self host path with all product features open sourced since June 2025.
Billing mechanics
Flat plan fee plus graduated usage overage on billable units (traces, observations, scores). Tokens are tracked for LLM cost analysis but do not count as billable units. Unlimited users on all paid tiers.
Cost watchouts
Each trace, observation and score draws on the unit allowance, so multi-step agents use several units per request; Enterprise SSO, SSO enforcement and fine-grained RBAC sit behind the $300 a month Teams add-on on top of Pro.
Variable cost rationale
Unit metering scales with trace volume, but graduated overage at $8 per 100K units keeps cost growth predictable and there are no per seat fees.
Additional watchouts
Buyers should watch post acquisition roadmap and licensing signals under ClickHouse ownership, though all public commitments to open source have held since the January 2026 close.
Sales call required
No, self serve available
Free / trial
Free Hobby plan, no card (50K units/mo, 2 users)
Lowest paid plan
Core $29/mo (100K units, unlimited users, 90 day data access)
Commercial notes
Langfuse has joined ClickHouse and continues as its own product, with an MIT-licensed core for self-hosting and paid Enterprise additions. Its homepage cites use by 21 of the Fortune 50 and about 35,000 GitHub stars.
Key ambiguities
Enterprise custom volume pricing requires annual commitment. Self hosting is free on the MIT core but ClickHouse operations are real work.
Missing data
Enterprise custom volume rates unpublished.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Langfuse
The closest documented capability profiles to Langfuse among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Confident AI8.5 / 14Matches Langfuse across all 14 documented capabilitiesLangfuse vs Confident AI →
- HoneyHive8.5 / 14Matches Langfuse across all 14 documented capabilitiesLangfuse vs HoneyHive →
- Arize AI9.0 / 14Fuller documented coverage on Prebuilt Agents, Templates & PacksLangfuse vs Arize AI →
- F5 AI Guardrails9.0 / 14Fuller documented coverage on Human Oversight & Guardrails
- Freeplay8.0 / 14A lighter documented profile than Langfuse
- Galileo9.0 / 14Fuller documented coverage on Human Oversight & GuardrailsLangfuse vs Galileo →
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded