Patronus AI
Also known as: Patronus, Patronus API
Evaluation, tracing and agent debugging platform with its own evaluator models, Experiments, production logs and traces, and the Percival agent debugger, sold through a free Developer plan and a per-call evaluator API.
Patronus AI sells a platform for evaluating, tracing and debugging AI applications and agents. Teams run Experiments that score prompts, models and data configurations against Datasets, compare results over time, and record production Logs and Traces.
Its evaluator API scores outputs with Patronus's own small and large evaluator models, including the Lynx hallucination detection model and the GLIDER judge, returns explanations on request, and is priced per call.
Percival, an agent built by Patronus, reads a customer's agent traces, detects more than twenty failure modes such as planning mistakes and incorrect tool use, and suggests fixes, with integrations for smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph and crewAI.
The Developer plan is free with no credit card and includes $10 in API credits; Enterprise adds on-prem or dedicated VPC deployment, SSO, custom data retention, webhooks, higher rate limits and custom evaluator fine-tuning. The company's homepage now leads with its research lab, which builds Digital World Models and reinforcement learning environments for training frontier models, and its platform documentation now requires a login.
Vendor details
Canonical URL
https://www.patronus.ai
Category
Agent infrastructure
Subcategory
Evaluation and guardrails
Funding status
Founded in 2023 in San Francisco by former Meta AI (FAIR) researchers Anand Kannappan (CEO) and Rebecca Qian (CTO). Raised a $17M Series A and a $50M round in 2026 to expand its agent simulation platform. Customers include AngelList, Pearson, HP, and Fortune 500 companies in finance, healthcare, and legal, with partners including NVIDIA, MongoDB, and IBM. Independent.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
Embedded into application code through a programming language agnostic API and Python SDK, with custom LLM judges, webhooks, and a web dashboard for logs and experiments. Evaluates RAG and agent pipelines, and ships open research models including the Lynx hallucination detector and the GLIDER judge.
In practice
Your RAG chatbot sometimes states facts not in the source documents. You wire Patronus Lynx in as a real time guardrail, and it flags the hallucinated span before the answer reaches the user.
Your agent fails intermittently across multi step traces and you cannot tell why. Percival inspects the execution traces, identifies which of twenty plus failure modes occurred, and suggests prompt and workflow fixes.
You are choosing between two prompts and three models for a regulated use case. Patronus Experiments runs them side by side against your criteria so you can pick the configuration that scores best before shipping.
Sources & related URLs
Related / legacy domains
Agentic Index coverage score
6.5 / 14 capabilities · 46%
| Integrations & Tool Calling | Partial |
|---|---|
|
Trace capture works with smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph, crewAI and custom clients, feeding a customer's agent traces to Percival, and Enterprise adds webhooks. These bring the customer's agent data in and send results out; the customer's agent reaches its own tools, and no integration through which an agent reads or writes a real system is described. SourcePatronus AI, patronus.ai/percival and patronus.ai/pricingread 2026-09-21 |
|
| Workflow Orchestration | Not documented |
|
Customers build and run their agents in frameworks such as LangGraph, crewAI and the OpenAI Agents SDK, and Patronus evaluates and debugs them there. It offers no sequencing, branching, retries or routing of agent steps of its own. SourcePatronus AI, patronus.ai/percival and patronus.ai/pricingread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
The Lynx model checks retrieval augmented answers for hallucination, and Patronus scores a customer's agents. No document ingestion, index, retrieval layer or knowledge API that grounds an agent in company data is offered. SourcePatronus AI, patronus.ai and patronus.ai/pricingread 2026-09-21 |
|
| Human Oversight & Guardrails | Not documented |
|
People confirm issues and annotate missed failure modes in agent traces on the Percival page, and the homepage presents the GLIDER evaluator as suited to guardrails. The public pages do not say whether Patronus provides a guardrail, approval step or checkpoint that holds or blocks an agent's output, rather than a verdict the customer's code acts on. SourcePatronus AI, patronus.ai/percival and patronus.ai; docs.patronus.ai/docsread 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
SSO comes on the Enterprise plan, with identity integration through the customer's own provider, alongside custom data retention and on premises or dedicated VPC deployment. The pricing page carries AICPA SOC, TISAX and HIPAA badges as an asserted attestation; no report, type or auditor is named and no trust page is linked. SourcePatronus AI, patronus.ai/pricingread 2026-10-01 |
|
| Observability & Auditability | Full |
|
Production Logs and Traces are recorded, and the Percival agent reads them step by step to find planning mistakes, incorrect tool use and context misunderstanding in a customer's agent, with trace integrations for smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph, crewAI and custom clients. Logs and traces are kept for the last two weeks on the Developer plan, and Enterprise offers custom data retention. SourcePatronus AI, patronus.ai/percival and patronus.ai/pricingread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Logs, traces, datasets and experiment results are stored for evaluation. No session, workflow or long term memory that a customer's agent reads and writes is offered. SourcePatronus AI, patronus.ai and patronus.ai/pricingread 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
The platform is hosted, and the Enterprise plan adds on premises or dedicated VPC deployment with custom data retention. SourcePatronus AI, patronus.ai/pricingread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Not documented |
|
Evaluator models, datasets and research benchmarks are Patronus's own, and its Percival agent debugs a customer's agents inside the platform. No ready made agents, templates or packaged workflows a buyer adopts for its own work are offered. SourcePatronus AI, patronus.ai/pricing, patronus.ai/percival and patronus.airead 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Teams monitor and receive real time alerts on LLM and agent interactions in production through tracing, logging and alerts, according to the documentation home, and the pricing page lists webhooks on Enterprise. Those checks and alerts fire on production interactions without a person asking. SourcePatronus AI, docs.patronus.ai (search index copy) and patronus.ai/pricingread 2026-09-21 |
|
| Model Flexibility & Routing | Not documented |
|
Customers can call small or large evaluators, which are Patronus's own models such as Lynx and GLIDER. Whether they can choose among foundation models from several providers, or bring their own providers and keys for judges, is not published. SourcePatronus AI, patronus.ai/pricing and patronus.ai; docs.patronus.ai/docsread 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Patronus sells programmatic access to its platform through the Patronus API, metered per thousand evaluator calls and evaluation explanations with an API key from sign-up, and publishes integration docs for smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph, crewAI and custom clients; Enterprise adds webhooks and higher rate limits. SourcePatronus AI, patronus.ai/pricing and patronus.ai/percivalread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
The product line covers Experiments across projects, Comparisons, Datasets and an evaluator API with small and large evaluators and paid evaluation explanations, with Evaluation Runs on Enterprise, and the Percival agent reads a customer's agent traces to detect more than twenty failure modes and suggest fixes. Research evaluators include the Lynx hallucination detection model and the GLIDER judge. SourcePatronus AI, patronus.ai/pricing and patronus.ai/percivalread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
Patronus sells evaluation, tracing and agent debugging for a customer's AI systems, and separately builds simulated digital environments for training models. No browser, desktop or computer control by a customer's agent through Patronus is offered. SourcePatronus AI, patronus.ai and patronus.ai/pricingread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Pricing
Developer free · API from $10 per 1,000 evaluator calls
Pay as you go per evaluator API call, priced by evaluator size, plus a charge per evaluation explanation; platform features free on the Developer plan within limits; Enterprise custom
Included quota
No subscription. $5 in free credits on signup. Published pay as you go rates: about $10 per 1,000 API calls for small evaluators and $20 per 1,000 for large evaluators. Enterprise adds higher rate limits, custom evaluation models, webhooks, and professional services.
What is public
The pay as you go model, $5 free credits, and launch era per call rates are public; current exact rates and enterprise pricing are not fully itemized.
Billing mechanics
Consumption based per evaluation API call, priced by evaluator size (small versus large), with no monthly subscription floor. Enterprise contracts add higher limits, custom models, and services.
Cost watchouts
Every evaluator call and every explanation is billed, so evaluating all production traffic scales cost with volume, and large evaluators cost twice the small rate; the free Developer plan keeps only two weeks of logs and traces.
Variable cost rationale
API spend is pure consumption: cost rises linearly with evaluator calls and explanations, and doubles per call on large evaluators, so screening all production traffic needs cost modeling.
Additional watchouts
Real time guardrailing screens every response, so consumption and cost scale directly with production traffic, and the larger Lynx evaluator costs more and adds latency. Model cost carefully for high volume deployments.
Overage / add-ons
Pure consumption pricing; you pay per evaluation API call with no monthly commitment, and cost scales linearly with call volume and evaluator size. The larger Lynx 70B evaluator costs more per call than the smaller evaluators.
Sales call required
Mixed (some tiers require a call)
Free / trial
Developer plan free with no credit card, and $10 in free API credits
Lowest paid plan
Pay as you go: $10 per 1,000 small evaluator calls, $20 per 1,000 large evaluator calls, $10 per 1,000 explanations
Commercial notes
Self-serve Developer plan with free platform access and pay-as-you-go evaluator API credits, with an Enterprise plan arranged through a call for private deployment, SSO and custom models.
Key ambiguities
Enterprise pricing is not published and is arranged by booking a call. The pricing page also shows a per-page plan slider (Individual free, Base $25 a month) that does not describe Patronus's evaluation products and is not relied on.
Cancellation / refund
Pay as you go has no commitment to cancel. Enterprise terms are contractual.
Support SLA / resale
Self serve and community for pay as you go; higher rate limits, professional services, and enterprise support on enterprise contracts.
Missing data
Current exact per call rates, enterprise pricing, and volume discount tiers are not fully public.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Patronus AI
The closest documented capability profiles to Patronus AI among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Voker6.0 / 14Matches Patronus AI across all 14 documented capabilities
- Hamming AI6.5 / 14Adds documented Human Oversight & Guardrails
- AgentOps5.0 / 14Adds documented Model Flexibility & Routing
- Freeplay8.0 / 14Adds documented Human Oversight & Guardrails and Model Flexibility & Routing
- Monte Carlo8.0 / 14Adds documented Human Oversight & Guardrails and Prebuilt Agents, Templates & Packs
- Traceloop8.0 / 14Adds documented Human Oversight & Guardrails and Model Flexibility & Routing
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded