Back to vendors
P

Patronus AI

Also known as: Patronus, Patronus API

Visit site
Entry priceDeveloper free · API from $10 per 1,000 evaluator callsFull pricing detail

Evaluation, tracing and agent debugging platform with its own evaluator models, Experiments, production logs and traces, and the Percival agent debugger, sold through a free Developer plan and a per-call evaluator API.

Patronus AI sells a platform for evaluating, tracing and debugging AI applications and agents. Teams run Experiments that score prompts, models and data configurations against Datasets, compare results over time, and record production Logs and Traces.

Its evaluator API scores outputs with Patronus's own small and large evaluator models, including the Lynx hallucination detection model and the GLIDER judge, returns explanations on request, and is priced per call.

Percival, an agent built by Patronus, reads a customer's agent traces, detects more than twenty failure modes such as planning mistakes and incorrect tool use, and suggests fixes, with integrations for smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph and crewAI.

The Developer plan is free with no credit card and includes $10 in API credits; Enterprise adds on-prem or dedicated VPC deployment, SSO, custom data retention, webhooks, higher rate limits and custom evaluator fine-tuning. The company's homepage now leads with its research lab, which builds Digital World Models and reinforcement learning environments for training frontier models, and its platform documentation now requires a login.

Vendor details

Canonical URL

https://www.patronus.ai

Category

Agent infrastructure

Subcategory

Evaluation and guardrails

Funding status

Founded in 2023 in San Francisco by former Meta AI (FAIR) researchers Anand Kannappan (CEO) and Rebecca Qian (CTO). Raised a $17M Series A and a $50M round in 2026 to expand its agent simulation platform. Customers include AngelList, Pearson, HP, and Fortune 500 companies in finance, healthcare, and legal, with partners including NVIDIA, MongoDB, and IBM. Independent.

Company status

independent

Use cases & customers

Primary use cases

hallucination detectionAI guardrailsagent evaluationagent debuggingagent simulation

Target customers

AI engineering teamsenterprise

Deployment options

SaaSself-hosted

Integrations

Embedded into application code through a programming language agnostic API and Python SDK, with custom LLM judges, webhooks, and a web dashboard for logs and experiments. Evaluates RAG and agent pipelines, and ships open research models including the Lynx hallucination detector and the GLIDER judge.

In practice

Your RAG chatbot sometimes states facts not in the source documents. You wire Patronus Lynx in as a real time guardrail, and it flags the hallucinated span before the answer reaches the user.

Your agent fails intermittently across multi step traces and you cannot tell why. Percival inspects the execution traces, identifies which of twenty plus failure modes occurred, and suggests prompt and workflow fixes.

You are choosing between two prompts and three models for a regulated use case. Patronus Experiments runs them side by side against your criteria so you can pick the configuration that scores best before shipping.

Agentic Index coverage score

6.5 / 14 capabilities · 46%

Integrations & Tool Calling Partial

Trace capture works with smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph, crewAI and custom clients, feeding a customer's agent traces to Percival, and Enterprise adds webhooks. These bring the customer's agent data in and send results out; the customer's agent reaches its own tools, and no integration through which an agent reads or writes a real system is described.

SourcePatronus AI, patronus.ai/percival and patronus.ai/pricingread 2026-09-21

Workflow Orchestration Not documented

Customers build and run their agents in frameworks such as LangGraph, crewAI and the OpenAI Agents SDK, and Patronus evaluates and debugs them there. It offers no sequencing, branching, retries or routing of agent steps of its own.

SourcePatronus AI, patronus.ai/percival and patronus.ai/pricingread 2026-09-21

Knowledge Grounding & RAG Not documented

The Lynx model checks retrieval augmented answers for hallucination, and Patronus scores a customer's agents. No document ingestion, index, retrieval layer or knowledge API that grounds an agent in company data is offered.

SourcePatronus AI, patronus.ai and patronus.ai/pricingread 2026-09-21

Human Oversight & Guardrails Not documented

People confirm issues and annotate missed failure modes in agent traces on the Percival page, and the homepage presents the GLIDER evaluator as suited to guardrails. The public pages do not say whether Patronus provides a guardrail, approval step or checkpoint that holds or blocks an agent's output, rather than a verdict the customer's code acts on.

SourcePatronus AI, patronus.ai/percival and patronus.ai; docs.patronus.ai/docsread 2026-09-21

Security, Identity & Governance Full

SSO comes on the Enterprise plan, with identity integration through the customer's own provider, alongside custom data retention and on premises or dedicated VPC deployment. The pricing page carries AICPA SOC, TISAX and HIPAA badges as an asserted attestation; no report, type or auditor is named and no trust page is linked.

SourcePatronus AI, patronus.ai/pricingread 2026-10-01

Observability & Auditability Full

Production Logs and Traces are recorded, and the Percival agent reads them step by step to find planning mistakes, incorrect tool use and context misunderstanding in a customer's agent, with trace integrations for smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph, crewAI and custom clients. Logs and traces are kept for the last two weeks on the Developer plan, and Enterprise offers custom data retention.

SourcePatronus AI, patronus.ai/percival and patronus.ai/pricingread 2026-09-21

Memory & State Persistence Not documented

Logs, traces, datasets and experiment results are stored for evaluation. No session, workflow or long term memory that a customer's agent reads and writes is offered.

SourcePatronus AI, patronus.ai and patronus.ai/pricingread 2026-09-21

Deployment & Data Residency Full

The platform is hosted, and the Enterprise plan adds on premises or dedicated VPC deployment with custom data retention.

SourcePatronus AI, patronus.ai/pricingread 2026-09-21

Prebuilt Agents, Templates & Packs Not documented

Evaluator models, datasets and research benchmarks are Patronus's own, and its Percival agent debugs a customer's agents inside the platform. No ready made agents, templates or packaged workflows a buyer adopts for its own work are offered.

SourcePatronus AI, patronus.ai/pricing, patronus.ai/percival and patronus.airead 2026-09-21

Triggers & Channel Coverage Full

Teams monitor and receive real time alerts on LLM and agent interactions in production through tracing, logging and alerts, according to the documentation home, and the pricing page lists webhooks on Enterprise. Those checks and alerts fire on production interactions without a person asking.

SourcePatronus AI, docs.patronus.ai (search index copy) and patronus.ai/pricingread 2026-09-21

Model Flexibility & Routing Not documented

Customers can call small or large evaluators, which are Patronus's own models such as Lynx and GLIDER. Whether they can choose among foundation models from several providers, or bring their own providers and keys for judges, is not published.

SourcePatronus AI, patronus.ai/pricing and patronus.ai; docs.patronus.ai/docsread 2026-09-21

APIs, SDKs & MCP Extensibility Full

Patronus sells programmatic access to its platform through the Patronus API, metered per thousand evaluator calls and evaluation explanations with an API key from sign-up, and publishes integration docs for smolagents, Pydantic AI, the OpenAI Agents SDK, LangGraph, crewAI and custom clients; Enterprise adds webhooks and higher rate limits.

SourcePatronus AI, patronus.ai/pricing and patronus.ai/percivalread 2026-09-21

Testing, Debugging & Optimization Full

The product line covers Experiments across projects, Comparisons, Datasets and an evaluator API with small and large evaluators and paid evaluation explanations, with Evaluation Runs on Enterprise, and the Percival agent reads a customer's agent traces to detect more than twenty failure modes and suggest fixes. Research evaluators include the Lynx hallucination detection model and the GLIDER judge.

SourcePatronus AI, patronus.ai/pricing and patronus.ai/percivalread 2026-09-21

Browser & Computer Use Not documented

Patronus sells evaluation, tracing and agent debugging for a customer's AI systems, and separately builds simulated digital environments for training models. No browser, desktop or computer control by a customer's agent through Patronus is offered.

SourcePatronus AI, patronus.ai and patronus.ai/pricingread 2026-09-21

The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Pricing

Developer free · API from $10 per 1,000 evaluator calls

Pay as you go per evaluator API call, priced by evaluator size, plus a charge per evaluation explanation; platform features free on the Developer plan within limits; Enterprise custom

Free tier

Included quota

No subscription. $5 in free credits on signup. Published pay as you go rates: about $10 per 1,000 API calls for small evaluators and $20 per 1,000 for large evaluators. Enterprise adds higher rate limits, custom evaluation models, webhooks, and professional services.

What is public

The pay as you go model, $5 free credits, and launch era per call rates are public; current exact rates and enterprise pricing are not fully itemized.

Billing mechanics

Consumption based per evaluation API call, priced by evaluator size (small versus large), with no monthly subscription floor. Enterprise contracts add higher limits, custom models, and services.

Cost watchouts

Every evaluator call and every explanation is billed, so evaluating all production traffic scales cost with volume, and large evaluators cost twice the small rate; the free Developer plan keeps only two weeks of logs and traces.

Variable cost rationale

API spend is pure consumption: cost rises linearly with evaluator calls and explanations, and doubles per call on large evaluators, so screening all production traffic needs cost modeling.

Additional watchouts

Real time guardrailing screens every response, so consumption and cost scale directly with production traffic, and the larger Lynx evaluator costs more and adds latency. Model cost carefully for high volume deployments.

Overage / add-ons

Pure consumption pricing; you pay per evaluation API call with no monthly commitment, and cost scales linearly with call volume and evaluator size. The larger Lynx 70B evaluator costs more per call than the smaller evaluators.

Sales call required

Mixed (some tiers require a call)

Free / trial

Developer plan free with no credit card, and $10 in free API credits

Lowest paid plan

Pay as you go: $10 per 1,000 small evaluator calls, $20 per 1,000 large evaluator calls, $10 per 1,000 explanations

Commercial notes

Self-serve Developer plan with free platform access and pay-as-you-go evaluator API credits, with an Enterprise plan arranged through a call for private deployment, SSO and custom models.

Key ambiguities

Enterprise pricing is not published and is arranged by booking a call. The pricing page also shows a per-page plan slider (Individual free, Base $25 a month) that does not describe Patronus's evaluation products and is not relied on.

Cancellation / refund

Pay as you go has no commitment to cancel. Enterprise terms are contractual.

Support SLA / resale

Self serve and community for pay as you go; higher rate limits, professional services, and enterprise support on enterprise contracts.

Missing data

Current exact per call rates, enterprise pricing, and volume discount tiers are not fully public.

Agentic Index verified 2026-09-21

Alternatives to Patronus AI

The closest documented capability profiles to Patronus AI among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • Voker6.0 / 14Matches Patronus AI across all 14 documented capabilities
  • Hamming AI6.5 / 14Adds documented Human Oversight & Guardrails
  • AgentOps5.0 / 14Adds documented Model Flexibility & Routing
  • Freeplay8.0 / 14Adds documented Human Oversight & Guardrails and Model Flexibility & Routing
  • Monte Carlo8.0 / 14Adds documented Human Oversight & Guardrails and Prebuilt Agents, Templates & Packs
  • Traceloop8.0 / 14Adds documented Human Oversight & Guardrails and Model Flexibility & Routing

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.