Galileo
Also known as: Galileo AI, Splunk Agent Observability
AI reliability platform that evaluates, observes, and guardrails GenAI apps and agents using its own low latency Luna evaluation models.
Galileo is an AI observability, evaluation and guardrail platform for agents and LLM applications, built on the idea that offline evals should become production guardrails. It traces every agent step with agent graph visualization, sessions and OpenTelemetry-based distributed tracing, runs more than 20 out-of-the-box evals for RAG, agents, safety and security plus custom evaluators, tunes LLM-as-a-judge metrics from live feedback, and distills them into its own Luna-2 small models so all production traffic can be scored at low latency and cost. Signals detect failure patterns automatically, and Agent Control applies the same checks as inline runtime protection.
Developers use Python and TypeScript SDKs, a REST API and a Galileo MCP server, with integrations for CrewAI, LangGraph, Google ADK, the OpenAI Agents SDK, Strands, Mastra and others. Galileo states SOC 2 Type 2 certification, offers RBAC and SSO, and deploys hosted, in a VPC or on premises. Its site lists a free plan with 5,000 traces a month, Pro at $100 a month billed yearly, and custom Enterprise.
Galileo's documentation now states that as of August 7, 2026 Galileo is Splunk Agent Observability, that the Galileo documentation applies to customers who onboarded before that date, and that newer customers use Splunk's agent observability documentation; galileo.ai itself still presents its own pricing and signup. Cisco has completed its acquisition of Galileo, and Splunk now sells the platform as Splunk Agent Observability while galileo.ai still sells the original Galileo plans. This is the evaluation company at galileo.ai, not the similarly named design tool.
Vendor details
Canonical URL
https://galileo.ai
Category
Agent infrastructure
Subcategory
Evaluation and observability
Funding status
Founded in 2021 by Vikram Chatterji, Atindriyo Sanyal, and Yash Sheth, engineers from Google AI, Google Brain, Apple Siri, and Uber. Headquartered in San Francisco. Has raised about $68M, headlined by a $45M Series B in October 2024 led by Scale Venture Partners, with Premji Invest, Databricks Ventures, ServiceNow Ventures, Citi Ventures, and Battery Ventures participating. Customers include HP, Twilio, Reddit, and Comcast.
Company status
acquired
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
Native integrations with agent frameworks like CrewAI, NVIDIA NeMo and NIM for guardrails, and the MongoDB MAAP ecosystem, plus CI/CD pipelines for unit testing AI before production. Evaluations run on Galileo's own Luna and Luna-2 small language models for low latency scoring.
In practice
Your agent fails intermittently in production and debugging means digging through thousands of traces by hand. Galileo's insights engine clusters similar failures and surfaces the root cause, so you fix the pattern, not single runs.
You want to evaluate every production response, not a sample, but general purpose judge models make that too expensive. Galileo's Luna models score at sub 200 millisecond latency and low cost, making 100 percent traffic evaluation affordable.
Your regulated application cannot risk shipping an unsafe response. Galileo's Agent Control turns your evaluations into runtime guardrails that apply inline protection to agent requests and responses before anything reaches a customer.
Sources & related URLs
Related / legacy domains
Agentic Index coverage score
9.0 / 14 capabilities · 64%
| Integrations & Tool Calling | Partial |
|---|---|
|
Tracing integrations cover A2A, CrewAI, Google ADK, LangChain and LangGraph, Mastra, Microsoft Agent Framework, the OpenAI SDKs, Pydantic AI, Strands and the Vercel AI SDK, and Galileo connects model providers for its metrics. These bring telemetry in; no connectors that let an agent take authenticated actions in outside systems are documented. SourceGalileo, docs.galileo.ai/what-is-galileo and galileo.ai llms.txtread 2026-09-21 |
|
| Workflow Orchestration | Not documented |
|
Agents that run elsewhere are what Galileo observes, evaluates and guards; no workflows that sequence, branch or retry an agent's steps are documented. SourceGalileo, galileo.ai llms.txtread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
Galileo evaluates the grounding of RAG systems built elsewhere with context adherence and RAG metrics, and no retrieval structure over the customer's knowledge that grounds an agent's answers is documented. SourceGalileo, galileo.ai llms.txtread 2026-09-21 |
|
| Human Oversight & Guardrails | Full |
|
Evaluations become production guardrails through Agent Control, which applies inline runtime protection to agent requests and responses, while Luna-2 models run the checks at low latency across all traffic. Together they apply structured guardrails and policy constraints at run time. SourceGalileo, galileo.ai llms.txt and homepage, and docs.galileo.ai/what-is-galileoread 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
A SOC 2 Type 2 certification is stated, and Galileo's plans and documentation cover project access control, with standard RBAC on Pro and RBAC with SSO integration on Enterprise. SourceGalileo, galileo.ai llms.txt and pricing, and docs.galileo.ai/what-is-galileo (Security section)read 2026-09-21 |
|
| Observability & Auditability | Full |
|
Every agent step is captured in end to end traces, with agent graph visualization, sessions and distributed tracing over OpenTelemetry. Real time monitoring raises alerts, Signals detect failure patterns automatically, and the SDKs carry a logger, decorators and context helpers. Together that shows step by step what the agent did. SourceGalileo, galileo.ai llms.txt and docs.galileo.ai/what-is-galileoread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Traces, sessions and datasets are kept for observability and evaluation, but no session, conversation or long-term memory that an agent reads back as context is documented. SourceGalileo, docs.galileo.ai/what-is-galileoread 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
On the Enterprise plan, Galileo deploys hosted, in the customer's VPC or on premises, and its own summary states SaaS, on-premises and in-VPC deployment. SourceGalileo, galileo.ai pricing and llms.txtread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
One AI Assistant inside the app, in beta, ships with sample projects to start from, alongside Galileo's evaluator library. There is no set of prebuilt agents or templates for a buyer to select from. SourceGalileo, docs.galileo.ai/what-is-galileoread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Luna-2 models evaluate all production traffic as it arrives, log stream metrics score incoming traces, Signals detect failures proactively, and real-time monitoring raises alerts, so Galileo's evaluation work starts from incoming telemetry with no person initiating each run. SourceGalileo, galileo.ai llms.txt and homepage, and docs.galileo.ai/what-is-galileoread 2026-09-21 |
|
| Model Flexibility & Routing | Full |
|
Model integrations in the documentation connect the customer's LLM providers, with costs per model, and Galileo's LLM-as-a-judge metrics and playground experiments run on them beside its own Luna-2 models. The customer chooses the models behind Galileo's own evaluation features. SourceGalileo, docs.galileo.ai/what-is-galileo and galileo.ai llms.txtread 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Galileo documents a Python SDK, a TypeScript SDK, a REST API reference with authentication and endpoints, and a Galileo MCP server for AI-enabled IDEs such as Cursor and VS Code. SourceGalileo, docs.galileo.ai/what-is-galileo and galileo.ai llms.txtread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
Galileo ships more than 20 out-of-the-box evals for RAG, agents, safety and security plus custom evaluators, LLM-as-a-judge metrics tuned from live feedback with Autotune and distilled into Luna-2 models that score all production traffic, and experiments run in code, playgrounds or unit tests against datasets with side-by-side comparison and annotations. That tests the customer's agent before production and scores quality over time. SourceGalileo, galileo.ai homepage and docs.galileo.ai/what-is-galileoread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
Galileo observes, evaluates and guards agents, and no browser, desktop or computer control by an agent is documented. SourceGalileo, galileo.ai llms.txtread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Galileo has officially transitioned to become Splunk Agent Observability. The existing platform and documentation now only apply to legacy customers who onboarded prior to this date, while new users are directed to Splunk's dedicated infrastructure.
Bears on: Funding / partnership
View sourceGalileo released Annotation Queues to general availability, alongside updates to OpenAI model integrations and enhancements to the AI Assistant beta. The release also introduces flexible custom model integrations for connecting proprietary or third-party models into observability workflows.
Bears on: Human approval / guardrails
View sourcePricing
From $100/mo billed yearly · free tier
Monthly traces within subscription tiers; Pro pricing scales with trace volume
Included quota
Free 5,000 traces/mo (unlimited users, unlimited custom evals). Pro 50,000 traces/mo. Enterprise unlimited traces.
What is public
Galileo publishes Free and Pro pricing with trace limits and feature differences. Enterprise pricing and exact trace tier steps above Pro are custom.
Billing mechanics
A usage based subscription metered on monthly traces. Free includes 5,000 traces; Pro starts at $100 a month billed yearly for 50,000 traces and scales with volume; Enterprise is a custom contract with unlimited traces. Runtime guardrails and dedicated inference are Enterprise only.
Cost watchouts
Trace volume is the main cost driver and grows with production traffic. Capabilities many teams consider essential, such as runtime guardrails and private deployment, require the Enterprise tier.
Variable cost rationale
Tiers are predictable monthly subscriptions, but Pro pricing scales with trace volume, so a busy production system that logs 100 percent of traffic can move up the trace tiers quickly. Galileo's low cost Luna evaluators are designed to keep that scaling affordable.
Additional watchouts
Pro pricing scales with trace volume, so 100 percent traffic evaluation on a high volume app can climb tiers. Runtime guardrails (Galileo Protect), SSO, and VPC or on premises deployment are gated to Enterprise.
Overage / add-ons
Pro pricing scales with trace volume above the base allotment. Enterprise offers unlimited traces under a custom contract.
Sales call required
No, self serve available
Free / trial
Free tier: 5,000 traces/month, unlimited users, unlimited custom evals, no card
Lowest paid plan
Pro $100/mo billed yearly (50,000 traces, RBAC, advanced analytics)
Commercial notes
Sold both bottom up through a generous free and Pro self serve tier and top down to large enterprises that need VPC or on premises deployment, runtime guardrails, and SSO. Customers skew to large companies including HP, Twilio, Reddit, and Comcast.
Key ambiguities
galileo.ai still lists Free, Pro and Enterprise, but Galileo's documentation states that from 7 August 2026 Galileo is Splunk Agent Observability and newer customers onboard through Splunk, so the purchase path for a new buyer may now run through Splunk.
Cancellation / refund
Free and Pro are self serve subscriptions, cheaper billed yearly; standard cancellation. Enterprise terms are contractual.
Support SLA / resale
Community support on Free, dedicated Slack support on Pro, and a dedicated customer success manager and SLA on Enterprise.
Missing data
Enterprise pricing, the exact trace tier steps and per trace overage above Pro, and the precise cost of dedicated inference servers are not public.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Galileo
The closest documented capability profiles to Galileo among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- F5 AI Guardrails9.0 / 14Matches Galileo across all 14 documented capabilities
- Opik9.0 / 14Matches Galileo across all 14 documented capabilities
- W&B Weave9.0 / 14Matches Galileo across all 14 documented capabilitiesGalileo vs W&B Weave →
- Confident AI8.5 / 14A lighter documented profile than GalileoGalileo vs Confident AI →
- Fiddler AI8.5 / 14A lighter documented profile than Galileo
- HoneyHive8.5 / 14A lighter documented profile than Galileo
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded