Causely
Causal reasoning platform that infers root cause and blast radius from a causal model of the environment, for SRE teams and their AI agents through an MCP server and a GraphQL API.
Causely builds a causal model of the customer's environment and uses causal reasoning, rather than an LLM running queries step by step, to infer the root cause and blast radius when symptoms start cascading. Correlation based tools cluster alerts that happen together and rank suspects; Causely models the causal relationships between services, infrastructure and data systems, so a storm of downstream symptoms collapses to the cause that produced them.
A mediator and agents run in the customer's environment, discover the topology from telemetry sources such as Datadog, Dynatrace, Prometheus, OpenTelemetry and the cloud provider APIs, and keep raw telemetry local, sending distilled symptom state and topology to the causal engine, which runs in Causely's cloud or the customer's own. Diagnoses reach responders through Slack, Teams, incident.io, Splunk On-Call and webhooks, and reach AI agents through an MCP server with 45 tools and six prebuilt skills, alongside a GraphQL API. For resource contention diagnoses, a person can apply the scaling fix with one click once the in-cluster executor is enabled, and the action is recorded.
Professional is published at $2,000 per month for up to 500 services, with a free 30 day trial; Enterprise is custom and adds on premises and air gapped deployment. Automated remediation is narrow, limited to resource contention, so teams wanting broader automated fixes or on call management will pair it with other tools.
Vendor details
Canonical URL
https://www.causely.ai
Category
SRE / DevOps agent
Subcategory
AI SRE agent
Funding status
Independent, venture backed. Differentiates on causal reasoning rather than correlation, auto discovering the environment and delivering insights in seconds from existing telemetry with no setup or tuning required.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
Connects to 20+ telemetry sources, including Datadog, Dynatrace, Prometheus, OpenTelemetry, Splunk, Elasticsearch and the AWS, Azure and GCP APIs, through a mediator in the customer's environment. Sends diagnoses to Slack, Microsoft Teams, incident.io, Splunk On-Call and generic webhooks, and exposes them to agents such as Claude Code, Cursor, Codex and GitHub Copilot through an MCP server, plus a GraphQL API.
In practice
One failure triggers a hundred downstream alerts and your team chases symptoms. Causely models causal relationships and collapses the storm to the underlying root cause and its blast radius.
Your AI coding or ops agents guess during incidents. Causely's MCP server gives them root cause, blast radius and remediation guidance from a causal model rather than raw telemetry.
Your correlation based AIOps ranks suspects but never tells you which one actually broke. Causely reasons over causality to name the cause, not a list.
Sources & related URLs
Agentic Index coverage score
9.5 / 14 capabilities · 68%
| Integrations & Tool Calling | Full |
|---|---|
|
Named catalog of 20+ telemetry sources (Datadog, Dynatrace, Prometheus, Splunk, Elasticsearch, AWS, Azure, GCP, Kubernetes, Snowflake) with outbound workflows to Slack, Microsoft Teams, incident.io, Splunk On-Call and a generic webhook; the executor on the in-cluster mediator applies scaling actions in the customer's cluster. Sourcedocs.causely.airead 2026-09-29 |
|
| Workflow Orchestration | Partial |
|
A causal reasoning engine produces root cause and blast radius in one analysis; no multi step workflow builder or multi agent runtime is documented. The six MCP skills sequence tool calls inside the customer's own agent, which is that agent's orchestration rather than Causely's. Sourcedocs.causely.ai/agent-integration/mcp-serverread 2026-09-29 |
|
| Knowledge Grounding & RAG | Full |
|
Causely maintains a causal model of the customer's environment. The mediator discovers services, infrastructure and dependencies, and agents periodically forward topology and symptom data to the backend, which keeps the causal graph current and queryable. Sourcedocs.causely.ai/getting-started/architectureread 2026-09-29 |
|
| Human Oversight & Guardrails | Full |
|
Remediation commits only when a person triggers it from the UI ('you can trigger automated remediation or apply a guided fix with one click'), and automated actions require the executor to be enabled on the mediator in that cluster. Nothing published describes remediation running without a person. Sourcedocs.causely.ai/in-action/automate-remediationread 2026-09-29 |
|
| Security, Identity & Governance | Full |
|
Access surface: workspace SSO, user management with assigned roles, personal and tenant wide API tokens, and MCP configuration tools limited to Developer and Admin roles; an admin audit log supports it. Compliance: SOC 2, report on request, with encryption in transit and at rest. Sourcecausely.ai/securityread 2026-09-29 |
|
| Observability & Auditability | Full |
|
Each diagnosis carries its causal chain and the targeted telemetry sent as evidence, and remediation keeps an 'auditable action history tied to the entity' it changed, a per action record readable after the fact. Sourcedocs.causely.ai/in-action/automate-remediationread 2026-09-29 |
|
| Memory & State Persistence | Not documented |
|
What persists is application data rather than agent memory. Diagnosis history is kept under 30 day retention, and thresholds, service tiers, custom causes and snoozed issues are settings the product applies. The causal model is the product's picture of the environment, and no separate memory layer is documented. Sourcedocs.causely.ai/getting-started/architectureread 2026-09-29 |
|
| Deployment & Data Residency | Full |
|
The mediator and agents always run in the customer environment, the causal engine runs in the customer's cloud (BYOC) or in Causely managed infrastructure, and Enterprise offers on premises and air gapped deployment; raw telemetry stays local. No region list is published. Sourcedocs.causely.ai/getting-started/architectureread 2026-09-29 |
|
| Prebuilt Agents, Templates & Packs | Full |
|
Six named skills ship with the MCP server, each with a distinct job. Three handle live incidents (causely-alert-triage, causely-correlated-incidents and causely-k8s-investigation), and the other three cover change impact (causely-change-impact), health reporting (causely-health-reporting) and postmortems (causely-postmortem). Sourcedocs.causely.ai/agent-integration/mcp-serverread 2026-09-29 |
|
| Triggers & Channel Coverage | Full |
|
Analysis runs continuously on streaming telemetry and on alerts routed in from existing tools ('An alert fires, Causely answers why'), and diagnoses are pushed to Slack, Teams, incident.io, Splunk On-Call or a webhook with nobody typing. Sourcedocs.causely.airead 2026-09-29 |
|
| Model Flexibility & Routing | Not documented |
|
Diagnosis comes from a deterministic causal reasoning engine and no LLM provider or customer model choice is documented. Customers can bring their own agent through the MCP server, which is an extension point rather than a choice of model. Sourcecausely.airead 2026-09-29 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Documented GraphQL API at api.causely.app/query with bearer token authentication, personal and tenant API tokens, and query guides (defects and root causes, hiding root causes, SLOs on paths); an MCP server at api.causely.app/mcp adds 45 tools, including configuration of thresholds, service tiers and snoozing. Sourcedocs.causely.ai/api/graphql-clientsread 2026-09-29 |
|
| Testing, Debugging & Optimization | Not documented |
|
No evaluation harness, scored tests or quality gate for the diagnosis engine is documented. Post deploy validation compares the customer's services before and after a release, which tests the customer's change rather than the agent, and the 63 percent benchmark is the vendor's own marketing. Sourcedocs.causely.ai/agent-integration/mcp-serverread 2026-09-29 |
|
| Browser & Computer Use | Not documented |
|
Works through telemetry sources, the in-cluster executor, a GraphQL API and an MCP server; no browser, desktop or computer control is documented. Sourcedocs.causely.airead 2026-09-29 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Pricing
Professional $2,000 per month (up to 500 services); Enterprise custom
Monthly by service count on Professional (up to 500 services); Enterprise on a custom plan
Included quota
Professional: up to 500 services, all published integrations, 30 days data retention, standard hybrid deployment.
What is public
Professional at $2,000 per month for up to 500 services, with all published integrations, a shared Slack or Teams support channel, 30 days of data retention and standard hybrid deployment. Enterprise is custom. A free 30 day trial.
Billing mechanics
Monthly subscription sized by the number of services on Professional; Enterprise is a custom plan with no service limit, custom retention, custom integrations built by the vendor and on premises or air gapped deployment.
Cost watchouts
Services above 500, longer data retention, custom integrations and on premises or air gapped deployment all sit in the unpriced Enterprise plan.
Variable cost rationale
A flat monthly subscription sized by service count; no usage metering is published.
Additional watchouts
The published tier caps at 500 services and 30 days of retention, so larger estates and longer history move to a custom Enterprise plan whose price is not published.
Overage / add-ons
Not published; services above the Professional allowance move to Enterprise
Sales call required
Mixed (some tiers require a call)
Free / trial
Free 30 day trial
Lowest paid plan
Professional, $2,000 per month (up to 500 services, 30 days data retention, standard hybrid deployment)
Commercial notes
Independent, venture backed. Differentiates on causal reasoning rather than correlation, with auto discovery and near instant time to insight.
Key ambiguities
The Professional price is published; the Enterprise price and any overage for services beyond 500 are not.
Related vendors
- AlertD — AI agents for AWS operations that run read only in the customer's…
- Anyshift — AI SRE built on a versioned knowledge graph of infrastructure, apps…
- Avesha — Obliq, Avesha's autonomous AI SRE for Kubernetes and agentic…
- Better Stack — Better Stack offers incident management with built in on call,…
- Bluebricks — Context and control layer that lets AI agents operate cloud…
- Cased — AI workflows for infrastructure and platform engineers: default and…
Alternatives to Causely
The closest documented capability profiles to Causely among SRE and DevOps agents tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Cased8.0 / 14A lighter documented profile than Causely
- Middleware8.0 / 14A lighter documented profile than Causely
- PagerDuty11.0 / 14Adds documented Memory & State Persistence and Testing, Debugging & Optimization
- Better Stack9.5 / 14Adds documented Memory & State Persistence
- Bluebricks10.5 / 14Adds documented Memory & State Persistence and Testing, Debugging & Optimization
- Chamber9.5 / 14Adds documented Memory & State Persistence and Testing, Debugging & Optimization
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded