RunWhen
Platform for SRE agent orchestration and automated troubleshooting workflows.
RunWhen is an AI SRE platform that automates troubleshooting and remediation for Kubernetes and cloud environments. In autonomous mode it continuously runs diagnostic tasks in the background to build always-current production insights, and in assistant mode an AI agent selects relevant tasks from a large library of expert-authored automation and runs them in your environment via a lightweight in-cluster runner, returning root-cause findings and next steps while credentials and cluster data never leave your network. Its open-source CodeCollections provide a public registry of 7 collections, 215 Skill Templates and 918 tools across SRE, DevOps and FinOps, tailored per discovered resource. RunWhen ships a built-in read-only Model Context Protocol server so agents like Claude Code and Cursor can use its Skills, which also run as standalone CLI tools with no long-term commitment.
Vendor details
Canonical URL
https://www.runwhen.com
Category
SRE / DevOps agent
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
Kubernetes, AWS, GCP and Azure with a native Azure indexer; 918 CLI and API tools plus a read-only MCP server.
Sources & related URLs
Research sources
Capability coverage
10.5 / 14 capabilities · 75%
| Integrations & Tool CallingConnects to Kubernetes clusters and AWS, GCP and Azure accounts with a native Azure indexer, exposing 918 CLI and API tools across the stack, RunWhen site 2026-07-22 | Full |
|---|---|
| Workflow OrchestrationAn AI assistant selects and runs tasks from a library of expert-authored automation, and an autonomous mode continuously runs diagnostics, analyzing results into root-cause findings and next steps, RunWhen site 2026-07-22 | Full |
| Knowledge Grounding & RAGContinuously builds always-current production insights, structured findings about service health, configuration and behavior, ready for the AI to reference, RunWhen site 2026-07-22 | Full |
| Human Oversight & GuardrailsReturns root-cause findings and recommended next steps for engineers and runs safe-for-production skills, keeping humans in control, RunWhen site 2026-07-22 | Full |
| Security, Identity & GovernanceCredentials and cluster data never leave your network and the read-only runner stays in-cluster, with RBAC and IAM audit skills, though formal certifications are not documented, RunWhen site 2026-07-22 | Partial |
| Observability & AuditabilityContinuously builds structured production insights about service health, configuration and behavior, though it runs diagnostic tasks over your systems rather than providing its own telemetry, RunWhen site 2026-07-22 | Partial |
| Memory & State PersistenceContinuously builds and maintains always-current production insights the AI references, a persistent knowledge store though not cross-incident learning memory, RunWhen site 2026-07-22 | Partial |
| Deployment & Data ResidencyA lightweight runner stays inside your cluster and credentials and cluster data never leave your network, with a self-hostable RunWhen Local discovery mode, RunWhen site 2026-07-22 | Full |
| Prebuilt Agents, Templates & PacksShips a large library of prebuilt skills, 7 CodeCollections, 215 Skill Templates and 918 tools, open source and tailored per discovered resource, RunWhen site 2026-07-22 | Full |
| Triggers & Channel CoverageRuns diagnostics continuously in the background in autonomous mode and supports on-call alert triage and investigation on demand, RunWhen site 2026-07-22 | Full |
| Model Flexibility & RoutingSkills are model and agent agnostic, usable with Claude, Cursor or LangGraph or as standalone CLI tools rather than locked to one model, RunWhen site 2026-07-22 | Partial |
| APIs, SDKs & MCP ExtensibilityShips a built-in read-only Model Context Protocol server plus open-source CodeCollections and standalone CLI tools, a strong extensibility surface, RunWhen site 2026-07-22 | Full |
| Testing, Debugging & OptimizationSkills are safe-for-production and tailored per resource and can be run standalone to validate, though not a dedicated agent evaluation harness, RunWhen site 2026-07-22 | Partial |
| Browser & Computer UseOperates through CLI, tasks and an MCP server rather than browser or computer interface control, RunWhen site 2026-07-22 | Unable to verify |
Pricing
Open-source Skills + free RunWhen Local; platform pricing not public
What is public
Open-source CodeCollections and RunWhen Local are free; the managed platform's pricing is not public, with no long-term commitment required.
Variable cost rationale
Open-source Skills and RunWhen Local are free to self-host; managed platform pricing is not public.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free open-source CodeCollections + RunWhen Local; no long-term commitment
Related vendors
- Alertd — Agentic AI SRE and DevOps platform for AWS that analyzes cost,…
- Anyshift — AI SRE agent that investigates production incidents by tracing…
- Avesha — Autonomous AI SRE platform whose coordinated agent swarm monitors,…
- Beeps — Agent native on call platform that routes incidents to humans and AI…
- Better Stack — AI native incident management with built in on call and status…
- Bluebricks — Context and control layer that lets AI agents operate cloud…