Bernstein
Also known as: bernstein.run, Bernstein orchestrator, bernstein-orchestrator
Open source deterministic orchestrator that runs 40 plus CLI coding agents such as Claude Code, Codex, and Gemini CLI in parallel git worktrees with no model in the coordination loop, so runs replay byte identically behind an HMAC signed audit chain.
Bernstein is an open-source governance layer for AI coding agents. It takes a software goal or a plan file, decomposes it into a task graph, and hands coordination to deterministic Python: scheduling, worktree isolation, verification gates and merging are auditable code rather than model responses, so the same plan replays byte-identically. Each task runs a CLI coding agent in its own git worktree, with adapters for Claude Code, Codex, Gemini CLI, Cursor, Aider, GitHub Copilot CLI, the OpenAI Agents SDK and more than 40 others, so local and cloud models can mix in one run and an optional router picks a cheaper model when the task allows.
A janitor stage checks tests, types and lint, with optional cross-model review and tournament selection, before any diff lands, and destructive actions pause mid-run for approval in the TUI, the web dashboard, the CLI or chat platforms such as Slack, Telegram, Discord and Teams. An always-on signed lineage spine, a replay journal, content-addressed evidence bundles and an opt-in HMAC-chained audit log let a reviewer verify what happened offline without rerunning anything. Runs start from the CLI, plan files, tickets, chat commands, MCP clients, dashboard webhooks and an autofix daemon that fixes failing CI on its own pull requests.
Bernstein runs wherever Python runs: on a laptop, on premises, air-gapped, as a Kubernetes cluster or on Cloudflare Workers, with per-agent credential scoping and pluggable sandboxes such as Docker, E2B, Modal and Daytona, and it exposes a REST task API with an OpenAPI spec, an MCP server and an A2A agent card. It is Apache 2.0, free, built by a single maintainer, Alex Chernysh, and described by its own site as beta software.
Vendor details
Canonical URL
https://bernstein.run
Category
Agent infrastructure
Subcategory
Deterministic multi-agent orchestration
Funding status
No funding disclosed; solo maintained open source project
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
More than 40 CLI agent adapters including Claude Code, Codex CLI, Gemini CLI, Cursor, Aider, GitHub Copilot CLI, OpenAI Agents SDK, Goose, OpenHands, and AWS Q Developer, plus a generic CLI wrapper; pluggable sandboxes (Docker, E2B, Modal); artifact storage on S3, GCS, and R2; chat bridges for Slack, Discord, and Telegram; MCP server mode over stdio, HTTP, and SSE; a task server REST API.
In practice
A platform team points Bernstein at a refactor goal; it splits the work across Claude Code, Codex, and Aider in parallel worktrees, gates every merge on tests and lint, and leaves a clean git history.
A consultancy stands up an AI engineering crew on a client repo in minutes, keeps its own keys scoped per agent, and hands the client an HMAC signed audit record replayable for compliance review.
A reviewer checks what a multi agent run actually did by replaying its journal and verifying the signed lineage offline, without rerunning any agent or trusting a model narrative.
Sources & related URLs
Related / legacy domains
Agentic Index coverage score
10.5 / 14 capabilities · 75%
| Integrations & Tool Calling | Full |
|---|---|
|
More than 40 CLI agent adapters (Claude Code, Codex, Gemini CLI, Cursor, Aider, GitHub Copilot CLI, OpenAI Agents SDK, AWS Q Developer and more) plus a generic wrapper for any CLI tool, pluggable sandboxes (Docker, E2B, Modal), artifact storage on S3, GCS, and R2, and chat bridges for Slack, Discord, and Telegram. Bernstein docs and PyPI Aug 2026 Sourcebernstein.readthedocs.io/en/latest/adapters/adapter_guideread 2026-09-21 |
|
| Workflow Orchestration | Full |
|
Orchestration is the product: a goal decomposes into a task graph, deterministic Python schedules agents in parallel git worktrees, a documented lifecycle FSM governs task and agent states, multi stage plans run from plan.yaml, and verified results merge automatically. Bernstein docs Aug 2026 Sourcebernstein.readthedocs.io/en/latestread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
No knowledge base, retrieval, or grounding layer is documented; agents ground in the repository itself and the.sdd workspace holds run state rather than knowledge. Bernstein docs Aug 2026 Sourcebernstein.readthedocs.io/en/latestread 2026-09-21 |
|
| Human Oversight & Guardrails | Full |
|
When an agent wants to take a destructive action mid-run, Bernstein pauses it for approval in the TUI, a web dashboard button, the CLI from a second terminal, or approve and reject buttons in Telegram, Slack, Discord or Teams, and always_allow rules decide which calls skip the prompt; every merge also waits on the janitor gate. That is a shipped approval control where the buyer decides which actions need sign-off and routes the request into existing channels. SourceBernstein, bernstein.run/blog/operator-commands and homepageread 2026-09-21 |
|
| Security, Identity & Governance | Partial |
|
Per-agent credential scoping gives each agent only the environment it declared needing, with SPIFFE workload identity and sandbox backends isolating execution, a least-privilege access model for agents. Bernstein is open-source software from a single maintainer with no hosted service, and its homepage offers its hash-chained audit log as proof in place of a SOC 2 report, so no attestation exists. SourceBernstein, bernstein.run homepage and llms.txtread 2026-09-21 |
|
| Observability & Auditability | Full |
|
Auditability is the thesis: an always on signed lineage spine, a deterministic replay journal, content addressed evidence bundles, an opt in HMAC chained audit log a reviewer can verify offline without rerunning, a live TUI dashboard, and an observability extra with OpenTelemetry support. Bernstein site and docs Aug 2026 Sourcebernstein.runread 2026-09-21 |
|
| Memory & State Persistence | Partial |
|
State persistence is strong: a durable ledger resumes runs on another machine, a replay journal and content addressed evidence bundles preserve every step, and workspace state lives in.sdd with no server to provision. Learned memory across runs in the agent sense is not documented. Bernstein site release notes Aug 2026 Sourcebernstein.runread 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
Runs entirely where the user chooses: installs via pip, pipx, uv, brew, and npm, state local in.sdd with no server to provision, cluster mode, a Cloudflare Workers backend, Kubernetes and Docker extras, and a documented air gap installation profile with a wheelhouse build and deny all egress. Bernstein docs Aug 2026 Sourcebernstein.readthedocs.io/en/latest/installation/air-gapread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
Progressive skills and reusable plan.yaml multi stage plans ship with the tool, and the adapter guide documents adding agents, but there is no gallery of prebuilt agents or packaged templates; the adapters are workers rather than prebuilt roles. Bernstein docs Aug 2026 Sourcebernstein.readthedocs.io/en/latestread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
The web dashboard receives GitHub, Linear and Slack webhooks, a daemon installer keeps the chat bridge running after reboot, and the autofix daemon turns a failing CI run on a Bernstein-opened pull request into a fix commit, so work reaches agents from events without a person starting it; runs also start from the CLI, from tickets, from chat commands and from MCP clients. SourceBernstein, bernstein.run/blog/operator-commands and llms.txtread 2026-09-21 |
|
| Model Flexibility & Routing | Full |
|
The customer declares each agent's adapter and model in bernstein.yaml, adapters are interchangeable across 40 plus CLI agents, local and cloud models can mix in one run, and an optional contextual bandit router picks a model and effort level per task under USD cost caps. That is customer-controlled model choice. SourceBernstein, bernstein.run llms.txt and bernstein.readthedocs.ioread 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
MCP server mode over stdio, HTTP, and SSE plus a stateless chain anchored MCP core and an MCP Tasks surface, a task server REST API with an OpenAPI reference, a documented adapter guide for adding custom agents, and a generic CLI wrapper; Apache 2.0 source is itself extensible. Bernstein docs Aug 2026 Sourcebernstein.readthedocs.io/en/latest/reference/openapi-referenceread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
Every agent diff passes a janitor verification stage before merge, running the repository's tests, types and lint plus an optional cross-model review, and tournament runs compare candidate outputs and select one with no model in the decision path, so only verified diffs land. That is a quality gate in the release path returning a verdict on the agent's work. SourceBernstein, bernstein.run homepage and llms.txtread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
No browser or computer use capability documented; the domain is CLI coding agents operating on repositories in worktrees and sandboxes. Bernstein docs Aug 2026 Sourcebernstein.readthedocs.io/en/latestread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Pricing
Free and open source (Apache 2.0)
none
Included quota
No quotas or tiers; full feature set ships in the Apache 2.0 build including cluster mode, air gap profile, and the MCP and REST surfaces
What is public
Everything: the tool is Apache 2.0 with public source, public docs, public benchmarks, and no paid tier anywhere
Billing mechanics
No billing; costs accrue only in the orchestrated agents (Claude Code, Codex, Gemini CLI and others) and any paid sandbox or storage backends the user configures; built in cost aware dispatch can enforce a USD cap per run
Cost watchouts
Parallel orchestration multiplies underlying agent subscription and API costs; sandbox backends E2B and Modal are paid third party services if used; solo maintainer with no commercial support offering
Variable cost rationale
The orchestrator is free but total run cost is the metered spend of the underlying coding agents it fans out in parallel, which scales with concurrency and model choice; the project exists because parallel agent bills reached hundreds of dollars a month, and it ships cost aware dispatch with USD caps as the control
Overage / add-ons
Not applicable; no paid product exists
Sales call required
No, self serve available
Free / trial
Entirely free; installs via pip, pipx, uv, brew, or npm
Lowest paid plan
None; free and open source
Commercial notes
Open source under Apache 2.0, built by a single maintainer, Alex Chernysh, with GitHub Sponsors and OpenCollective sponsorship and no commercial support offering; the homepage calls it beta software.
Key ambiguities
None on price; the open question for buyers is support and continuity, since no commercial entity stands behind the project
Missing data
No commercial support, SLA, or paid escalation path exists to price
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Bernstein
The closest documented capability profiles to Bernstein among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Braintrust10.0 / 14Fuller documented coverage on Security, Identity & Governance
- Inngest10.0 / 14Fuller documented coverage on Security, Identity & Governance
- LangWatch10.0 / 14Fuller documented coverage on Security, Identity & Governance
- Thread AI11.0 / 14Adds documented Knowledge Grounding & RAG
- Windmill12.0 / 14Adds documented Browser & Computer Use
- LangSmith11.5 / 14Fuller documented coverage on Security, Identity & Governance and Memory & State Persistence
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded