Blackbox AI
Also known as: Blackbox, CyberCoder, Chairman LLM, Blackbox Router, k-agents
Coding agent platform and inference layer: an Agents API that sends Blackbox, Claude Code, Codex or Gemini agents to any repository over HTTP, an agentic CLI, and a router over 300+ models, deployable from single tenant SaaS to fully air gapped.
Blackbox AI, founded in 2019 and based in San Francisco, sells coding agents and the inference layer beneath them to organizations that need high trust AI. Its Agents API sends coding agents to any repository over HTTP: a team creates a run, streams live logs, reads and writes files in the sandbox, continues the conversation and gets a pull request with clean commits.
A multi agent task runs two to five agents, including Blackbox's own agent, Claude Code, Codex and Gemini, on the same task in parallel so their approaches and results can be compared, and Chairman LLM evaluates the implementations. An agentic command line tool reads a repository's conventions and handles coding, debugging, setup and deployment from the terminal, with skills teams write and share through version control, and a VS Code extension brings the agent into the editor.
Underneath, the Blackbox Router exposes more than 300 open and closed models through one OpenAI compatible endpoint with provider routing, prompt caching and cost aware routing, and Enterprise Inference runs an open weight model the customer chooses on reserved capacity.
The platform is built for regulated buyers: prompts are encrypted before they leave the developer's machine, zero data retention is enforced at the gateway, PII can be stripped before prompts reach a closed model, and deployment runs from multi tenant SaaS through a dedicated VPC in AWS, Azure or GCP to on premises and air gapped installs, pinned to US or EU regions. Enterprise adds SAML SSO, SCIM, fine grained RBAC over repositories, models and agent capabilities, and audit logs streamed to a SIEM; SOC 2 Type II and ISO 27001 audits are in progress. Pricing is enterprise only, on committed annual spend.
Vendor details
Canonical URL
https://blackbox.ai
Category
Coding agent
Subcategory
Coding — assistant
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
The Agents API reaches any GitHub repository over HTTP to create runs, stream logs and open pull requests, and a remote MCP server at agent.blackbox.ai lets the CLI coordinate cloud agents. The Blackbox Router exposes 300+ open and closed models through one OpenAI compatible endpoint with provider routing, audit logs stream to Splunk, Datadog, Sumo Logic and Elastic, and a Microsoft Azure partnership is documented for enterprise deployments.
In practice
You want to try several coding agents on the same task before trusting one. A Blackbox multi agent task runs its own agent, Claude Code, Codex and Gemini in parallel on your repository so you can compare their approaches, commits and results.
Your platform team wants to start coding agents from its own tools. The Agents API creates runs against any repository over HTTP, streams live logs and opens pull requests with clean commits.
Your code can't leave your network or region. Blackbox deploys to a dedicated VPC, on premises or fully air gapped, pinned to US or EU regions, with zero data retention and prompts encrypted before they leave the machine.
Sources & related URLs
Related / legacy domains
Agentic Index coverage score
7.5 / 14 capabilities · 54%
| Integrations & Tool Calling | Partial |
|---|---|
|
Agents work against GitHub repositories named per task and open pull requests, audit logs stream to Splunk, Datadog, Sumo Logic and Elastic, and a remote MCP server lets the CLI coordinate cloud agents. No MCP client support or catalog of third party tool integrations for the agents is documented. Sourceblackbox.ai/agents, blackbox.ai/enterprise and docs.blackbox.airead 2026-09-29 |
|
| Workflow Orchestration | Full |
|
A multi agent task runs two to five agents (Blackbox's own, Claude Code, Codex or Gemini, each with a chosen model) on the same repository task in parallel and records each agent's status, commits and result, Chairman LLM evaluates the implementations, and the CLI orchestrates sub tasks and parallel tasks across cloud agents. Sourcedocs.blackbox.ai/api-reference/multi-agent-task, blackbox.ai/agents and docs.blackbox.ai/features/blackbox-cloud-mcpread 2026-09-29 |
|
| Knowledge Grounding & RAG | Partial |
|
The CLI reads the conventions of the repository, the VS Code agent analyzes the whole project context, and skills stored in the workspace add team knowledge. No indexing, embedding or retrieval architecture over the customer's code or documents is described. Sourceblackbox.ai, docs.blackbox.ai CLI and VS Code pagesread 2026-09-29 |
|
| Human Oversight & Guardrails | Partial |
|
Admins constrain what agents may do through fine grained RBAC over repositories, models and agent capabilities, model allow listing and per workspace content policies with DLP rules, prompt filters and redaction, and agent work lands as pull requests. No step where a person approves an agent action before it commits is documented. Sourceblackbox.ai/enterprise and blackbox.ai/agentsread 2026-09-29 |
|
| Security, Identity & Governance | Partial |
|
The control surface is deep, with SAML 2.0 SSO through Okta, Azure AD, Google Workspace and OneLogin, SCIM 2.0, fine grained RBAC, zero data retention, prompts encrypted before they leave the machine and PII stripping before closed models. The enterprise page lists SOC 2 Type II and ISO 27001 as in progress, with interim letters under NDA, so no completed attestation is asserted. Sourceblackbox.ai/enterprise and blackbox.airead 2026-09-29 |
|
| Observability & Auditability | Partial |
|
The Agents API streams live run logs and a multi agent task records each agent's status, commits, result and errors, and audit logs stream to a customer SIEM. What the logs and audit events contain is not described, so no step by step record of the agent's tool calls and reasoning is documented. Sourceblackbox.ai/agents, blackbox.ai/enterprise and docs.blackbox.ai/api-reference/multi-agent-taskread 2026-09-29 |
|
| Memory & State Persistence | Not documented |
|
No agent memory is documented on the home, agents, enterprise, CLI, skills or VS Code pages; a run can be continued as a conversation, which is state inside one task. Retention windows are data retention settings and instruction files are instructions, so neither is memory. Sourceblackbox.ai/enterprise and docs.blackbox.airead 2026-09-29 |
|
| Deployment & Data Residency | Full |
|
Four deployment modes are documented: multi tenant SaaS, a dedicated VPC in AWS, Azure or GCP, on premises in the customer's datacenter and air gapped with no outbound internet, and data can be pinned to US or EU regions across all modes. Sourceblackbox.ai/enterpriseread 2026-09-29 |
|
| Prebuilt Agents, Templates & Packs | Not documented |
|
No prebuilt agent library, template set or bundled skills are published: CLI skills are written by the customer, and the agents a task can run (Blackbox's own, Claude Code, Codex, Gemini) are interchangeable coding agents doing the same job, which is agent and model choice rather than a pack. The CyberCoder and App Builder agents no longer appear on the vendor's current pages. Sourceblackbox.ai/agents and docs.blackbox.ai/features/blackbox-cli/skillsread 2026-09-29 |
|
| Triggers & Channel Coverage | Partial |
|
Runs start when a person or the customer's own code asks: a task described in the browser, the CLI, the VS Code extension or an Agents API call over HTTP. No schedule, webhook or repository event that wakes an agent on its own is documented. Sourceblackbox.ai/agents and docs.blackbox.airead 2026-09-29 |
|
| Model Flexibility & Routing | Full |
|
The Blackbox Router exposes more than 300 open and closed models through one OpenAI compatible endpoint with the model named per request and provider routing, a multi agent task sets the model per agent, and Enterprise Inference runs the open weight model the customer chooses on reserved capacity. Sourceblackbox.ai and docs.blackbox.ai/api-reference/multi-agent-taskread 2026-09-29 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
The Agents API is documented endpoint by endpoint (create, get, status, list, cancel and continue a task, and create a multi agent task), creating runs, streaming logs and managing sandbox files over HTTP, alongside an OpenAI compatible inference endpoint and a remote MCP server at agent.blackbox.ai with bearer token auth. Sourcedocs.blackbox.ai api-reference and features/blackbox-cloud-mcpread 2026-09-29 |
|
| Testing, Debugging & Optimization | Partial |
|
A multi agent task lets a team compare approaches, code quality and results across two to five agents on the same task, and Chairman LLM evaluates the implementations. This arbitrates outputs for one task, and no harness to test, evaluate or regression check agent behavior over time is documented. Sourcedocs.blackbox.ai/api-reference/multi-agent-task and blackbox.ai/agentsread 2026-09-29 |
|
| Browser & Computer Use | Not documented |
|
Agents work through repositories, sandboxes and APIs; no capability for an agent to drive a browser or operate software through a graphical interface is documented. Sourceblackbox.ai/agents and docs.blackbox.airead 2026-09-29 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Blackbox AI introduced a two-model orchestration architecture that combines GPT-5.6 Sol and Claude Opus 4.8. The system uses a sandbox-executing critic agent to evaluate responses and gate the final output, achieving improved accuracy on the Terminal-Bench v2.1 benchmark.
Bears on: Agent capability
View sourcePricing
Enterprise only: custom pricing on committed annual spend; no self serve plan
Committed annual spend drawn down by token usage across models.
Included quota
The committed balance covers token usage across 300+ models, with a dedicated forward deployed engineer and implementation at no extra cost.
What is public
Only an Enterprise plan is listed, priced on committed annual spend with discounts of up to ten percent on open weight models and five percent on closed models; no self serve tiers, free plan or trial are published.
Billing mechanics
The customer commits an annual spend, and the balance decreases as teams consume tokens across models, with alerts at 75 and 90 percent of the commitment.
Cost watchouts
Spend follows token usage against the commitment, so model choice and agent volume drive how fast the balance runs down.
Variable cost rationale
A committed balance is drawn down by usage, so cost is bounded by the commitment but depends on model mix and volume.
Additional watchouts
There is no self serve entry point: a sales conversation and a committed spend come first.
Overage / add-ons
Usage draws down the committed balance, with alerts at 75 and 90 percent; terms beyond the commitment are not published.
Sales call required
Yes, required for paid access
Free / trial
No free plan or trial; engagement starts with sales.
Lowest paid plan
None published; Enterprise only.
Commercial notes
Enterprise includes single tenant deployment, end to end encryption, contractual zero data retention, PII removal before closed models, SAML SSO, SCIM, RBAC, audit logs and the Agents API.
Key ambiguities
No dollar figures are published; cost depends on the committed spend and model mix.
Cancellation / refund
Annual committed spend agreed with sales; terms are not published.
Support SLA / resale
Priority compute/support on higher tiers; Enterprise adds SAML SSO, training opt-out by default, advanced security controls, and on-prem deployment
Missing data
Commitment minimums and per model token rates under the plan are not published.
Related vendors
- Cognition — Maker of Devin, an autonomous AI software engineer
- 10Web — Agentic website platform whose specialized AI agents build, host,…
- AgentUI — Managed AI app builder for operations teams that generates and hosts…
- Aider — Open source, model agnostic terminal coding agent that edits your…
- Anthropic Claude Code — Anthropic's agentic coding system across terminal, desktop, IDE, web…
- AppFactor — Agentic platform that runs persistent agents across an enterprise…
Alternatives to Blackbox AI
The closest documented capability profiles to Blackbox AI among coding agents tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- 10Web8.5 / 14Adds documented Prebuilt Agents, Templates & Packs
- camelAI6.0 / 14Adds documented Memory & State Persistence
- OpenCode10.0 / 14Adds documented Memory & State Persistence and Prebuilt Agents, Templates & Packs
- Tabby6.0 / 14Fuller documented coverage on Knowledge Grounding & RAG
- Zed10.0 / 14Adds documented Memory & State Persistence and Prebuilt Agents, Templates & Packs
- Cosine8.5 / 14Adds documented Memory & State Persistence and Prebuilt Agents, Templates & Packs
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded