Cosine
Also known as: Cosine AI, Genie, Cosine Genie, Lumen, Lumen Scout, Lumen Outpost, Lumen Sovereign, Cosine CLI, Cosine Cloud, Red Team
UK sovereign AI lab whose Lumen coding models power a terminal CLI and cloud workspace, with Swarm mode multi-agent orchestration, a 21-model menu across eight providers, cross-session memory, and fully air-gapped deployment for regulated industries.
Cosine describes itself as The Sovereign AI Lab: it trains AI models and specialized coding agents for organizations that need frontier capability inside secure environments. Founded in 2022 and backed by Y Combinator, it is now based in London and positions around regulated and sovereign deployment rather than the open benchmark race it first became known for. Its Genie agent drew wide attention in 2024 for topping SWE-bench at roughly 30 percent, the highest by any company at the time; the Genie branding has since been retired and the company now leads with the Lumen model family.
Lumen is a family of coding models rather than a single one. Scout is post-trained from Devstral 123B and runs cheaply and on-device for the work around the coding loop. Outpost is post-trained from Kimi K2.6 and handles everyday production implementation, trained with behavioral reinforcement learning specifically against the failure modes developers complain about: messy patches, duplicated logic, unnecessary abstraction, and agents that claim a task is finished before it is.
Sovereign, aimed at frontier-scale reasoning, is announced but not yet shipped, targeted for late 2026 and to be trained on UK soil using the Isambard-AI supercomputer under the UK Government's sovereign AI program. The models are proprietary and post-trained for enterprise and niche languages including COBOL, Fortran, Verilog, Rust and complex SQL. Cosine evaluates them on three internal benchmarks of its own design: Niche-Bench for legacy-language robustness, Slop-Bench for maintainability and Vibe-Bench for collaboration.
The agent reaches developers through two surfaces. The CLI installs via Homebrew and runs terminal-native with a local-to-remote execution model, moving between planning, implementation and review modes. Every agent turn is a lightweight git commit, so any change can be undone instantly.
A live todo list tracks complex jobs, Memory persists project conventions and architecture decisions across sessions, language servers supply symbol-level understanding for references, definitions, renames and diagnostics, and Swarm mode spawns specialized child agents that break a task down and execute in parallel. Cosine Cloud runs multiple tasks at once in a shared project across engineers, product managers and stakeholders, and remote agents keep the terminal local while long-running work executes in the cloud. MCP connections are auto-installed with zero configuration, covering GitHub, Jira, Postgres, Figma, Stripe, Slack and Linear.
The company is explicit about not locking customers into its own models: the CLI ships a menu of twenty-one models across eight providers, with per-model credit multipliers ranging from Lumen Scout at 0.1x to GPT 5.5 and Claude Opus at 2.75x.
Deployment is the commercial center of gravity. Cosine offers fully managed public cloud, a managed single-tenant environment with enterprise controls, and fully air-gapped installation inside the customer's own perimeter, with Scout available on-device. That posture underpins an industry coalition co-designing Lumen Sovereign whose members include BAE Systems, Babcock, BT, HSBC, Lloyds, LSEG, NatWest, PwC, Thales UK, Leonardo UK, QinetiQ, Fujitsu, Deloitte, Vodafone, EY and The Alan Turing Institute.
Pricing is credit-based across three published tiers: Starter at $19 a month with 4M credits, Team at $199 with 47M, and Enterprise at $999 with 240M, with add-on credits priced per million and enterprise or private deployment scoped through sales. Credits are consumed by agent work, model calls and cloud execution, so cost varies with task size and model choice. Notable gaps for a vendor selling into regulated industries: no security attestation, trust center, SSO or RBAC documentation appears anywhere on the site, and no public API or SDK for embedding Cosine is published.
Vendor details
Canonical URL
https://cosine.sh
Category
Coding agent
Subcategory
Sovereign coding agent with proprietary models and air-gapped deployment
Funding status
Independent. Founded 2022, Y Combinator backed, and now describing itself as the UK's sovereign AI research lab, based in London. Its Genie agent topped SWE-bench in 2024 at roughly 30 percent, the highest by any company at the time. Cosine is backed by the UK Government's sovereign AI program and leads an industry coalition co-designing Lumen Sovereign, with MoUs reported from BAE Systems, Babcock International, BT, HSBC, Lloyds Banking Group, LSEG, NatWest, PwC, Thales UK, Leonardo UK, Telefonica Tech UK&I, QinetiQ, Fujitsu, Deloitte, Vodafone, EY and The Alan Turing Institute. Lumen Sovereign is to be trained on UK soil using the Isambard-AI supercomputer.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
MCP connections are auto-installed with zero configuration, covering GitHub, Jira, Postgres, Figma, Stripe, Slack and Linear across source control, ticketing, database, design, payments and chat. Language servers give the agent symbol-level understanding for references, definitions, renames and diagnostics. The CLI installs via a Homebrew tap and works alongside the editor, tests and package manager with a local-to-remote execution model; Cosine Cloud runs tasks in parallel in a shared project. The model menu spans twenty-one models across eight providers with per-model credit multipliers.
In practice
You have a backlog of well scoped bug tickets. Assign them to Genie in Jira or GitHub and it plans, writes, and tests each fix, then opens a pull request for your team to review.
Your code cannot leave your perimeter for compliance reasons. Cosine offers air gapped deployment and on device Lumen models, so an autonomous coding agent can run entirely inside your own security boundary.
A new engineer needs to understand a sprawling legacy codebase. Cosine indexes it as a graph plus semantic search and answers natural language questions about how files and functions actually connect.
Sources & related URLs
Related / legacy domains
Agentic Index coverage score
8.5 / 14 capabilities · 61%
| Integrations & Tool Calling | Full |
|---|---|
|
Auto installed MCP connections cover GitHub, Jira, Postgres, Figma, Stripe, Slack and Linear, language servers supply references, definitions, renames and diagnostics, and the agent runs terminal commands in the developer's shell. Sourcecosine.sh/cli and cosine.sh/docs/customizing/modesread 2026-09-29 |
|
| Workflow Orchestration | Full |
|
Swarm mode makes the primary agent an orchestrator that delegates work to subagents running in parallel, and Cosine Cloud runs long tasks remotely while the terminal stays local. Sourcecosine.sh/docs/customizing/modes and cosine.sh/cliread 2026-09-29 |
|
| Knowledge Grounding & RAG | Partial |
|
Context is gathered on demand, loading relevant files and using language server operations such as go to definition and find references, with MCP connections for outside systems; no persistent index or embeddings layer over the codebase is documented. Sourcecosine.sh/docs/concepts/memory-and-context and cosine.sh/cliread 2026-09-29 |
|
| Human Oversight & Guardrails | Full |
|
Manual mode, the default, asks for confirmation before every mutating action (edits, file operations, terminal commands and MCP tool calls), and Plan mode is read only until the user chooses how the plan is carried out; auto mode is an opt in, and every turn is a revertible git commit. Sourcecosine.sh/docs/customizing/modesread 2026-09-29 |
|
| Security, Identity & Governance | Partial |
|
Security rests on deployment control, a managed single tenant cloud with enterprise controls and a fully air gapped option. No SOC 2 or ISO certification, SSO or role based access is published. Sourcecosine.sh and cosine.sh/docsread 2026-09-29 |
|
| Observability & Auditability | Partial |
|
A live task list shows what the agent intends and has finished, and every turn is recorded as a git commit in the customer's repository; no vendor side run history, trace of tool calls or audit log is documented. Sourcecosine.sh/cliread 2026-09-29 |
|
| Memory & State Persistence | Partial |
|
The agent saves reusable facts with a save_memory tool into .cosine/agents.md, a project scoped file that persists across sessions and loads at the start of each, beside the team's own AGENTS.md; no lifetime, expiry or purge path is published. Sourcecosine.sh/docs/customizing/memory and concepts/memory-and-contextread 2026-09-29 |
|
| Deployment & Data Residency | Full |
|
Three deployment tiers are published: fully managed public cloud, a managed single tenant cloud with enterprise controls, and fully air gapped inside the organization's own security perimeter, with Lumen Scout sized for on device use. Sourcecosine.shread 2026-09-29 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
Swarm mode's specialized subagents are chosen by the orchestrator rather than selected by the customer, so they are the vendor's own machinery; no catalog of prebuilt agents or templates a customer adopts is documented. Sourcecosine.sh/docs/customizing/modes and cosine.sh/cliread 2026-09-29 |
|
| Triggers & Channel Coverage | Partial |
|
Work starts from the three documented surfaces, the terminal CLI, Cosine Cloud and Desktop, each invoked by a person; no schedule, webhook or ticket event that wakes an agent on its own is documented. Sourcecosine.sh/docs/ways-to-use-cosineread 2026-09-29 |
|
| Model Flexibility & Routing | Full |
|
The CLI offers a model menu of more than 25 models across providers including OpenAI, Anthropic, Google, Kimi and Qwen alongside Cosine's own Lumen models, described as model agnostic by design, with the user choosing the model. Sourcecosine.sh/cliread 2026-09-29 |
|
| APIs, SDKs & MCP Extensibility | Not documented |
|
No API, SDK, headless mode or MCP server that lets an outside caller drive Cosine is documented; the docs name three surfaces, the CLI, Cloud and Desktop. MCP connections let Cosine reach other systems, which runs the other way. Sourcecosine.sh/docs/ways-to-use-cosine and cosine.sh/docsread 2026-09-29 |
|
| Testing, Debugging & Optimization | Partial |
|
Every agent turn is a git commit that can be reverted, and the vendor publishes its own Niche-Bench, Vibe-Bench and Slop-Bench results for its Lumen models; these measure Cosine's models and let a change be undone, and no harness a customer runs against agent behavior on its own work is documented. Sourcecosine.sh and cosine.sh/cliread 2026-09-29 |
|
| Browser & Computer Use | Not documented |
|
The agent works through shell commands, file edits, language servers and MCP connections; no browser control or operation of software through a graphical interface is documented. Sourcecosine.sh and cosine.sh/cliread 2026-09-29 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Pricing
Starter $19/mo (4M credits)
Flat monthly subscription with a bundled credit pool per tier, not per seat. Credits are consumed across agent work, model calls and cloud execution, so burn varies with task size, model choice and runtime; per-model credit multipliers range from Lumen Scout at 0.1x to GPT 5.5 and Claude Opus at 2.75x. Add-on credits are purchasable at any time on all tiers. Enterprise and private deployment pricing is scoped with sales because infrastructure, support and security requirements vary.
Included quota
Credit-based across three published tiers, not seat-based. Starter $19/mo: 4M credits/month, add-on credits $6.50 per 1M. Team $199/mo: 47M credits/month, add-on credits $5.00 per 1M. Enterprise $999/mo: 240M credits/month, add-on credits $4.50 per 1M. Credits represent usage across agent work, model calls and cloud execution, with actual consumption depending on task size, model choice and runtime. All tiers can buy top-ups at any time. Enterprise, private tenant and air-gapped deployments are scoped through sales.
What is public
All three tier prices and their credit allocations are published exactly, along with per-million add-on credit rates: Starter $19 with 4M credits, Team $199 with 47M, Enterprise $999 with 240M. Per-model credit multipliers are published on the CLI page. Not public: credit cost per task, and pricing for enterprise, single-tenant or air-gapped deployment, which is scoped with sales.
Billing mechanics
Monthly subscription tiers, each with a credit pool that does not roll over. Credits are spent per agent action, model call and cloud execution, with per-model multipliers, and add-on credits are priced per million.
Cost watchouts
Credit pool exhaustion pauses work (top-ups add cost); no rollover; Professional seats + top-ups scale quickly; BYO path shifts cost to your own model subscription.
Variable cost rationale
Cost combines a per-seat fee with a credit pool consumed by every read, plan, and write, and included credits do not roll over, so heavy or exploratory use can exhaust the pool and require top-ups. The gap between the Hobby (5M) and Professional (60M) credit allotments is large, and the $200 Professional seat plus top-ups can add up for active teams. The BYO-subscription path shifts model cost to your own Claude/OpenAI/Copilot plan, which can lower or change the exposure.
Additional watchouts
Credits do not roll over and can pause work when exhausted (top-ups needed). Team at $199 a month plus top-ups can be costly for active teams. No free tier. Enterprise and air-gapped pricing is opaque. Proprietary (not open source).
Overage / add-ons
Credits are consumed per agent action (read/plan/write); when the pool is exhausted, inference pauses until a top-up or the next cycle. Top-ups purchasable anytime. Unused included credits do not roll over.
Sales call required
No, self serve available
Free / trial
No free tier published; entry is the $19/month Starter plan
Lowest paid plan
Starter $19/month, including 4M credits
Commercial notes
Task-outcome oriented: credits track work done. The step from Starter (4M credits) to Team (47M) reflects heavier autonomous workloads. Enterprise value is in air-gapped and sovereign deployment and Lumen models for regulated industries.
Key ambiguities
Credit cost per task is not published and depends on task size, model choice and runtime, though per-model multipliers are published on the CLI page. Enterprise, single-tenant and air-gapped deployment pricing is quoted rather than listed.
Cancellation / refund
Monthly seat-based subscription; included credits do not roll over. Specific cancellation/refund terms not detailed.
Support SLA / resale
Enterprise agreements for dedicated tenant and air-gapped deployments; standard support on Hobby/Professional. UK Sovereign AI participant. No reseller/whitelabel program surfaced.
Missing data
Exact credit consumption per task type and enterprise/air-gapped pricing are not public; no free tier surfaced.
Related vendors
- Cognition — Maker of Devin, an autonomous AI software engineer
- 10Web — Agentic website platform whose specialized AI agents build, host,…
- AgentUI — Managed AI app builder for operations teams that generates and hosts…
- Aider — Open source, model agnostic terminal coding agent that edits your…
- Anthropic Claude Code — Anthropic's agentic coding system across terminal, desktop, IDE, web…
- AppFactor — Agentic platform that runs persistent agents across an enterprise…
Alternatives to Cosine
The closest documented capability profiles to Cosine among coding agents tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- OpenCode10.0 / 14Adds documented APIs, SDKs & MCP Extensibility
- Zed10.0 / 14Adds documented APIs, SDKs & MCP Extensibility
- Blitzy8.5 / 14Fuller documented coverage on Knowledge Grounding & RAG and Security, Identity & Governance
- JetBrains AI10.5 / 14Adds documented APIs, SDKs & MCP Extensibility
- Charm10.0 / 14Adds documented APIs, SDKs & MCP Extensibility
- CodeRabbit11.0 / 14Adds documented APIs, SDKs & MCP Extensibility
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded