Agentic Index

cubic vs Graphite (2026)

Two of the most affordable entries in the cluster, at 8.5 and 7 of 14, solving different halves of the pull request problem. That verdict is the Agentic Index coverage score, graded from each vendor's own published materials.

Cubic reviews every pull request with whole codebase context, enforces plain English rules and runs background scans that find and fix, free for twenty reviews a month and free for public repositories. Graphite restructures the workflow around stacked pull requests with a stack aware merge queue and an agent that reviews, chats, fixes and merges, with a free hobby tier. Cubic makes reviews better; Graphite makes them smaller, which is arguably the more effective fix.

On the Agentic Index coding agent ranking, neither cubic nor Graphite clears the bar, which asks for all five merge loop capabilities documented in full. cubic does not document observability and auditability in full; Graphite does not document observability and auditability in full, nor human oversight and guardrails. 4 of the 63 vendors in the lane clear it. See the coding agent ranking

This comparison is published by Agentic Index, an independent agentic AI vendor research platform. cubic and Graphite are each graded against the same 14 capability Agentic Index taxonomy, from the vendor's own public materials under the Agentic Index verification standard, alongside 969 researched vendors. No vendor pays for placement and no vendor has reviewed this page. How this evidence is graded

Choose cubic if

  • Documented coverage is slightly broader and review quality is the immediate need.
  • Plain English rules are how your team will keep conventions current.
  • Background scanning outside the pull request finds what review never sees.

Choose Graphite if

  • Your pull requests are too big to review properly and stacking is the structural answer.
  • A merge queue that understands stacks removes a real source of friction.
  • You want the workflow fixed, not just better commentary on it.
At a glance cubic Graphite
Category Coding agent Coding agent
Entry price Free · Team $30/dev/mo billed annually ($40 monthly) Free (Hobby: stacking CLI, VS Code, limited AI reviews) · Starter $20/user/mo annual · Team $40/user/mo annual (unlimited Graphite Agent + merge queue) · Enterprise custom
Free / trial Free Starter plan: 20 PR reviews/month, up to 5 custom agents, automatic PR descriptions, custom context, Jira/Linear/Asana/Notion connectors and unlimited AI wikis. Unlimited and free for public repositories. Paid tiers carry a 7-day trial. Free Hobby tier includes the stacking command line tool, the VS Code extension, and a limited amount of Graphite Agent AI review. Every paid plan includes a 30 day free trial that does not require a credit card.
Pricing confidence public exact public exact
Feature
c
cubic
G
Graphite
Action & orchestration

Integrations & Tool Calling

Ability to connect agents to real systems through native integrations, OAuth-authenticated actions, custom tools, APIs, webhooks, or MCP-compatible tools.

Full / Explicit

Stands at F, re-based off first-party pages after the June basis cited aitools.inc. Breadth across classes is met on ticketing, documentation, chat, source control and IDE, and the connector set is now confirmed as named plan features with Confluence gated to Pro rather than available on all tiers as the record implied.

Full / Explicit

Stands at F, re-based off first-party pages after the June basis cited git-tower.com. Recorded rather than credited to another axis: Cursor Cloud Agents now run inside the Graphite pull request, which is a post-acquisition integration and the clearest product evidence that the two are converging.

Workflow Orchestration

Ability to sequence, branch, retry, route, and combine deterministic workflow nodes with autonomous agent steps.

Partial

Stands at P, re-based off first-party pages after the June basis cited a Medium post. Held at P rather than F on the same reading applied to blink-new: parallel scanning agents with no documented coordination are not multi-agent orchestration. The review-to-fix-to-ticket chain is a genuine multi-step pipeline, and Fix with coding agents delegating to an external agent is a real handoff, which is why this is not lower.

Full / Explicit

Upgraded from P. The June basis called this human-collaborative and rules-based rather than an orchestration engine, sourced to a vendor blog post. Reading the plan matrix, the orchestration is real and layered: a stack-aware merge queue with basic and advanced tiers, an automations engine, a CI optimizer, and automatic rebasing of dependent branches as changes land. Sequencing dependent pull requests to keep trunk green while rebasing the chain is genuine multi-step workflow execution over a dependency graph, which is what the axis measures. Recorded honestly: this orchestrates the merge pipeline rather than coordinating multiple agents, so it earns F on the multi-step half rather than the multi-agent half.

Triggers & Channel Coverage

How agents wake up and where they work: schedules, webhooks, message events, CRM events, inbox events, chat, email, voice, and collaboration tools.

Partial

Stands at P, re-based off first-party pages. Trigger coverage is genuinely good for a review product: pull request events, scheduled scans, interactive tagging, local CLI and IDE hosts. Held below F because every path is anchored to GitHub and the local development loop, with no chat-initiated or ticket-initiated invocation; Slack and email are notification outputs rather than inbound channels, which is the distinction that separates this from warp and goose at F.

Partial

Stands at P, re-based off the pricing matrix. Trigger coverage is real within one envelope: pull request events, merge queue events, CI events, plus CLI, VS Code, MCP, web and inbox surfaces, with Slack as a notification output. Held below F because everything is anchored to GitHub and the local development loop, with no chat-initiated or ticket-initiated invocation and no scheduled runs, which is the distinction separating this from warp, goose and ellipsis at F.

Knowledge & context

Knowledge Grounding & RAG

Ability to ground agent behavior in company data through document ingestion, retrieval, external knowledge APIs, semantic search, or RAG layers.

Full / Explicit

Stands at F, re-based off the docs after the June basis cited a YC profile. Cross-repo reviews and current library documentation are both new to the record and both strengthen the grade: grounding extends past the repository boundary in two directions, into companion repositories and into upstream library docs.

Full / Explicit

Stands at F but confidence is lowered from high to medium, because the June basis was graded at high confidence off braintrust.dev, a third-party vendor's customer case study, which the conventions exclude. Whole-codebase context and code indexing are confirmed first-party, but the mechanism is not documented anywhere I reached: no index architecture, retrieval method or scope statement appears on the pricing page or in navigation. The grade holds on the documented existence of code indexing plus admin controls over it; the confidence reflects that the depth claim rests on marketing rather than docs. A fetch of graphite.com/features/ai-reviews would settle it.

Memory & State Persistence

Ability to persist context across a run, conversation, workflow, user, team, or longer-term memory layer.

Partial

Upgraded from N. The June basis reasoned that self-learning was feedback adaptation to be counted under Know rather than memory, which was a defensible reading of the axis but is now overtaken by the docs, which describe it as a discrete Learns from you feature: the AI remembers reactions and corrections and applies them to reduce false positives over time. That is learned state accumulating across sessions, which is the axis. The one fact does not do double duty either, because Know rests separately on repository-wide context, cross-repo reads and library documentation.

No / Not documented

Stands at N on re-retrieval. The reviewer tuning itself from accepted and rejected comments is real but is feedback-loop adaptation of a model's behaviour rather than a documented memory layer, and imported style guides and rules are static configuration the customer maintains. Contrast cubic, moved to P because the docs name a Learns from you feature with a described mechanism, and cosine, moved to F for a named Memory feature persisting architecture decisions. Nothing comparable is named here. The distinction is the vendor documenting persistence as a capability rather than a training characteristic.

Control & trust

Human Oversight & Guardrails

Approval steps, consent checkpoints, escalation rules, structured guardrails, policy constraints, and pause/resume controls.

Full / Explicit

Upgraded from P, and the mechanism the June basis missed is the strongest oversight evidence in this lane after warp. Auto-approval with a shadow mode is a genuine shipped approval gate with a safe rollout path: the customer watches which pull requests cubic would have approved before granting it authority to approve any. That is the shipped mechanism the axis asks for, not a governance posture. Ultrareview adds customer-controlled escalation of scrutiny on risky changes.

Partial

Stands at P. Oversight is largely structural, since a reviewer's output is advisory findings a human accepts or rejects, which is the same reading applied to qodo. What lifts it toward the gate the axis wants is the merge queue: approved pull requests are sequenced and merged under rules the team configures, with advanced settings on Enterprise, and ACLs bound who can do what. Held below F because no per-action approval gate on the agent's own actions is documented, and no shadow or preview mode exists for the kind cubic ships on auto-approval.

Security, Identity & Governance

RBAC, SSO, auditability, encryption, least-privilege tool access, compliance posture, and data handling policy.

Full / Explicit

Upgraded from P, and the June basis explicitly conceded what it could not find: it said a full identity and governance matrix was not enumerated, sourced to aiagentslist.com. The pricing page enumerates it. The axis conjunction is now met on both halves, an attestation plus multiple named customer-facing controls, which is the same reading applied to qodo and adopt-ai in this review. Confidence stays high but the SOC 2 claim itself comes from vendor marketing rather than a trust centre; no trust portal or named audit firm was retrieved, and that is recorded rather than smoothed over.

Full / Explicit

Upgraded from N. The June basis said no enumerated identity and governance matrix is documented, sourced to the privacy and security docs page; the pricing page's Admin section enumerates it in full. The axis conjunction is met twice over: SOC 2 Type II plus continuous penetration testing on the attestation side, and SAML/SSO, ACLs, SIEM audit log export, code indexing controls, AI privacy controls and GHES support on the named-control side. This is the Sec understatement pattern the process note predicts, and it was found by reading the plan comparison table rather than the security page.

Observability & Auditability

Traces, logs, execution histories, metrics, audit events, and debugging detail for production agent behavior.

Partial

Stands at P, re-based off first-party pages after the June basis cited aitools.inc. Analytics and an Analytics API are real reporting surfaces and exportable compliance audits are a genuine governance artefact, so this is comfortably above N. Held below F on the lane-wide reading: these report on code, review throughput and AI coding usage, not on what the agent itself did and why. There is no retained per-action trace of a review or a background fix. Same distinction that took augment-code and sourcegraph to P.

Partial

Stands at P, re-based off the pricing matrix after the June basis cited a vendor blog post. Insights and custom analytics are genuine reporting, and the Enterprise SIEM audit log is a real governance artefact, but both measure people and access rather than agent behaviour. Under the lane-wide reading applied to augment-code and sourcegraph, reporting on throughput and cycle time is not a retained per-action trace of what the reviewer did and why. Would move to F on a documented per-review execution record.

Deployment & Data Residency

Deployment modes and options, including SaaS, dedicated cloud, VPC, on-prem, hybrid, local runtime, and self-hosting.

No / Not documented

Stands at N on re-retrieval. The CLI running locally is a real addition the June basis missed, but it is a client against cubic's cloud rather than a deployment option, and the vendor is explicit that CLI findings differ from cloud review. No self-hosted, VPC, on-premises or region-selectable option appears on the pricing page including Enterprise, which is where such an option would be listed and where GitHub Enterprise support and custom MSA and DPA terms are listed instead. Searched: pricing including Enterprise, the docs key features page and the enterprise page in navigation.

No / Not documented

Stands at N, but the basis is corrected: the June note said no deployment option beyond hosted is documented, sourced to git-tower.com, and GitHub Enterprise Server support is in fact listed on Enterprise. That is not a deployment option for Graphite, though, it is support for a customer-hosted code host, so the grade does not move. This is the same distinction drawn on blink-new, where portability of the artifact did not confer residency on the platform. No region selection, VPC, private instance or self-host of Graphite itself is documented.

Solution readiness

Prebuilt Agents, Templates & Packs

Ready-made workflows, packaged employees, templates, blueprints, industry solutions, and role-specific agents that reduce time-to-value.

Partial

Stands at P, re-based off first-party pages after the June basis cited a Medium post. Custom agents are genuinely reusable and the docs show a rules library UI, but they are authored by the customer rather than shipped as a vendor catalogue, and the tier caps of five and ten make clear these are configuration slots rather than a marketplace. That is the same reading that held cosine at N for vendor-built components and lifted codebuff to F for a public Agent Store; cubic sits between the two.

No / Not documented

Stands at N on re-retrieval, and this is a considered hold rather than a default. AI review customization covering automations, filters and rules is genuine reusable configuration, and it is what lifted comparable cells to P elsewhere; but the axis rewards prebuilt assets the customer installs or adopts, and every one of these is authored by the customer from scratch. Graphite ships a single reviewer agent with no library, gallery, template set or marketplace. Same reading that held cosine at N for vendor-built components. Searched: the pricing matrix, the features navigation and the homepage.

Platform extensibility

Model Flexibility & Routing

Ability to work across multiple foundation models, route tasks to different models, or let buyers bring their own providers and keys.

No / Not documented

Stands at N on re-retrieval, and the temptation to move it was real. Enterprise lists bring your own API keys, which is adjacent to this axis, but the axis measures whether the customer chooses the model that powers the product, and nothing documents model selection: Ultrareview escalates to cubic's most capable models by cubic's choice, not the customer's. Supplying a key is not choosing a model. Recorded rather than credited, and this is the best-evidenced form of N: the vendor discusses models extensively in its own research and still exposes no picker.

No / Not documented

Stands at N on re-retrieval, and the temptation to move it was checked. AI privacy controls and code indexing controls are admin capabilities over data handling, not model selection, and the same distinction was applied to cubic, where bring-your-own API keys did not lift Model because supplying a key is not choosing a model. No provider is named as powering reviews on any first-party surface. Worth flagging as a likely mover: Cursor's stated plan on acquisition was to leverage its coding models to make Graphite's AI features more intelligent and to merge Graphite's reviewer with Cursor's Bugbot, so the model layer here may change materially.

APIs, SDKs & MCP Extensibility

Composability layer: stable APIs, SDKs, MCP tool consumption/serving, custom tools, and integration into internal systems.

Full / Explicit

Upgraded from N, which was graded from absence off the marketing homepage. Two MCP servers plus an Analytics API plus a documented CLI is a real outward extensibility surface, and the MCP direction test is satisfied cleanly: cubic ships servers that other agents call, which is Ext credit. Bring your own API keys is recorded here but deliberately not used to lift Model, since it appears only as an Enterprise line item with no documented model selection.

Partial

Upgraded from N. The June basis said no MCP server is documented, sourced to git-tower.com, a third-party blog; the pricing page lists MCP as a Hobby-tier feature on every plan, grouped under Stacking alongside the CLI and VS Code extension. Held at P rather than F because the grouping indicates an MCP server exposing Graphite's stacking operations to other agents rather than a general platform API, and no REST API, SDK or webhook surface is documented anywhere. F would need a documented API or SDK for building on Graphite.

Testing, Debugging & Optimization

Testing, debugging, scoring, retries, fallbacks, quality gates, and optimization loops for improving agent workflows before and after deployment.

Full / Explicit

Stands at F, re-based off first-party pages after the June basis cited cubic.dev, and the grade is on firmer ground than it was. This is one of the rarer cases where the axis and the product coincide, as with snyk and qodo elsewhere in this lane: the customer points cubic at their own code and it tests it. Recorded but not credited: the vendor's #1 ranking on Martian's benchmark at 61.8 percent F1 is a vendor-reported claim about a third-party benchmark, which is marketing rather than a customer-facing harness.

Full / Explicit

Stands at F on the qodo and cubic precedent, where the axis and the product coincide: the customer points a shipped review product at their own code and it finds defects. Recorded honestly and worth noting against the seven downgrades this axis has taken in this lane: those were agents running a project's existing tests, whereas this is a defect-detection product with customer-configurable rules and filters. No harness for evaluating agent behaviour exists, which is the F bar on the other reading of this axis, so this F rests on the product-coincides reading alone.

Specialist automation

Browser & Computer Use

Browser, desktop, or remote/local computer control for workflows that cannot be handled through stable APIs alone.

No / Not documented

DOWNGRADED from P, the fifth correction of this same error in this cohort after blink-new, codebuff, compyle and cosine. The June basis credited reviewing code in an isolated container while navigating with grep and jump-to-definition as sandbox computer use; a container and a code index are programmatic interfaces, which is exactly what the axis excludes. Both passes reached the docs key features page, the pricing page and the marketing site and found no browser, screenshot or GUI capability. Graded as not documented rather than asserted absent.

No / Not documented

DOWNGRADED from P, the seventh correction of this identical axis error in the 30 June cohort. The June basis credited reviewing and navigating code, applying commits, rebasing branches via the CLI and self-healing CI as computer use; every one of those runs through git and the GitHub API, which are programmatic interfaces and exactly what the axis excludes. Both passes reached the pricing matrix, the features navigation and the homepage and found no browser, screenshot or GUI automation anywhere. Worth watching rather than assuming: Cursor Cloud Agents now run inside the pull request, and if those agents carry browser tooling this could change, but that capability belongs to cursor's record and is not documented on any Graphite surface.

Pricing snapshot

Sourced from the Index pricing dataset · open each vendor's profile for full detail.

Pricing cubic logocubic Graphite logoGraphite

Entry price

Lowest public entry point

Free · Team $30/dev/mo billed annually ($40 monthly) Free (Hobby: stacking CLI, VS Code, limited AI reviews) · Starter $20/user/mo annual · Team $40/user/mo annual (unlimited Graphite Agent + merge queue) · Enterprise custom

Pricing confidence

How public the numbers are

Public, exact Public, exact

Billing

Primary billing axis

Per-developer/month subscription tiers with a monthly reviewed-line allowance (added/deleted diff lines cubic reads); free tier capped by monthly review count; enterprise custom Per user per month subscription tiers billed annually, with a free tier and a 30 day trial. The Team tier includes unlimited AI reviews and chat with no per review metering; enterprise is custom.

Variable cost

Workload / overage exposure

Medium variable cost Low variable cost

Free tier / trial

Try before you buy

Free tierTrial
Free tierTrial

Buying motion

Self-serve vs sales call

Self-serve Self-serve

Other matchups in coding agents

Not the pairing you were after? These compare a different set of coding agents on the same 14 capabilities.

See all 93 coding agents comparisons

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.