Agentic Index

Baz vs cubic (2026)

Both review every pull request with whole codebase context and the gap is wide, 12 of 14 against 8.5. That verdict is the Agentic Index coverage score, graded from each vendor's own published materials.

Cubic reviews each pull request against the whole codebase, catches bugs, enforces rules written in plain English and runs background scans that find and fix issues, free for twenty reviews a month and free for public repositories. Baz covers the lifecycle with a benchmark leading review agent, specification validation and a pre code plan gateway. Cubic is the cheaper way to start; plain English rules are its best idea, because a rule your team can read is a rule your team will maintain.

On the Agentic Index coding agent ranking, Baz clears the bar and cubic does not. Baz documents all five merge loop capabilities in full; cubic does not document observability and auditability in full. 4 of the 63 vendors in the lane clear it. See the coding agent ranking

This comparison is published by Agentic Index, an independent agentic AI vendor research platform. Baz and cubic are each graded against the same 14 capability Agentic Index taxonomy, from the vendor's own public materials under the Agentic Index verification standard, alongside 969 researched vendors. No vendor pays for placement and no vendor has reviewed this page. How this evidence is graded

Choose Baz if

  • Documented coverage is nearly twice as deep across the matrix.
  • Governing agent written code across the lifecycle is what you are buying.
  • The pre code plan gateway is capability nothing else in this cluster has.

Choose cubic if

  • Free for twenty reviews a month and free for public repositories removes the decision.
  • Rules written in plain English is how your team will actually keep standards current.
  • Background scans that find and fix without waiting for a pull request suit your workflow.
At a glance Baz cubic
Category Coding agent Coding agent
Entry price Free trial and on demand starter on the GitHub Marketplace; paid team and enterprise pricing not fully public Free · Team $30/dev/mo billed annually ($40 monthly)
Free / trial Free trial and an on demand starter plan on the GitHub Marketplace Free Starter plan: 20 PR reviews/month, up to 5 custom agents, automatic PR descriptions, custom context, Jira/Linear/Asana/Notion connectors and unlimited AI wikis. Unlimited and free for public repositories. Paid tiers carry a 7-day trial.
Pricing confidence public partial public exact
Feature
B
Baz
c
cubic
Action & orchestration

Integrations & Tool Calling

Ability to connect agents to real systems through native integrations, OAuth-authenticated actions, custom tools, APIs, webhooks, or MCP-compatible tools.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Breadth across classes is comfortably met with three source platforms, six ticketing systems, design, and two knowledge bases, all dated in the changelog. The detail worth carrying to comparison pages is the Works alongside positioning: Baz publishes dedicated solution pages for Claude Code, OpenAI Codex, Cursor and Devin, which is the super-harness framing, integrating with coding agents rather than competing with them, the same shape as sonar in this cohort.

Full / Explicit

Stands at F, re-based off first-party pages after the June basis cited aitools.inc. Breadth across classes is met on ticketing, documentation, chat, source control and IDE, and the connector set is now confirmed as named plan features with Confluence gated to Pro rather than available on all tiers as the record implied.

Workflow Orchestration

Ability to sequence, branch, retry, route, and combine deterministic workflow nodes with autonomous agent steps.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Two things lift this beyond a standard multi-agent claim: the Planner loop runs BEFORE code exists, so orchestration covers plan generation, review and approval as well as implementation, and the Fixer agent closes its own loop by building an environment, validating and only committing when validation succeeds. Named infrastructure, the Agent Harness and Context Broker, is documented rather than implied, and Sessions trace the coordination.

Partial

Stands at P, re-based off first-party pages after the June basis cited a Medium post. Held at P rather than F on the same reading applied to blink-new: parallel scanning agents with no documented coordination are not multi-agent orchestration. The review-to-fix-to-ticket chain is a genuine multi-step pipeline, and Fix with coding agents delegating to an external agent is a real handoff, which is why this is not lower.

Triggers & Channel Coverage

How agents wake up and where they work: schedules, webhooks, message events, CRM events, inbox events, chat, email, voice, and collaboration tools.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.55 to 0.85. All three trigger classes are documented and dated: events through pull request webhooks, schedules through Skills Maintainer scans every three days, weekly or fortnightly, and on-demand through comment commands, the CLI and a slash command invoked from inside another coding agent. That last one is the distinguishing detail, since it means a developer starts a Baz planning session without leaving Claude Code or Cursor.

Partial

Stands at P, re-based off first-party pages. Trigger coverage is genuinely good for a review product: pull request events, scheduled scans, interactive tagging, local CLI and IDE hosts. Held below F because every path is anchored to GitHub and the local development loop, with no chat-initiated or ticket-initiated invocation; Slack and email are notification outputs rather than inbound channels, which is the distinction that separates this from warp and goose at F.

Knowledge & context

Knowledge Grounding & RAG

Ability to ground agent behavior in company data through document ingestion, retrieval, external knowledge APIs, semantic search, or RAG layers.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Grounding here is unusually broad because it spans three kinds at once: code structure through embeddings and graphs, organisational knowledge through Notion and Fibery, and product intent through six ticketing systems plus Figma. The Voyage AI code embeddings detail is a rare instance of a vendor naming its retrieval model. Cross-repository context with a visible cross repo tag on findings is the part most comparable to potpie-ai and greptile at F.

Full / Explicit

Stands at F, re-based off the docs after the June basis cited a YC profile. Cross-repo reviews and current library documentation are both new to the record and both strengthen the grade: grounding extends past the repository boundary in two directions, into companion repositories and into upstream library docs.

Memory & State Persistence

Ability to persist context across a run, conversation, workflow, user, team, or longer-term memory layer.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Two independent named memory systems is more than any other vendor in this lane documents, and both are dated and mechanically described rather than asserted: Reviewer Memory accumulates from repeated feedback and rewrites the reviewer's own system prompt with version history and rollback, and Module Memory persists implementation detail per module. The prompt-versioning-and-rollback detail is the part that makes this unambiguously F, since it means memory is a managed, inspectable artefact rather than a black box.

Partial

Upgraded from N. The June basis reasoned that self-learning was feedback adaptation to be counted under Know rather than memory, which was a defensible reading of the axis but is now overtaken by the docs, which describe it as a discrete Learns from you feature: the AI remembers reactions and corrections and applies them to reduce false positives over time. That is learned state accumulating across sessions, which is the axis. The one fact does not do double duty either, because Know rests separately on repository-wide context, cross-repo reads and library documentation.

Control & trust

Human Oversight & Guardrails

Approval steps, consent checkpoints, escalation rules, structured guardrails, policy constraints, and pause/resume controls.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Plan review approvals shipped 1 July 2026 and change the picture from the July basis, which credited only suggest-without-blocking and chat: a plan now goes through a dedicated approval workflow with versioning, pinned comments and approve or request-changes before implementation starts, which is a shipped gate ahead of code generation rather than review after it. Eighth distinct oversight architecture in this lane and the only one that gates the PLAN rather than the diff or the tool call.

Full / Explicit

Upgraded from P, and the mechanism the June basis missed is the strongest oversight evidence in this lane after warp. Auto-approval with a shadow mode is a genuine shipped approval gate with a safe rollout path: the customer watches which pull requests cubic would have approved before granting it authority to approve any. That is the shipped mechanism the axis asks for, not a governance posture. Ultrareview adds customer-controlled escalation of scrutiny on risky changes.

Security, Identity & Governance

RBAC, SSO, auditability, encryption, least-privilege tool access, compliance posture, and data handling policy.

Full / Explicit

Stands at F but the basis is materially rewritten, because the July basis graded partly from FOUNDER PEDIGREE, founded by cloud application security veterans, which is not evidence of anything the product or company does and should never have carried a grade. Re-retrieval found the real evidence the July build missed: SOC 2 certification stated in the site footer with a published trust centre at trust.baz.co and a security, privacy and compliance documentation page, plus role-based access control shipped September 2025. The conjunction is now met properly on attestation plus named controls. Confidence raised from a non-canonical 0.55 to 0.85 accordingly.

Full / Explicit

Upgraded from P, and the June basis explicitly conceded what it could not find: it said a full identity and governance matrix was not enumerated, sourced to aiagentslist.com. The pricing page enumerates it. The axis conjunction is now met on both halves, an attestation plus multiple named customer-facing controls, which is the same reading applied to qodo and adopt-ai in this review. Confidence stays high but the SOC 2 claim itself comes from vendor marketing rather than a trust centre; no trust portal or named audit firm was retrieved, and that is recorded rather than smoothed over.

Observability & Auditability

Traces, logs, execution histories, metrics, audit events, and debugging detail for production agent behavior.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.55 to 0.85. Sessions are a retained, structured execution record covering invocation, tool calls, stages, outcome and cost, which is the F bar on this axis as applied to ellipsis and poolside, and they extend across every agent rather than one. The product-level timeline framing rather than raw logs is worth recording: the vendor is deliberately building auditability as a reviewable artefact. The July basis graded this F on the marketing phrase observable and explainable, not a black box; the grade was right and the basis was not.

Partial

Stands at P, re-based off first-party pages after the June basis cited aitools.inc. Analytics and an Analytics API are real reporting surfaces and exportable compliance audits are a genuine governance artefact, so this is comfortably above N. Held below F on the lane-wide reading: these report on code, review throughput and AI coding usage, not on what the agent itself did and why. There is no retained per-action trace of a review or a background fix. Same distinction that took augment-code and sourcegraph to P.

Deployment & Data Residency

Deployment modes and options, including SaaS, dedicated cloud, VPC, on-prem, hybrid, local runtime, and self-hosting.

Partial

Upgraded from N and confidence normalised from a non-canonical 0.45. The July basis said no on-premise or self-hosted option is documented, which was graded from absence and is now contradicted by two dated first-party entries the July builder did not reach: Private Mode shipped 25 December 2025, and private application review running inside the customer's own Kubernetes cluster shipped 29 July 2026. Held at P rather than F under the early access convention, since private application review is explicitly a limited technical preview, and because Private Mode governs analysis and retention rather than giving the customer a full self-hosted deployment. Would move to F when private application review reaches general availability.

No / Not documented

Stands at N on re-retrieval. The CLI running locally is a real addition the June basis missed, but it is a client against cubic's cloud rather than a deployment option, and the vendor is explicit that CLI findings differ from cloud review. No self-hosted, VPC, on-premises or region-selectable option appears on the pricing page including Enterprise, which is where such an option would be listed and where GitHub Enterprise support and custom MSA and DPA terms are listed instead. Searched: pricing including Enterprise, the docs key features page and the enterprise page in navigation.

Solution readiness

Prebuilt Agents, Templates & Packs

Ready-made workflows, packaged employees, templates, blueprints, industry solutions, and role-specific agents that reduce time-to-value.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Eleven named prebuilt agents and reviewers is the largest shipped agent suite in this lane, and the Awesome Reviewers corpus is an unusual second layer: a free public library of 5,132 review instructions across twelve domains that a customer fetches as a markdown bundle straight into a skill folder. That is vendor-supplied adoptable assets in exactly the form this axis rewards, and it is published openly rather than gated behind the product.

Partial

Stands at P, re-based off first-party pages after the June basis cited a Medium post. Custom agents are genuinely reusable and the docs show a rules library UI, but they are authored by the customer rather than shipped as a vendor catalogue, and the tier caps of five and ten make clear these are configuration slots rather than a marketplace. That is the same reading that held cosine at N for vendor-built components and lifted codebuff to F for a public Agent Store; cubic sits between the two.

Platform extensibility

Model Flexibility & Routing

Ability to work across multiple foundation models, route tasks to different models, or let buyers bring their own providers and keys.

No / Not documented

Stands at N and confidence normalised from a non-canonical 0.5 to 0.65. The July basis said Bedrock foundation models with no customer choice, which re-retrieval supports but also refines: the changelog names GPT-5-Codex with Sonnet 4.5 reflections as the reviewer stack, so the vendor mixes providers, but that is the vendor's routing decision, not the customer's. Nothing documents model selection, bring-your-own-key or a provider setting anywhere. Recorded honestly: like sonar in this cohort, a low Model grade here reflects a deliberate product design where the vendor owns model choice as part of guaranteeing review quality, and the vendor does disclose which models it uses, which is better than the provider-undisclosed case the conventions treat as N.

No / Not documented

Stands at N on re-retrieval, and the temptation to move it was real. Enterprise lists bring your own API keys, which is adjacent to this axis, but the axis measures whether the customer chooses the model that powers the product, and nothing documents model selection: Ultrareview escalates to cubic's most capable models by cubic's choice, not the customer's. Supplying a key is not choosing a model. Recorded rather than credited, and this is the best-evidenced form of N: the vendor discusses models extensively in its own research and still exposes no picker.

APIs, SDKs & MCP Extensibility

Composability layer: stable APIs, SDKs, MCP tool consumption/serving, custom tools, and integration into internal systems.

Partial

DOWNGRADED from F on Mike's ruling of 2026-08-30, overturning my own August grade. Same error as zencoder: I held F while the basis itself recorded that no public REST API or SDK was retrieved, and flagged it as the weakest F on this axis in the lane rather than acting on that. Under the axis bar, full coverage means the platform is composable from outside through a documented API or SDK, with MCP as one form of evidence rather than the bar. The Baz surface is genuinely strong and that is why this is P rather than N: an MCP server exposing review tools to any IDE agent, a CLI with an interactive review loop, a slash command that starts Planner from inside another coding agent, and marketplace distribution. The Awesome Reviewers raw endpoints are a real documented API, but they serve a public corpus of review instructions rather than the Baz platform, which is the same distinction that correctly kept appfactor at P: MCP Bridge exposes the customer's systems, not AppFactor. Grade would move to F on a documented platform API or SDK. baz.ai/docs was not fetched on either pass and remains the cheap check at lane close.

Full / Explicit

Upgraded from N, which was graded from absence off the marketing homepage. Two MCP servers plus an Analytics API plus a documented CLI is a real outward extensibility surface, and the MCP direction test is satisfied cleanly: cubic ships servers that other agents call, which is Ext credit. Bring your own API keys is recorded here but deliberately not used to lift Model, since it appears only as an Enterprise line item with no documented model selection.

Testing, Debugging & Optimization

Testing, debugging, scoring, retries, fallbacks, quality gates, and optimization loops for improving agent workflows before and after deployment.

Full / Explicit

Stands at F and confidence normalised from a non-canonical 0.55 to 0.85, and this is one of the very few genuine F grades on this axis in the lane. The Eval rule is that the axis measures what the CUSTOMER can test, not what the vendor tests internally, and Prompt Playground is exactly that: a shipped agent test lab where a customer edits a reviewer's prompt and runs it against real change requests before rollout. Evaluations closes the loop by classifying every comment as accepted, rejected or unaddressed and feeding that back into reviewer behaviour. Contrast the nine vendors held at P in this lane for agents that merely verify their own output.

Full / Explicit

Stands at F, re-based off first-party pages after the June basis cited cubic.dev, and the grade is on firmer ground than it was. This is one of the rarer cases where the axis and the product coincide, as with snyk and qodo elsewhere in this lane: the customer points cubic at their own code and it tests it. Recorded but not credited: the vendor's #1 ranking on Martian's benchmark at 61.8 percent F1 is a vendor-reported claim about a third-party benchmark, which is marketing rather than a customer-facing harness.

Specialist automation

Browser & Computer Use

Browser, desktop, or remote/local computer control for workflows that cannot be handled through stable APIs alone.

Full / Explicit

Upgraded from P and confidence raised from a non-canonical 0.45 to 0.85. The July basis said computer use is limited to validation, which undersold it: this is the most thoroughly documented browser-driven capability found anywhere in this lane, built up across four dated releases from November 2025 to July 2026, and it is the product's differentiator rather than an accessory. The axis test is met exactly, since a rendered user interface has no programmatic interface and the agent operates it as a human would, clicking through flows and reading state off the screen. Contrast the thirteen downgrades in the June cohort, every one of which credited a terminal or sandbox; here the sandbox is where the browser runs, and the browser is the point.

No / Not documented

DOWNGRADED from P, the fifth correction of this same error in this cohort after blink-new, codebuff, compyle and cosine. The June basis credited reviewing code in an isolated container while navigating with grep and jump-to-definition as sandbox computer use; a container and a code index are programmatic interfaces, which is exactly what the axis excludes. Both passes reached the docs key features page, the pricing page and the marketing site and found no browser, screenshot or GUI capability. Graded as not documented rather than asserted absent.

Pricing snapshot

Sourced from the Index pricing dataset · open each vendor's profile for full detail.

Pricing Baz logoBaz cubic logocubic

Entry price

Lowest public entry point

Free trial and on demand starter on the GitHub Marketplace; paid team and enterprise pricing not fully public Free · Team $30/dev/mo billed annually ($40 monthly)

Pricing confidence

How public the numbers are

Public, partial Public, exact

Billing

Primary billing axis

per seat team and enterprise tiers, with a free trial and on demand starter option Per-developer/month subscription tiers with a monthly reviewed-line allowance (added/deleted diff lines cubic reads); free tier capped by monthly review count; enterprise custom

Variable cost

Workload / overage exposure

Medium variable cost Medium variable cost

Free tier / trial

Try before you buy

Free tierTrial
Free tierTrial

Buying motion

Self-serve vs sales call

Mixed Self-serve

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.