Agentic Index
Baz vs Greptile (2026)
Both review pull requests with real codebase understanding and they differ on scope, at 12 and 11.5 of 14. That verdict is the Agentic Index coverage score, graded from each vendor's own published materials.
Greptile builds a graph of your repository to review with full codebase context, catching cross file bugs that diff only tools miss, free for qualified open source then thirty dollars per developer monthly. Baz spans the lifecycle with a benchmark leading review agent, a specification validator against Figma and Jira, and a pre code gateway that catches risky plans. Greptile does one thing with unusual depth; Baz does more and asks you to adopt more.
On the Agentic Index coding agent ranking, Baz clears the bar and Greptile does not. Baz documents all five merge loop capabilities in full; Greptile does not document observability and auditability in full. 4 of the 63 vendors in the lane clear it. See the coding agent ranking
This comparison is published by Agentic Index, an independent agentic AI vendor research platform. Baz and Greptile are each graded against the same 14 capability Agentic Index taxonomy, from the vendor's own public materials under the Agentic Index verification standard, alongside 969 researched vendors. No vendor pays for placement and no vendor has reviewed this page. How this evidence is graded
Choose Baz if
- Documented coverage is materially broader and pre code review is capability you want.
- Specification validation against design and ticket is a gap in your process today.
- Governing agent written code across the lifecycle is the problem, not just reviewing it.
Choose Greptile if
- Cross file bugs are what your reviews miss, and repository graphs are how you catch them.
- Published per developer pricing at thirty dollars makes the cost obvious.
- You want a review tool that slots in, not a platform that changes your process.
| At a glance | Baz | Greptile |
|---|---|---|
| Category | Coding agent | Coding agent |
| Entry price | Free trial and on demand starter on the GitHub Marketplace; paid team and enterprise pricing not fully public | Free for qualified open source · Pro $30/developer/mo (50 reviews included, then $1/review) · Enterprise custom (self hosting, SSO/SAML, air gapped) · 14 day free trial |
| Free / trial | Free trial and an on demand starter plan on the GitHub Marketplace | Free Starter tier for individual developers, launched 29 June 2026: 1 active developer, 50 credits per month, unlimited repositories, no team creation. Open source projects also qualify for free use. 14-day free trial on paid plans. |
| Pricing confidence | public partial | public exact |
| Feature | B Baz |
G Greptile |
|---|---|---|
| Action & orchestration | ||
|
Integrations & Tool Calling Ability to connect agents to real systems through native integrations, OAuth-authenticated actions, custom tools, APIs, webhooks, or MCP-compatible tools. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Breadth across classes is comfortably met with three source platforms, six ticketing systems, design, and two knowledge bases, all dated in the changelog. The detail worth carrying to comparison pages is the Works alongside positioning: Baz publishes dedicated solution pages for Claude Code, OpenAI Codex, Cursor and Devin, which is the super-harness framing, integrating with coding agents rather than competing with them, the same shape as sonar in this cohort. |
Full / Explicit
Stands at F, re-based off the vendor's own changelog after the June basis cited aicodereview.cc. Breadth across classes is comfortably met, and the notable addition since the record was built is Fix with your Agent, which routes findings outward into five named coding agents through a local bridge CLI. That is an unusual integration direction for a reviewer and worth recording: Greptile positions as the reviewer that hands work to whichever agent the customer already runs. |
|
Workflow Orchestration Ability to sequence, branch, retry, route, and combine deterministic workflow nodes with autonomous agent steps. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Two things lift this beyond a standard multi-agent claim: the Planner loop runs BEFORE code exists, so orchestration covers plan generation, review and approval as well as implementation, and the Fixer agent closes its own loop by building an environment, validating and only committing when validation succeeds. Named infrastructure, the Agent Harness and Context Broker, is documented rather than implied, and Sessions trace the coordination. |
Full / Explicit
Upgraded from P. The June basis described parallel agents and multi-hop passes but predates v5, shipped 5 August 2026, which makes the architecture explicit: a swarm of narrowly scoped agents each exploring a single hypothesis, run in parallel and aggregated into one review. Graded F on the cosine precedent, where Swarm mode spawning specialised child agents earned F, and distinguished from blink-new, held at P because its parallel agents had no documented coordination. Here the coordination is evidenced by measured aggregate outcomes: median review time halved and comment-addressed rate rose from 52 to 66 percent, which only makes sense if the swarm's output is filtered and merged rather than concatenated. |
|
Triggers & Channel Coverage How agents wake up and where they work: schedules, webhooks, message events, CRM events, inbox events, chat, email, voice, and collaboration tools. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.55 to 0.85. All three trigger classes are documented and dated: events through pull request webhooks, schedules through Skills Maintainer scans every three days, weekly or fortnightly, and on-demand through comment commands, the CLI and a slash command invoked from inside another coding agent. That last one is the distinguishing detail, since it means a developer starts a Baz planning session without leaving Claude Code or Cursor. |
Partial
Stands at P, re-based off the changelog. Surface coverage grew materially since the record was built, with a CLI in June and CLI onboarding in July adding a terminal path that did not previously exist, alongside PR events, mention and re-trigger invocation, MCP from four IDEs and Slack delivery. Held below F because every path still anchors to the pull request or the local development loop: Slack is a delivery destination rather than an invocation channel, and no scheduled or cron-driven review is documented anywhere. Same line that holds cubic and graphite at P while warp, goose and ellipsis reach F on chat, ticketing or event-driven invocation. |
| Knowledge & context | ||
|
Knowledge Grounding & RAG Ability to ground agent behavior in company data through document ingestion, retrieval, external knowledge APIs, semantic search, or RAG layers. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Grounding here is unusually broad because it spans three kinds at once: code structure through embeddings and graphs, organisational knowledge through Notion and Fibery, and product intent through six ticketing systems plus Figma. The Voyage AI code embeddings detail is a rare instance of a vendor naming its retrieval model. Cross-repository context with a visible cross repo tag on findings is the part most comparable to potpie-ai and greptile at F. |
Full / Explicit
Stands at F and is now among the best-evidenced grounding cells in the lane, alongside cubic and cosine. Two mechanisms are new to the record and both extend grounding past the repository boundary: Repo Clusters reading up to seven related repositories per review, and the Partner Program supplying maintained context for third-party APIs. The files.json mechanism is worth noting as a design choice, since it points the reviewer at existing schemas and architecture docs rather than requiring a separate knowledge base. |
|
Memory & State Persistence Ability to persist context across a run, conversation, workflow, user, team, or longer-term memory layer. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Two independent named memory systems is more than any other vendor in this lane documents, and both are dated and mechanically described rather than asserted: Reviewer Memory accumulates from repeated feedback and rewrites the reviewer's own system prompt with version history and rollback, and Module Memory persists implementation detail per module. The prompt-versioning-and-rollback detail is the part that makes this unambiguously F, since it means memory is a managed, inspectable artefact rather than a black box. |
Full / Explicit
Upgraded from N. The June basis reasoned that learning from comments and reactions was feedback-loop adaptation belonging under Know, sourced to a dev.to post; Memory and Learning is in fact a named system with its own documentation page, its own dashboard section, five documented learning signals and an inference step that proposes new rules from observed behaviour. Graded F on the cosine precedent, where a named Memory feature persisting conventions across sessions earned F. The distinction from graphite, held at N in this same batch, is exactly that the vendor documents persistence as a capability with a mechanism rather than as a training characteristic. |
| Control & trust | ||
|
Human Oversight & Guardrails Approval steps, consent checkpoints, escalation rules, structured guardrails, policy constraints, and pause/resume controls. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Plan review approvals shipped 1 July 2026 and change the picture from the July basis, which credited only suggest-without-blocking and chat: a plan now goes through a dedicated approval workflow with versioning, pinned comments and approve or request-changes before implementation starts, which is a shipped gate ahead of code generation rather than review after it. Eighth distinct oversight architecture in this lane and the only one that gates the PLAN rather than the diff or the tool call. |
Full / Explicit
Upgraded from P. The June basis described advisory comments plus configurable rules, which was accurate but predates auto-approve, shipped 26 June 2026. Graded F on the cubic precedent: a documented autonomy boundary with a customer-set risk ceiling and a published never-approve list is a shipped guardrail mechanism, which is what this axis grades. Beta is not a bar to F under the early access treatment, since it is open to all users rather than gated by application, unlike warp Factories. The hard exclusion list is the strongest part: auth, secrets, billing, database migrations, infrastructure, CI and public APIs are never auto-approved regardless of configuration, which is a guardrail the customer cannot switch off. |
|
Security, Identity & Governance RBAC, SSO, auditability, encryption, least-privilege tool access, compliance posture, and data handling policy. |
Full / Explicit
Stands at F but the basis is materially rewritten, because the July basis graded partly from FOUNDER PEDIGREE, founded by cloud application security veterans, which is not evidence of anything the product or company does and should never have carried a grade. Re-retrieval found the real evidence the July build missed: SOC 2 certification stated in the site footer with a published trust centre at trust.baz.co and a security, privacy and compliance documentation page, plus role-based access control shipped September 2025. The conjunction is now met properly on attestation plus named controls. Confidence raised from a non-canonical 0.55 to 0.85 accordingly. |
Full / Explicit
Stands at F. The axis conjunction is met on both halves independently, and the self-hosted air-gapped path with customer-supplied models is the part that matters most for the regulated buyers this vendor targets, since it removes the trust question rather than attesting to it. Recorded honestly: the SOC 2 Type II claim and the defence, healthcare and financial services customer base come from vendor marketing pages rather than a trust portal, and no named audit firm or published report was retrieved. |
|
Observability & Auditability Traces, logs, execution histories, metrics, audit events, and debugging detail for production agent behavior. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.55 to 0.85. Sessions are a retained, structured execution record covering invocation, tool calls, stages, outcome and cost, which is the F bar on this axis as applied to ellipsis and poolside, and they extend across every agent rather than one. The product-level timeline framing rather than raw logs is worth recording: the vendor is deliberately building auditability as a reviewable artefact. The July basis graded this F on the marketing phrase observable and explainable, not a black box; the grade was right and the basis was not. |
Partial
Stands at P, re-based off the changelog after the June basis cited greptile.com/what-is-ai-code-review, a marketing explainer. The analytics dashboard shipped 15 April 2026 and is a real reporting surface with export, and the review footer's counter and last-reviewed-commit link add per-PR traceability. Held below F on the lane-wide reading applied to graphite, cubic and sourcegraph: these measure review outcomes and team throughput, not a retained per-action record of what the agent did and why. Contrast ellipsis, which earned F because every step, tool call and message is retained and replayable. Confidence stays medium because the analytics docs page was not fetched directly. |
|
Deployment & Data Residency Deployment modes and options, including SaaS, dedicated cloud, VPC, on-prem, hybrid, local runtime, and self-hosting. |
Partial
Upgraded from N and confidence normalised from a non-canonical 0.45. The July basis said no on-premise or self-hosted option is documented, which was graded from absence and is now contradicted by two dated first-party entries the July builder did not reach: Private Mode shipped 25 December 2025, and private application review running inside the customer's own Kubernetes cluster shipped 29 July 2026. Held at P rather than F under the early access convention, since private application review is explicitly a limited technical preview, and because Private Mode governs analysis and retention rather than giving the customer a full self-hosted deployment. Would move to F when private application review reaches general availability. |
Full / Explicit
Stands at F and is the strongest Dep cell reviewed in this lane so far. Unlike warp and cosine, where air-gapped deployment is described on a marketing page and quoted through sales, Greptile publishes the actual deployment mechanics: named services, sizing thresholds, a public repository, a Terraform path and a documented migration route between deployment methods. Both halves of the axis are met independently, deployment surface and data location control, with bring-your-own-LLM closing the inference path. |
| Solution readiness | ||
|
Prebuilt Agents, Templates & Packs Ready-made workflows, packaged employees, templates, blueprints, industry solutions, and role-specific agents that reduce time-to-value. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.6 to 0.85. Eleven named prebuilt agents and reviewers is the largest shipped agent suite in this lane, and the Awesome Reviewers corpus is an unusual second layer: a free public library of 5,132 review instructions across twelve domains that a customer fetches as a markdown bundle straight into a skill folder. That is vendor-supplied adoptable assets in exactly the form this axis rewards, and it is published openly rather than gated behind the product. |
Partial
Upgraded from N. The June basis said custom rules are user-defined configuration and no vendor-supplied library exists, which was true then; the Partner Program shipped 22 June 2026 and is exactly a vendor-curated pack, supplying partner-maintained review rules for eight named third-party APIs, enabled by default. Held at P rather than F because these are context packs applied automatically rather than a browsable catalogue of installable agents or templates, and there is still no marketplace or gallery. AI rules import is customer-owned configuration and is recorded rather than credited. |
| Platform extensibility | ||
|
Model Flexibility & Routing Ability to work across multiple foundation models, route tasks to different models, or let buyers bring their own providers and keys. |
No / Not documented
Stands at N and confidence normalised from a non-canonical 0.5 to 0.65. The July basis said Bedrock foundation models with no customer choice, which re-retrieval supports but also refines: the changelog names GPT-5-Codex with Sonnet 4.5 reflections as the reviewer stack, so the vendor mixes providers, but that is the vendor's routing decision, not the customer's. Nothing documents model selection, bring-your-own-key or a provider setting anywhere. Recorded honestly: like sonar in this cohort, a low Model grade here reflects a deliberate product design where the vendor owns model choice as part of guaranteeing review quality, and the vendor does disclose which models it uses, which is better than the provider-undisclosed case the conventions treat as N. |
Full / Explicit
Upgraded from P, and this is a retrieval failure rather than product movement: Configurable Models has been documented since 26 September 2025, nine months before the record was built, and the June basis cited sacra.com rather than the vendor. Three independent forms of the axis are present, which is as strong as this cell gets: explicit customer selection, documented routing, and bring your own model on self-host. Model Inversion is a genuinely unusual routing rule and worth a comparison-page note, since it routes away from the authoring model on the vendor's own research that models catch more bugs in code written by a different model. |
|
APIs, SDKs & MCP Extensibility Composability layer: stable APIs, SDKs, MCP tool consumption/serving, custom tools, and integration into internal systems. |
Partial
DOWNGRADED from F on Mike's ruling of 2026-08-30, overturning my own August grade. Same error as zencoder: I held F while the basis itself recorded that no public REST API or SDK was retrieved, and flagged it as the weakest F on this axis in the lane rather than acting on that. Under the axis bar, full coverage means the platform is composable from outside through a documented API or SDK, with MCP as one form of evidence rather than the bar. The Baz surface is genuinely strong and that is why this is P rather than N: an MCP server exposing review tools to any IDE agent, a CLI with an interactive review loop, a slash command that starts Planner from inside another coding agent, and marketplace distribution. The Awesome Reviewers raw endpoints are a real documented API, but they serve a public corpus of review instructions rather than the Baz platform, which is the same distinction that correctly kept appfactor at P: MCP Bridge exposes the customer's systems, not AppFactor. Grade would move to F on a documented platform API or SDK. baz.ai/docs was not fetched on either pass and remains the cheap check at lane close. |
Full / Explicit
Upgraded from P, which said no public SDK is documented, sourced to sacra.com. The surface is broader than an SDK gap implies: four named REST endpoints, a hosted MCP server with a documented bearer-token setup across four IDEs, an npm CLI with machine-readable and agent-oriented output modes, a plugin in Anthropic's official marketplace, webhooks and Zapier. The MCP direction test is satisfied outward, since other assistants call Greptile to query rules and trigger reviews. |
|
Testing, Debugging & Optimization Testing, debugging, scoring, retries, fallbacks, quality gates, and optimization loops for improving agent workflows before and after deployment. |
Full / Explicit
Stands at F and confidence normalised from a non-canonical 0.55 to 0.85, and this is one of the very few genuine F grades on this axis in the lane. The Eval rule is that the axis measures what the CUSTOMER can test, not what the vendor tests internally, and Prompt Playground is exactly that: a shipped agent test lab where a customer edits a reviewer's prompt and runs it against real change requests before rollout. Evaluations closes the loop by classifying every comment as accepted, rejected or unaddressed and feeding that back into reviewer behaviour. Contrast the nine vendors held at P in this lane for agents that merely verify their own output. |
Full / Explicit
Stands at F on the qodo and cubic precedent, where the axis and the product coincide, and the case is stronger than it was in June because TREX and the security agent both shipped after the record was built. TREX is the unusual part and worth pairing on comparison pages: writing and executing targeted tests against the repository's real stack rather than a mock environment, and attaching execution evidence to the comment, is closer to a customer-facing test harness than anything else reviewed in this lane. Recorded honestly: TREX is in public beta and the security agent's benchmark claims are vendor-reported. |
| Specialist automation | ||
|
Browser & Computer Use Browser, desktop, or remote/local computer control for workflows that cannot be handled through stable APIs alone. |
Full / Explicit
Upgraded from P and confidence raised from a non-canonical 0.45 to 0.85. The July basis said computer use is limited to validation, which undersold it: this is the most thoroughly documented browser-driven capability found anywhere in this lane, built up across four dated releases from November 2025 to July 2026, and it is the product's differentiator rather than an accessory. The axis test is met exactly, since a rendered user interface has no programmatic interface and the agent operates it as a human would, clicking through flows and reading state off the screen. Contrast the thirteen downgrades in the June cohort, every one of which credited a terminal or sandbox; here the sandbox is where the browser runs, and the browser is the point. |
No / Not documented
DOWNGRADED from P, the eighth correction of this identical axis error in the 30 June cohort. The June basis credited running TREX in a sandbox, navigating the repository graph and applying click-to-accept fixes as computer use; sandboxes, code graphs and the GitHub API are all programmatic interfaces, which is what the axis excludes. One genuine ambiguity is recorded rather than resolved: TREX attaches screenshots and videos as failure evidence, which implies browser-driven end-to-end tests, but running the customer's own test framework is executing their harness rather than operating software that lacks a programmatic interface, and no browser tool, computer-use capability or GUI automation is named on the TREX page or anywhere in the changelog. Would move on a documented browser tool. |
Pricing snapshot
Sourced from the Index pricing dataset · open each vendor's profile for full detail.
| Pricing | ||
|---|---|---|
|
Entry price Lowest public entry point |
Free trial and on demand starter on the GitHub Marketplace; paid team and enterprise pricing not fully public | Free for qualified open source · Pro $30/developer/mo (50 reviews included, then $1/review) · Enterprise custom (self hosting, SSO/SAML, air gapped) · 14 day free trial |
|
Pricing confidence How public the numbers are |
Public, partial | Public, exact |
|
Billing Primary billing axis |
per seat team and enterprise tiers, with a free trial and on demand starter option | Per developer per month base subscription of thirty dollars including fifty reviews, then one dollar per additional review. Free for qualified open source projects. Enterprise is a custom annual or multi year contract, including self hosted deployment. |
|
Variable cost Workload / overage exposure |
Medium variable cost | High variable cost |
|
Free tier / trial Try before you buy |
Free tierTrial
|
Free tierTrial
|
|
Buying motion Self-serve vs sales call |
Mixed | Self-serve |
More comparisons with Baz or Greptile
Other matchups in coding agents
Not the pairing you were after? These compare a different set of coding agents on the same 14 capabilities.