← All issues

The Agentic Index Brief

September 19 to September 26, 2026 · Published September 27, 2026

The week in one line

The agent now asks permission only when it matters, which is more self restraint than most of us manage. GitHub Copilot began waving through low risk tool calls, Factorial made every tool declare whether it reads, writes or destroys, and Factory nudged its default toward autonomy. Windmill, meanwhile, now bills a service account as half a seat.

This issue covers September 19 to 26. The log recorded 84 entries across 84 vendors, 67 of them Verified against a primary source and 17 Partially Verified. MCP and tool calling tied with plain agent capability for the most entries, at 20 each. Human approval and guardrails came next with 13, and that is where the week's story starts.

The permission prompt learned triage

On August 30 this Brief said approval had stopped meaning stop. On September 19 the confirmation click came off at two vendors. This week the click came back, but only for the calls that deserve it.

GitHub Copilot for JetBrains introduced assisted approvals in public preview. Lower risk tool calls are approved automatically, and higher risk actions still wait for a person. The same release added a rewind that undoes an earlier request along with the file changes it caused.

Factorial made the risk a declaration rather than a guess. Every tool in Factorial Actions now states whether its effect is a read, a write or something destructive. Reads run without asking. Writes ask first. Destructive actions ask with a warning that the step cannot be undone.

The detail worth borrowing is the fallback. A tool that declares nothing is treated as destructive. Most software assumes the best about code it has never met, so this is a refreshing change of manners.

Other vendors drew the line with a number. Microsoft 365 Copilot can now delete, move and file email on request. It asks first only when an action touches more than five messages or changes an inbox rule. Five emails is, apparently, the point where tidying up becomes a decision.

Dust pauses an agent once a single message spends 600 credits and asks whether to carry on. AutoGPT added approvals that trigger at a spending threshold. cubic will now approve a pull request on its own while findings at or below a chosen priority are still open.

Then the defaults moved. Factory changed its app's default session autonomy from Off to Medium for anyone who never picked a level. Medium lets the agent edit files, install dependencies, build and commit locally without asking each time.

Claude Code 2.1.283 now starts interactive sessions in auto mode when no permission mode is configured, on third party providers or with telemetry turned off. In both products an explicit setting still wins. That matters less than it sounds, because most people never open the setting.

The ways an agent decides to stop and ask got more expressive too. Browserbase agents now accept a pause condition written in plain English. They stop inside the same browser session and pick up again when a person supplies what was missing, such as a login code.

LangGraph lets an interrupt describe the exact shape of the answer it needs, and checks the reply before resuming. xpander.ai's gated mode now holds web searches and page fetches for approval. Reading the internet, it turns out, is also an action.

Our read: the approval prompt is being rebuilt as a risk classifier, and the category is settling on three tiers. Do it, ask first, or ask with a warning. What now separates products is who decides which tier a call belongs in. At GitHub it is the vendor's judgment. At Factorial it is the developer's declaration. At Microsoft and Dust it is a published number. Meanwhile the default keeps creeping toward autonomy, and a default is the one product decision that reaches every user who never opened settings. Expect autonomy level to become something every agent product has to name and explain, the way privacy settings once did.

The session became a job

For most of this year an agent's work ended when the conversation did. This week a cluster of releases gave it a goal, a clock and somewhere to put unfinished work.

Kilo Code added a goal tool that lets its agent set a session goal and keep working after the current turn ends. The same release lets the agent schedule its own recurring tasks, which expire after seven days. It may be the first colleague ever to book a standing meeting with itself and then actually show up.

Exa launched Agent Ultra for exhaustive research, with runs lasting up to three hours. The controls that shipped with it are the telling part. You can cap the spend, cap the duration, or stop early and keep whatever it found.

Kiro gave its crews a durable queue, so accepted work survives a gateway restart. It also added AWS Fargate as a place to run remote crews, with a six hour task lifetime. DBOS added workflow rewind, which restarts a finished workflow from a chosen step rather than from the top.

The runtime layer filled in underneath. Temporal began running serverless workers on Amazon Bedrock AgentCore in prerelease. Prime Intellect made its sandboxes generally available, each one its own Linux microVM, with room for 1,024 at once on a new account. LiteLLM previewed LiteAgents, one interface across five agent frameworks with optional recovery from recorded checkpoints.

The job also started moving between machines. CodeRabbit can now hand a local session to a cloud agent, along with its plan and transcript. Codex CLI 0.156.0 added worktree sessions and a usage dashboard that reports credits and tokens.

Cursor pushed the job past the merge. Its new Rollouts feature follows a pull request into production and checks deployment health against logs, metrics and traces. It can propose a revert or hand a regression to a cloud agent. It does not merge or roll back by itself, and that is the line worth noticing.

Our read: the unit of agent work is moving from the turn to the job. A turn needs a prompt. A job needs a goal, a schedule, a queue that survives a restart, a budget and a way back to an earlier step, and each of those shipped this week at a different vendor. That changes which numbers belong on a spec sheet. Run length, concurrency and recovery are becoming the figures vendors lead with, the way context window size used to be.

The trace became the training set

Salesforce opened a beta of Agent Optimizer, which puts an agent's configuration, its test conversations and its production sessions in one place. It finds patterns in the failures and proposes fixes. It then builds regression tests and stages each change behind review points the customer chooses.

Ada shipped the same idea as a single MCP tool. Give it a set of changes and a date range, and it compares the new version of an AI agent against the live one on containment, resolution and satisfaction. That is an A/B readout for an agent's behavior, delivered as a tool call.

LangSmith went a step further with smithtune, now in public beta. It turns recorded agent runs into a fine tuning dataset and sends the training job to Fireworks or Baseten. Then it compares the tuned model with the original through LangSmith evaluations and can deploy whichever wins.

Every mistake the agent made in production becomes a lesson plan. That is a more constructive relationship with past mistakes than most people manage.

Evaluation also reached the channels that were hardest to test. Confident AI launched voice evaluations with simulated callers who interrupt, talk over background noise and take turns badly. Seven audio metrics score the result. Intercom let teams build reports on Fin's ratings and count how often the AI agent hands a conversation to a human.

FLOWX.AI added a browser automation node that checks the outcome against the target site's own record, rather than the agent's account of what it did. Trusting the self report is how most bad status meetings begin.

The plumbing kept pace. Braintrust released Nitro, a query engine it says ran full text trace searches more than twice as fast. W&B Weave now links each call to the agent spans around it, in both directions. Boomi's Agent Control Tower can take in agents from Salesforce, Microsoft Copilot and Snowflake Cortex, including usage figures those vendors do not return through the standard sync.

Our read: the evaluation loop is closing, and it is closing inside the product rather than next to it. All summer the pitch was that a vendor could show you what the agent did. The pitch now being built is that it can show whether a change made the agent better, and that is much harder to fake in a demo. The census at the end of this issue shows how far most of the market still has to go.

The service account got half a seat

The August 30 issue asked a pricing question nobody had answered. When an agent does the work, whose seat is it?

Windmill now has an answer. In self hosted Enterprise deployments, each enabled service account counts as half an operator seat, whatever role it holds. Once the license cap is reached, new service accounts are refused. Half a seat is a remarkably precise answer to a philosophical question.

The more revealing packaging move came from TrueFoundry. On its managed control planes, nonpaying tenants now run into plan limits. Agents, MCP tool approval, custom roles and single sign on all require a higher plan. Existing setups keep running, and paying and self hosted customers are unaffected.

In the same week Voiceflow made live handoff to a human agent available on every plan. Read those two together. Reaching a person became a free feature. Controlling the agent became the reason to upgrade.

Time turned into a price lever as well. OpenRouter launched a Batch API across more than 70 models, where accepting up to 24 hours of waiting typically halves the token price. Vapi added a setting that lets an assistant request faster processing from OpenAI. Waiting now earns a discount and hurrying has a dial, which is roughly what airlines worked out decades ago.

Our read: agent packaging is sorting features into three bins. The handoff to a human is going free. The controls over the agent are becoming the upgrade. Everything else is metered, by credit, by token, by hours of patience and now by service account. The seat is not disappearing from agent pricing. It is being reassigned to things that are not people.

Market notes

The coding agent picked up an administrator. Databricks Mosaic AI added central configuration for coding agents to its Unity Gateway, covering approved models, MCP servers, skills and routing. A new command line launcher applies those settings and handles sign in when a developer starts Claude Code or Codex.

Claude Code itself gained a managed setting that denies specific models outright. OpenHands Enterprise now syncs automations with a Git repository, so changes go through pull request review before they deploy. Baz launched shared review of the plans coding agents write before they write any code. When a data platform ships the launcher for somebody else's coding agent, the coding agent has become infrastructure.

MCP kept separating reading from writing. Atlassian began rolling out capacity planning tools in Rovo MCP as two toolsets, one to read and one to write, each granted by an administrator. Arcade's new Snowflake toolkit queries under the requesting user's own identity and blocks writes, schema changes and role switches.

Smallest AI added short lived tokens for browser and mobile clients, which cannot clone a voice or mint more tokens. And the protocol acquired a past. Mastra's MCP package 2.0 supports only the July 28, 2026 revision of the specification and drops the older transport entirely. Nobody bothers to deprecate the version of a protocol that lost.

Several vendors decided backward compatibility had become a cost rather than a courtesy. Base44 set October 15 as the retirement date for account API keys, moving integrations to personal access tokens. In the interest of disclosure, this index is built on Base44, and it is named here for the same reason it is named on our ranking pages.

Cognition retired its standalone Devin Review command, and copies already installed stop working. Agent Zero renamed a remote execution setting and kept no alias for the old name. AutoGPT now refuses the default encryption key it once published, so installations still using it must migrate their stored credentials. None of these is dramatic alone. Together they say the category believes it finally has enough real users to annoy.

Agents took on more of the money, with the human kept exactly where the money leaves. Ramp launched Accounts Receivable agents that draft invoices from contracts, prepare collection messages and match incoming payments. Finance still reviews each draft and sends every collection email.

Chatbase agents can now create Shopify orders from WhatsApp, Instagram and Messenger conversations, though a person approves any exchange before it reaches the store. Delight.ai keeps a caller on the line while its agent reaches and briefs the human taking over. Where the person stays in these products is a fair map of where money or reputation leaves the building.

In the regulated verticals, coverage was the headline. Clio added Canadian case law and legislation to Vincent, including more than 567,000 judicial decisions. careCycle extended its Medicare plan data to every US county and cut benefit answers from 6.5 seconds to two. Corti now lets a clinician's edits override the transcript when it generates a note, and leaves out the facts they discard.

Elsewhere, the workhorse releases. Palo Alto Networks launched a continuous Unit 42 service that keeps testing a changing environment and routes each security task to a frontier or open weight model. Firecrawl launched Alexandria, one interface for agents to query official data providers alongside the live web.

Google Antigravity's SDK can now run models locally, and Cline shipped its desktop app for Linux. Cartesia released voices that keep one vocal identity across up to 25 languages. And Replit's agent now draws interactive charts right in the chat from a dataset you hand it.

What the week says about the category

Eighty four entries, and one idea under most of them. The agent is being given more rope, and the rope is being measured.

Defaults moved toward autonomy in the same week approvals learned to tell a delete from a read. Sessions turned into jobs in the same week evaluation learned to compare a change against the live version. The service account got a seat in the same week the controls became the upgrade.

Almost nothing shipped more autonomy this week without shipping a way to bound it, and the bounds are finally specific enough to compare. Five messages. Six hundred credits. Six hours. Read, write or destructive. A market that argues in numbers like those has stopped debating whether agents should act. It is negotiating how much.

Index Answer

Which AI agent platforms give you tracing and observability in production?

More than half of them, and that is the less useful half of the answer. Of the 548 platforms the Agentic Index grades across its four horizontal lanes, 317 document observability and auditability in full. That means tracing across agent runs, visibility into each tool call, and an audit log a compliance team can read.

The number underneath it is much smaller. Only 151 platforms document testing and evaluation in full, meaning a way to replay or simulate the agent and a stated way to tell whether a change made it better. 185 of the 317 platforms that trace in full do not document evaluation at all.

That gap is the finding. A trace tells you what the agent did. It cannot tell you whether the change you are about to ship will make it better or worse. Most of this market sells the first and leaves the second to you.

It is also where the bar breaks. Of the 185 platforms sitting exactly one capability short of the full reliability bar, 129 are short on testing alone. Only 81, or 14.8 percent of the field, document all three capabilities. Their agents can be tested before release, watched in production and stopped when they go wrong.

This week's releases are the missing half starting to arrive, from support, sales and developer tooling vendors at once. If that holds, the number that moves next is the second one in this answer, not the first.

The bar, the method and the 81 platforms that clear it are on the Agentic Index agent observability page.

The Agentic Index Brief is published weekly by Agentic Index, the verified directory of 946 agentic AI vendors. Compare platforms by capability at agenticindex.io/compare. Methodology at agenticindex.io/methodology.

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.