Agentic Index
Agentic AI Platform Change Log
Agentic Index tracks material changes across agentic AI platforms, including new agent capabilities, workflow orchestration updates, integrations, MCP/tool-calling changes, memory features, human approval controls, observability, deployment options, pricing, funding, and verification-status updates. This change log is the index that tracks changes to agentic AI platforms, updated weekly.
Public-source research. All entries are sourced from official vendor blogs, changelogs, documentation, or press materials. Support levels marked Needs Review or Publicly Claimed should be independently verified before relying on them for procurement decisions.
561 entries
MCP / tool calling / API
ServiceNow published an overview of Action Fabric, an interoperability framework connecting AI agents across the enterprise. It features an MCP Server Console for external clients to invoke ServiceNow tools, an MCP Client for ServiceNow agents to call external systems, and Agent2Agent protocol support for peer collaboration.
This prevents the creation of isolated AI agent silos by enabling cross-platform orchestration. Buyers can securely expose ServiceNow workflows to external assistants or empower Now Assist with specialized external capabilities while maintaining strict governance and access controls.
Pricing / packaging
OpenAI reduced the pricing for its GPT-5.6 Terra and Luna models and announced that GPT-5.4 and GPT-5.4 mini will be retired from Codex on August 31, 2026. Users must migrate their workspace defaults, saved model settings, and custom agents to the GPT-5.6 models before the cutoff.
IT administrators and developers must proactively update managed configurations to avoid service disruptions when the GPT-5.4 models are retired. The price reductions on the newer GPT-5.6 tier provide a cost efficiency incentive for the mandatory migration.
Agent capability
Hippocratic AI launched a new suite of Rapid Response Climate Agents designed to help health systems and payers perform proactive outreach to vulnerable patients during extreme weather events. The climate-responsive voice AI agents scale instantly to conduct clinical assessments, such as heat stroke checks, and route individuals to appropriate resources like cooling centers.
Healthcare organizations can now automate critical, high-volume outreach during sudden environmental crises, extending care team capacity when it is most strained. This allows payers and providers to rapidly check on high-risk members without needing to source emergency call center staffing.
Memory / state
Cobl introduced a major product update shifting the platform from single-document generation to comprehensive deal management. The release includes Deal Workspaces to group related sales documents, a Shared Drive acting as an automated knowledge base for agents, automatic extraction of reusable design systems from existing documents, and document analytics.
Sales teams can now manage complete deal lifecycles in a unified workspace where context is automatically shared across assets like decks and proposals. The automated knowledge retrieval and design system creation eliminate the need for reps to repeatedly input company background information or manually apply brand formatting.
Security / enterprise
Charm released version 0.88.0 of the Crush terminal coding agent, introducing a scriptable configuration system via crush.sh and crushrc. The update adds explicit tool permissions, allowing users to apply "deny" rules to completely hide specific tools from the AI agent. The release also includes verb-first setup commands for managing MCP servers and LSPs, as well as auto-retry logic for AWS SSO logins.
The addition of explicit tool deny rules provides teams with stronger security boundaries, ensuring the AI agent cannot access restricted features or sensitive system operations. Furthermore, the new scriptable configuration and streamlined MCP setup make it easier to enforce consistent development environments across larger engineering teams.
Funding / partnership
Tabnine was acquired by Tricentis, an enterprise software testing and quality engineering company. The acquisition integrates Tabnine's Enterprise Context Engine into the Tricentis Agentic Quality Engineering Platform.
Buyers evaluating Tabnine will see its capabilities merged with the Tricentis ecosystem, connecting AI-driven development directly to quality assurance. This combined approach provides deep architectural context to testing agents, helping organizations reduce hallucinations and deploy reliable AI workflows.
Integrations
Snyk integrated its Studio product directly into Snowflake's Cortex Code environment. The integration acts as a continuous security guardrail, scanning AI-generated code snippets, containers, and third-party dependencies for vulnerabilities in real time as developers build data applications.
This allows data engineers to safely leverage AI coding agents within Snowflake without relying on disconnected post-deployment security tools. By catching vulnerable or poisoned dependencies early in the development loop, enterprise teams reduce the risk of shadow AI and lower the costs of remediation.
Browser/computer use
OpenAI shipped version 26.727 of the ChatGPT and Codex desktop app, adding an improved built-in browser with address-bar search and enhanced Chrome extension integration. The update also introduces multi-repository code review and an expanded viewer for generated images.
Developers can blend web research, open Chrome tabs, and local cross-repository code diffs within a single task without switching contexts. Enterprise security teams should evaluate this expanded integration surface, as browsing history and multi-repo code are now merged into the agent's memory.
MCP / tool calling / API
Ironclad has released integrations with Anthropic, OpenAI, and Slack using the Model Context Protocol (MCP). The update routes user questions through an MCP server to the Ironclad Assistant, which translates queries into searches across the contract repository. This allows the AI tools to surface real-time contract data, workflow statuses, and key dates directly in the chat interface.
Users can retrieve specific contract details and status updates without leaving their preferred AI chat environments or Slack workspaces. This reduces friction for non-legal business teams who need quick answers about renewals and deal approvals. It effectively extends the contract system of record into everyday conversational tools.
Deployment / data residency
HireVue achieved certification under the US Department of Commerce Data Privacy Framework. This update provides a regulatory mechanism for transferring personal candidate data from Europe, the United Kingdom, and Switzerland to the United States.
Global enterprises evaluating this platform can now rely on the framework to ensure cross-border data transfers comply with European data protection laws. This milestone simplifies compliance reviews and reduces legal friction for multinational deployments.
Workflow orchestration
Element451 introduced a Repeatable toggle for Bolt Agent Jobs that allows contacts to be re-enrolled once their prior enrollment reaches a terminal state. This enables the agent to continuously respond every time a contact repeats a recurring action, such as submitting a new inquiry or opening a new support request.
Institutions using automated outreach can now configure agents to continuously handle repetitive behaviors from the same student. This prevents the system from permanently skipping a contact after their first interaction and reduces the need for manual duplicate job configurations.
Agent capability
Dropzone AI announced the general availability of AI Threat Hunter, a proactive threat hunting agent. The tool runs structured hunt packs across environments to uncover threats, emerging risks, and security coverage gaps missed by traditional alerts, leveraging more than 270 prebuilt hunt packs mapped to MITRE ATT&CK.
Security teams can operationalize proactive threat hunting without needing dedicated specialists or extensive manual queries. By automating federated searches and evidence correlation, the agent reduces hunt times from days to hours, allowing organizations to routinely surface stealthy intrusions and improve their overall security posture.
Agent capability
DepthFirst introduced dfs-large1, a proprietary security model built on GLM 5.2 and post-trained using reinforcement learning within its security agent harness. The model is now available in preview for vulnerability discovery and validation across large enterprise repositories.
Security teams can deploy a model specifically tuned for complex codebases instead of relying on general-purpose LLMs for vulnerability scanning. The reinforcement learning process is designed to penalize excessive tool use, which helps control compute costs from open-ended agent investigations.
Security / enterprise
CrowdStrike extended its Falcon AI Detection and Response (AIDR) module to provide active protection for AI agents built in Microsoft Copilot Studio and Anthropic's Claude Code. The update safeguards against threats originating from third-party AI tool usage.
Security teams can now safely enable internal employees to build custom agents and use conversational AI coding assistants without sacrificing visibility or risking data leaks.
Integrations
Celonis expanded authentication options for Delta Sharing connections by adding OAuth (Client Credentials) support. Customers can now integrate with identity providers like Microsoft Entra ID or Okta to manage access.
Buyers can use their existing identity providers to manage and dynamically rotate access credentials, replacing static bearer tokens with short-lived tokens for improved connection security.
Integrations
Wayfound introduced a direct integration with AI gateways like LiteLLM to add a Guardian Agent supervision layer to existing AI agents. By ingesting gateway telemetry, the platform automatically discovers agents via virtual keys and evaluates their activity against defined policies.
Governance teams and business leaders can now supervise AI agent behavior without requiring engineering teams to change client code or deploy new infrastructure. This provides an immediate inventory of running agents and ensures they align with expected business outcomes.
Agent capability
Sierra introduced Agency, a new infrastructure primitive that provides secure and scalable sandboxes for its AI agents. The Agency environment allows agents to execute long-running, complex tasks with dedicated computing resources, maintaining state and data isolation without interruption.
Enterprise buyers require strict data privacy and reliability when deploying autonomous systems. This backend architecture ensures customer data remains secure during execution while enabling agents to handle sophisticated workflows that take hours or days to complete without losing progress.
Agent capability
Resilinc rolled out its July 2026 platform update, which introduces enhancements to the DDRT FLC Agent to improve visibility into affected supplier locations and accelerate compliance investigations. The release also adds drag-and-drop file uploads for Dataflow Integration, part-level User Defined Attributes, and new REST API fields for supplier hierarchy data.
Supply chain risk teams can conduct faster compliance and disruption investigations using the upgraded FLC Agent. Furthermore, the new data and API features reduce IT bottlenecks by allowing users to directly upload and segment supply chain data without schema changes.
Workflow orchestration
Qualtrics added the ability to select AI Topics when building conditions within a workflow. This feature allows users to create automated workflows that are triggered based on the presence of specific AI-generated topics detected in the text.
This provides greater flexibility and speed in responding to customer feedback. Buyers can orchestrate automated follow-up actions or alerts exactly when the system identifies targeted themes, reducing the need for manual triage.
MCP / tool calling / API
OpenAI released Codex CLI version 0.146.0, introducing Agent Plugins 1.0 for portable Model Context Protocol (MCP) servers, remote Code Mode over WebSocket, and custom-provider web search routing. The release also stabilizes proxy routing across authentication and plugin downloads.
Engineering teams can cleanly separate the agent's user interface from its execution environment using remote Code Mode. The stabilized proxy routing ensures enterprise compliance, allowing organizations to securely deploy Codex within strict network boundaries.
Observability / auditability
Monte Carlo launched Automatic Conversation Clustering for Databricks AI/BI Genie and Snowflake Cortex agents. The platform natively reads trace data from the data warehouse to automatically group unstructured user queries into named conversational topics without requiring manual tagging or code instrumentation.
Data teams can now measure aggregate user intent and monitor share-of-traffic across AI agents instead of manually sampling individual conversational threads. This top-down visibility enables organizations to identify usage patterns, prioritize capability improvements, and run targeted evaluations on the highest-volume agent queries.
Security / enterprise
Monte Carlo introduced OAuth 2.0 client-credentials authentication for machine-to-machine interactions across its API, SDK, and CLI tools. The update provides a credential management interface that supports personal and authorization group-scoped service clients with dual-secret rotation capabilities.
Enterprise platform engineers gain a secure and standardized method for programmatically interacting with the Monte Carlo infrastructure. The ability to perform zero-downtime secret rotation helps teams comply with strict corporate security policies and safely integrate data observability into automated CI/CD pipelines.
Agent capability
LangChain released Deep Agents v0.7, an update to its open-source agent harness that reduces base input tokens by approximately 65 percent. The release removes the default system prompt, trims built-in tool descriptions, and makes TodoListMiddleware opt-in. It also introduces new configurability for overriding built-in middleware and optimizes filesystem tools with paginated reads and streamed grep outputs.
Developers can achieve significant cost savings and faster execution times due to the massive reduction in base token usage on every agent turn. The removal of prompts and the addition of middleware configurability give engineering teams full control over their agents' context and behavior without framework interference.
Agent capability
Kai launched Auto Remediation, a new capability for its agentic AI platform that automates the resolution of confirmed security vulnerabilities. The system evaluates findings to determine if they can be remediated without human intervention and applies fixes autonomously. If human approval is required, the platform builds a complete remediation plan and stages the change for technical teams to execute.
Security teams can significantly reduce their response times and alleviate alert fatigue by offloading routine fixes to the autonomous agent. Providing staged changes for complex issues allows organizations to retain control over critical infrastructure while still accelerating the overall remediation process.
MCP / tool calling / API
The July 2026 end-of-month release introduces rate limiting for MCP tools, allowing administrators to cap per-minute requests at the tool and tenant levels. The update also adds a centralized connections manager for reusing knowledge source credentials, richer input widgets for multi-turn chat, external workflow event notifications, and a new AWS London deployment region.
Enterprise teams gain stronger governance over API integrations to prevent external system overload and unexpected costs. Additionally, the unified knowledge connection manager reduces administrative overhead, while the new London region enables UK customers to meet strict data residency requirements.
Security / enterprise
Goose version 1.45.0 introduces several new capabilities, including support for the latest Gemini models and a configurable GOOSE_DOCS_ROOT environment variable for air-gapped documentation access. The release also adds structured summary output, the ability to disable built-in skills, and resolves a security vulnerability by patching the Nostr dependency.
Enterprise teams operating in secure environments can now host the agent's documentation fully offline using the new air-gapped mode. Furthermore, administrators gain tighter control by being able to disable built-in skills, and the security patch ensures compliance with the latest vulnerability disclosures.
Integrations
Cyware launched a new integration with Armis Centrix to deliver asset-centric threat contextualization. The partnership connects Armis asset visibility with the agentic AI capabilities of the Cyware Intelligence Suite to map global threat data directly onto an organization's device landscape.
This integration enables security teams to identify exactly which unmanaged or IoT assets are exposed to specific threat actors. By unifying threat intelligence with asset exposure, buyers can transition from reactive mitigation to automated defense workflows.
Observability / auditability
Cresta expanded its Cresta Insights platform with a new Real-Time Trends feature that detects conversational anomalies and spikes minute by minute. The release also upgrades the AI Analyst to autonomously execute multi-step research queries and enhances Topic Discovery to automatically generate multi-layered conversation taxonomies.
Contact center leaders can proactively catch emerging issues, such as outages or product faults, as they happen without relying on predefined keywords. The autonomous research tools reduce the manual labor required to categorize customer interactions and extract actionable insights.
Security / enterprise
Baz added the SAST-inside feature to its Advanced Security agent. This update runs specialized static analysis checks over modified code to nominate potential vulnerability hypotheses. The agent then investigates each signal against the broader repository context, including data flow and reachability, to determine if it is exploitable before creating a review finding.
Security and engineering teams can catch complex vulnerabilities during code review without increasing false positives. By combining the broad coverage of static analysis with the contextual reasoning of large language models, organizations reduce the noise of raw alerts while preventing insecure code from reaching production.
Security / enterprise
Automation Anywhere announced that its agentic AI platform is now HIPAA compliant, supporting the processing of protected health information (PHI) across administrative healthcare workflows. The compliance framework enforces strict access controls, data encryption, and comprehensive audit logging. It also includes the provision of business associate agreements (BAAs) for covered entities.
Healthcare organizations can now securely deploy these AI agents for sensitive multi-system tasks like claims processing and patient intake. This ensures compliance with regulatory data protection standards while reducing manual administrative overhead.
Memory / state
Torq introduced Torq SOC Brain, a new self-learning layer for its AI SOC Platform that trains on an organization's specific investigation history and analyst decisions. The release includes three core capabilities (Recall, Reflex, and Retrospect) that allow the system to reason from past precedent and continuously refine its judgment with every completed security investigation.
Security teams can rely on a private AI memory layer that continuously adapts to their unique operational practices and analyst corrections. This provides fewer repetitive reviews, more consistent threat classifications, and stronger trust in automated incident responses.
Workflow orchestration
Taskade Genesis applications can now accept file uploads and route them directly into automation flows. The release introduces a 'Convert File to Text' step with built-in OCR fallback for scanned PDFs, automatically storing the ingested files in the workspace media library for downstream processing.
Organizations building custom agent applications can natively ingest and parse user documents without wiring up third-party storage or text extraction tools. This simplifies the deployment of document-heavy AI workflows and speeds up document analysis tasks.
MCP / tool calling / API
Stacklok released ToolHive v0.41.0, introducing support for the stateless MCP 2026-07-28 specification revision. The update enables the platform to negotiate and bridge legacy session-based and new stateless clients, while adding an RFC 8693 OAuth 2.0 token exchange grant handler for agentic authorization.
Platform engineering teams can govern mixed generations of MCP clients and backends on a single enterprise gateway. The new token exchange functionality provides concrete identity controls, allowing organizations to securely delegate and audit permissions for autonomous agents acting on behalf of users.
Agent capability
Pydantic AI released version 2.20.0, introducing native support for the Claude Opus 5 and GPT-5.6 models. The update adds explicit prompt caching for GPT-5.6, reasoning context support for the broader GPT-5.x family, and enables DynamicCapability toolsets in durable execution workflows via DBOS and Prefect.
Engineering teams can now build agents utilizing the latest frontier models from Anthropic and OpenAI while benefiting from prompt caching to reduce latency and execution costs. The integration of DynamicCapability toolsets also simplifies deploying long-running, fault-tolerant agent architectures across specialized infrastructure.
Integrations
NICE finalized its 26.2 release cycle for CXone, introducing new workforce management and routing features. Key updates include the ability to block agent call transfers to ACD skills outside of their hours of operation, a Zendesk integration update adopting OAuth access tokens, and a new capability for CXone WFM to generate agent adherence reports directly from Amazon Connect data.
The updated transfer controls prevent interactions from being stranded in closed departments, improving routing efficiency and the end user experience. Furthermore, contact centers leveraging heterogeneous tech stacks benefit from unified workforce visibility through the Amazon Connect connection and more secure Zendesk authentication.
Integrations
My AskAI released its July 2026 changelog, introducing several new capabilities across its integration ecosystem. The AI agent now supports image reading and SMS replies in Gorgias, performs live inventory checks and product recommendations in Shopify, and includes new intervention controls for Freshdesk users.
Support teams using Gorgias can now process customer screenshots automatically without prompting for text descriptions. Shopify merchants gain an agent capable of dynamically checking stock levels to guide purchases, while Freshdesk users benefit from stronger human-in-the-loop oversight to halt or guide AI responses via internal notes.
Agent capability
LiteLLM released version 1.94.0, introducing router plugins configurable via proxy YAML and Auto-Router v2 featuring a soft-floor adaptive mode and session affinity. The update also adds interactive SSO sign-in for Model Context Protocol (MCP) client credentials, a unified DataTable UI for resource management, and a beta Cost Optimization dashboard.
Engineering teams gain programmatic control over model routing logic and enhanced security for handling MCP tool credentials. The new Cost Optimization dashboard provides visibility into API spending by tool and cache efficiency to help control infrastructure costs.
Integrations
Linx Security launched a new secure third-party agent access integration within the Snowflake security ecosystem. This integration enables enterprises to connect, govern, and audit the activity of autonomous AI agents by applying the same identity discipline used for human and system accounts.
Security teams face growing identity risks as AI agents increasingly operate outside of manual oversight. This integration allows organizations to enforce granular and real-time access controls to mitigate excessive privilege risks for third-party agents interacting with sensitive data inside Snowflake.
Human approval / guardrails
Kustomer launched a completely rebuilt agent workspace called the New Kustomer Experience, currently available to all existing users in early access. The redesigned interface features structural performance improvements to reduce page load times to under a second and natively embeds AI tools like Copilot, Signals, Summaries, and Tasks directly alongside the conversation timeline.
This overhauled workspace reduces cognitive load and micro-interruptions for human agents managing complex support escalations. By tightly integrating AI context into the core interface and drastically improving system speed, support teams can lower average handle times and resolve difficult interactions more effectively.
Agent capability
Kiro IDE version 1.0.242 adds a Kiro submenu to the editor context menu and introduces an Ask Kiro to Fix quick action for errors and warnings. The release includes a guided form for creating hooks, persists chat history across workspace folder changes, and updates the editor to Code OSS v1.108.2.
By integrating agent actions directly into the IDE context menus, developers can trigger code fixes and documentation generation with less manual effort. The guided hook creation and persistent chat history improve usability and ensure context is maintained across different project directories.
Integrations
Intercom added native support for Telegram, enabling businesses to connect their Telegram bot directly to the Intercom inbox. This integration allows the Fin AI Agent to be deployed on Telegram to automatically resolve customer questions and execute automated workflows.
Customer support teams can now deploy their Fin AI Agent across Telegram without requiring custom middleware or API development. This directly expands the agent's omnichannel reach, which is critical for brands serving international or mobile-first customer bases.
Workflow orchestration
Extend released an upgrade to its workflow orchestration layer that allows users to build complete document processing workflows in natural language using its Composer agent. The update also enables teams to validate these workflows with code or semantic rules and manage deployment directly through GitHub.
This bridging of natural language generation with GitOps allows non-technical operators to design document extraction logic while keeping engineering teams in control. Buyers can maintain version control, continuous integration, and pipeline reliability by managing the resulting workflows as code.
Workflow orchestration
Clay updated its TAM sourcing features with a faster, more credit-efficient Search powered by natural language, filters, and API or CLI access. The release also introduces a Lookalikes tool that continuously discovers up to 15,000 similar companies based on existing customer profiles, and a new Reporting dashboard that tracks how Clay-sourced records convert to pipeline and revenue.
Go-to-market and RevOps teams can now measure the direct ROI of their data enrichment efforts by linking sourced accounts to closed-won revenue. The improved search efficiency and automated lookalikes reduce both the time and credit spend required to maintain an active target account list.
Security / enterprise
Celonis introduced the Bring Your Own Key (BYOK) security model, enabling customers to manage their own encryption keys for data at rest using their own Key Management System.
Enterprise buyers can now retain complete ownership over their encryption keys, ensuring Celonis only has access for specific encryption and decryption operations. This significantly bolsters data protection and assists in meeting strict regulatory compliance requirements.
Integrations
Celonis and Valence Intelligence launched the Celonis Deductions and Disputes Manager, an AI-driven application available via the Celonis Platform Apps Programme. The tool uses AI agents to automate the aggregation of unstructured deduction data across fragmented sources, such as retailer portals and ERP systems.
Finance teams can drastically reduce manual effort in reconciling and disputing deduction claims, providing real-time visibility into their deductions position. The agentic application generates guided recommendations for claims processing with automated resolution options.
Observability / auditability
Automation Anywhere introduced a new enterprise AI evaluation framework and operational memory architecture designed to measure agent trajectory accuracy for multi-step tasks. The company also published a benchmark report demonstrating its base agent (aa-agent-v1) achieved top rankings across all pass levels on the public τ-bench evaluation.
Technical teams can leverage this evaluation framework to reliably benchmark AI agents on custom enterprise workflows instead of relying solely on public leaderboards. The validated τ-bench results provide quantitative confidence in the base agent's execution speed and reasoning capabilities.
Browser/computer use
Visualping released its Q2 2026 update, featuring a major upgrade to its Chrome extension that introduces action recording for clicks, scrolling, and logins, alongside one-click AI cloud monitor configuration. The update also adds a Model Context Protocol (MCP) server integration for AI assistants like Claude and ChatGPT, Scheduled Reports, Telegram alerts, and Google Single Sign-On.
Teams tracking dynamic or gated web pages can now easily record multi-step browser interactions directly via the extension for automated replay. The new MCP support allows developers and knowledge workers to query page changes and manage tracking tasks within their preferred AI assistant environments.
Agent capability
Unikraft detailed its agent infrastructure platform architecture, optimized for running AI sandboxes at scale. The platform leverages hardware-isolated microVMs to deliver 10-millisecond cold starts and stateful scale-to-zero capabilities, allowing agents to pause and resume exactly where they left off. It also features a dedicated network shield that enforces outbound traffic policies and isolates credentials from the agent environment.
Organizations building agentic AI can leverage this architecture to safely execute untrusted code without the overhead of traditional virtual machines. The scale-to-zero capability reduces compute costs during agent idle periods, while the network shield ensures credentials remain secure against prompt injection attacks.
Observability / auditability
Distyl AI introduced Canary, a self-improvement layer within its Distillery platform that continuously analyzes production AI outcomes to identify system failures. The feature automatically drafts and validates targeted improvements, such as prompt edits and routing changes, against real-world evaluation sets before domain experts approve them for deployment.
Buyers can significantly reduce the manual engineering time typically required to investigate and resolve AI system defects. By surfacing validated and evidence-backed improvements directly to domain experts, organizations can continuously optimize their production AI performance without risking unexpected regressions.
Workflow orchestration
Dify released version 1.16.1, adding a Tool Multi-Select Input for configuring multiple tool parameters simultaneously and a Workflow Node Locator that links run-log errors directly to the corresponding canvas node. The update also shifts the default OpenAI plugin API from Chat Completions to Responses to support the newly released GPT-5.6 model family.
Workflow builders gain faster debugging tools and more flexible parameter configuration directly on the canvas. The proactive shift to the Responses API prevents breaking errors and parameter restrictions for teams deploying the latest GPT-5.6 models.
Integrations
Demandbase has launched a strategic integration with DemandWorks, connecting its AI-powered buying signals and account intelligence directly to the DemandWorks managed multi-channel activation engine. This setup automatically transitions qualified accounts into coordinated engagement programs without manual list uploads.
B2B marketing and sales teams can now maintain a continuous workflow from account identification to audience activation. This integration reduces campaign launch lag and eliminates the need for manual data exports, allowing teams to act on buying signals much faster.
Deployment / data residency
camelAI completely redesigned its coding agent architecture, migrating from per-user virtual machines to a serverless model using Cloudflare Durable Objects. The updated execution stack stores project files in SQLite and R2, replacing open-ended bash commands with sandboxed JavaScript Code Mode running in V8 isolates.
This infrastructure shift removes the fixed costs of always-on virtual machines, leading to more scalable and cost-effective deployments. Additionally, restricting open-ended shell access to explicit platform methods improves credential control and system security.
Security / enterprise
Box rolled out a Classification-Based Access Policy in Box Shield to govern what internal and external AI agents can access. Administrators can now apply separate security rules for 'Download' and 'Read' actions based on specific file classification labels.
Because AI models do not need to physically download a file to comprehend and exfiltrate its data, the ability to block preview or read access is critical. This control ensures highly sensitive documents are protected against unauthorized processing by connected agents.
Observability / auditability
Box released the AI Insights Dashboard within the Admin Console to provide a comprehensive view of AI consumption across the organization. It tracks essential metrics such as AI queries over time, top agents by usage, and total AI units consumed.
Administrators can monitor usage patterns, identify heavy users, and effectively manage agent deployment costs. This visibility is essential for organizations looking to scale enterprise AI initiatives without risking unpredictable budget overruns.
Memory / state
Box introduced seamless integration between the inline AI composer and the sidebar in Box Notes. The AI conversation history is now maintained as a single session across both interfaces to prevent fragmented contexts.
End users no longer lose their train of thought or contextual data when switching between the inline text generation tool and the broader sidebar chat. This unified memory improves the user experience and the overall coherence of generated content.
Integrations
Spellbook launched a native integration for Google Docs, bringing its AI contract review, playbooks, and Ask Q&A features directly into the browser-based document editor. The integration allows legal teams to apply redlines and run reviews directly on the page. This bypasses the previous requirement to download documents and open them in Microsoft Word.
This expansion removes a major adoption barrier for organizations and early-stage startups standardized on Google Workspace. Buyers no longer have to compromise their foundational document management practices or maintain duplicate software licenses to access AI contract drafting capabilities.
MCP / tool calling / API
Sonar introduced seamless hooks for Claude Code and GitHub Copilot CLI within the SonarQube CLI, powered by a new Model Context Protocol (MCP) server. This update includes a pre-tool-use hook that scans for secrets locally, preventing hardcoded credentials from being transmitted to LLM providers during agentic workflows.
Developers can now securely integrate SonarQube's analysis directly into their local AI assistant workflows, enabling faster issue remediation while ensuring sensitive credentials remain safe. The MCP integration allows autonomous agents to fetch code quality context programmatically without leaving the IDE or terminal.
Integrations
SAP has announced the general availability of its generative AI assistant, Joule, within SAP Cloud ALM for all eligible customers. This initial release embeds a conversational AI assistant directly into the application lifecycle management platform to provide navigation assistance and contextual guidance on features and processes, laying the groundwork for future specialized AI agents.
Application lifecycle management teams can now use conversational AI natively within their platform to streamline workflow navigation and access contextual help. Because this initial Joule capability is provided at no additional cost and does not consume AI Units, buyers can immediately activate the assistant to improve productivity without impacting their current licensing costs.
MCP / tool calling / API
Oracle Fusion Cloud Applications introduced REST APIs that support the asynchronous invocation of published Oracle AI agent teams from external applications. This framework allows third-party tools to trigger agent workflows and use a returned job ID to retrieve status updates and final results without exposing raw JSON.
Developers can embed Oracle's specialized AI agents into custom portals, mobile apps, or collaboration tools instead of keeping them confined to the native Oracle UI. This provides teams the flexibility to build purpose-built interfaces while retaining Oracle AI agents as the underlying system of intelligence.
Agent capability
Optimizely updated its Opal AI platform, introducing Agent Builder and Skill Builder for configuring agents directly within chat. The release also includes nested workflows for multi-step automations, new remote MCP connectors for HubSpot, ZoomInfo, Gmail, and Google Search Console, as well as Safe URL Browsing powered by Google Web Risk.
Organizations can allow non-technical teams to build and iterate on specialized AI agents without writing code. The new enterprise MCP integrations and governance guardrails ensure that agents can securely access internal data while protecting against malicious web activity.
Observability / auditability
OpenRouter launched Classifiers in beta, enabling developers to automatically tag generation requests with structured metadata. Users can define custom taxonomies, such as task type or compliance category, and use a small model like Gemini 3.5 Flash Lite to tag the traffic. The resulting tags are aggregated and integrated directly into the Activity Explorer for usage reporting.
Teams gain powerful, built-in observability to track exactly what their AI agents are doing, which models they use for specific tasks, and where spend is accumulating. This eliminates the need to build and maintain custom logging, classification, and attribution pipelines.
Security / enterprise
OpenBox expanded its runtime governance platform into a framework-agnostic control layer. The update introduces integrations for over eight frameworks including LangChain, LangGraph, CrewAI, and Temporal. It adds capabilities such as pre-execution policy enforcement, session replays, human approval workflows, and tamper-proof audit trails mapped to the EU AI Act.
Organizations can enforce a single, unified governance standard across disparate AI agent frameworks without re-architecting their underlying systems. The new session replay and immutable audit logs give security and compliance teams the necessary evidence to verify agent decisions and satisfy regulatory requirements.
Integrations
Netlify has added support for Anthropic's Claude Opus 5 model to its AI Gateway and Agent Runners. Developers can now use the Anthropic SDK within Netlify Functions to access the model directly, benefiting from built-in caching, rate limiting, and authentication without managing separate API keys.
Engineering teams building agentic workflows can now leverage Anthropic's most capable model without building custom middleware for credential management. Centralizing model access through the platform also ensures organizations can enforce governance and track usage costs consistently.
Agent capability
Lyzr open-sourced SivaClaw, the AI agent it initially built to support its own Series B fundraising process. The agent runs on GitAgent using the OpenGAP protocol and handles automated workflows like investor communication, diligence material organization, and company briefings.
Startup teams and developers can now leverage an open-source AI assistant to manage operational tasks during their capital raises. By providing access to the underlying architecture, Lyzr allows users to inspect, modify, and extend the tool without being locked into a proprietary ecosystem.
Security / enterprise
Lovable rolled out a suite of enterprise governance features, including automated basic and deep codebase security scanning, Workspace Insights for activity tracking, and automated cleanup policies for abandoned applications. The platform also introduced reusable authentication controls via SSO.
IT administrators gain the necessary visibility and control to prevent application sprawl and shadow IT. The automated security scans and automated cleanup policies ensure that rapidly generated internal tools remain secure and compliant without requiring manual oversight for every deployment.
Integrations
LiteLLM added Day 0 support for Anthropic's Claude Opus 5 model across its integrations with Anthropic, Azure, Vertex AI, and Bedrock. The update natively supports Opus 5 features including Fast mode, adaptive thinking, prompt caching, tool calling, and computer use.
Teams can immediately upgrade to Anthropic's latest frontier model via a simple configuration change without altering application code. Supporting Opus 5 across multiple cloud backends allows organizations to balance rate limits and cloud spend while using the new capabilities.
Memory / state
Letta released the open-source trajectory package, which normalizes coding-agent sessions from Claude Code, Codex, and Letta Code into a single JSON-compatible format. The schema captures user messages, agent reasoning, and tool results while stripping redundant payloads to reduce token usage by up to fivefold.
Teams using multiple agent harnesses can now aggregate session data into a standardized format. This structure powers Letta's background dreaming process, allowing agents to extract lessons from past experiences across different tools and write them into persistent memory.
Observability / auditability
Legora introduced the Benchmark for Agentic Reasoning (BAR), an evaluation platform designed to test AI models across end-to-end legal workflows using real-world cases. The benchmark utilizes weighted scoring to continuously measure performance across different frontier models based on speed, cost, and citation accuracy.
Buyers evaluating Legora gain a transparent framework demonstrating how the platform performs on complex legal tasks compared to baseline models. This continuous observability builds trust by tracking whether new underlying models and system updates measurably improve output quality.
Security / enterprise
A critical remote code execution vulnerability (CVE-2026-0770) was disclosed for Langflow versions 1.7.3 and earlier. The flaw carries a 9.8 CVSS score, is reportedly under active exploitation in the wild, and allows unauthenticated attackers to remotely execute arbitrary code on affected systems.
Security and engineering teams must immediately identify and update exposed Langflow instances. Compromised AI workflow platforms present severe risks for lateral movement and credential theft due to their embedded access to external APIs, databases, and cloud environments.
Agent capability
Google Antigravity added support for defining custom agents using standard Markdown format via an agent.md file. The release also introduces progressive streaming for the codesearch command, an optional index argument for copy, and default read access to the system temporary directory.
These enhancements give developers a more accessible way to configure agents and improve the speed and efficiency of searching codebases and managing files within agentic workflows.
Integrations
Anthropic's Claude Opus 5 is now available as a selectable model across GitHub Copilot surfaces, including VS Code, Visual Studio, Copilot CLI, the cloud agent, and mobile apps. The new model is designed for complex, long-running agentic coding tasks and multi-step execution. Administrators for Copilot Enterprise and Business plans must enable the model policy in settings.
Providing access to Anthropic's newest model allows engineering teams to leverage stronger reasoning capabilities for intricate regression verification and autonomous workflows. This expansion ensures developers have the flexibility to route demanding tasks to an engine optimized for extensive tool coordination, improving overall reliability.
Memory / state
Cline updated its default session settings to use 'agentic compaction', which intelligently compresses long conversations to preserve the model's context window. Additionally, headless routines now default to 'YOLO mode' for fully unattended execution. The release also resolved an issue where the agent could not automatically locate CLI tools like the GitHub CLI.
Automated context compaction helps teams run longer and more complex autonomous tasks without exceeding context limits or incurring unnecessary token costs. Defaulting headless runs to unattended execution streamlines the integration of Cline into CI/CD pipelines by removing the need for manual approval steps.
Observability / auditability
Cline launched the first public release of 'Cline Code' for macOS, a standalone desktop application for running and inspecting agent sessions. The app natively supports Apple Silicon and Intel architectures and operates completely independently of an IDE. It features background auto-updates to ensure the client is always up to date.
A dedicated desktop app provides teams with a native interface for monitoring autonomous coding workflows outside of a code editor. This enhances observability and makes it easier for managers or reviewers to audit the agent's multi-step execution processes.
Agent capability
Arize announced Signal, a capability within Arize AX that utilizes autonomous agents to investigate and resolve software issues automatically. Signal surfaces failure patterns from production telemetry, uncovers root causes, and generates review-ready fixes as part of a self-improving feedback loop.
Engineering teams can move beyond manual debugging to autonomous software issue resolution to reduce downtime and operational overhead. This represents a shift from traditional observability to self-healing agentic systems.
Agent capability
Anthropic launched the Claude Opus 5 model, making it immediately available for use within Claude Code alongside a new fast mode that accelerates token generation 2.5 times at roughly double the cost. The system also introduces automatic fallbacks to Opus 4.8 when a request triggers Opus 5's heightened cybersecurity classifiers.
Developers can tackle highly complex agentic tasks with a more capable model while using fast mode to speed up iteration times when latency matters more than cost. The automatic downgrade fallback ensures that overly cautious security classifiers do not completely halt workflows.
Security / enterprise
Workday Adaptive Planning has achieved FedRAMP Authorization at the Moderate Impact Level. This certification confirms that the platform meets the federal security requirements for handling sensitive unclassified workforce and financial data.
Public sector and federal organizations can now adopt Workday's connected workforce and budget planning features. This allows government agencies to leverage Workday's operational forecasting tools within a fully compliant cloud environment.
Funding / partnership
Workato and L&T Technology Services jointly launched FinShield AI on the Workato Agentic Marketplace. The multi-agent assistant automates fraud investigation workflows, suspicious activity report generation, and fraud pattern analysis.
Financial institutions can reduce their fraud investigation times from hours to under ten minutes and cut manual reviewer workload by up to 50 percent. Because the assistant runs on the governed orchestration layer of the platform, all agentic actions remain tied to real identities and are logged for regulatory compliance.
Observability / auditability
TrueFoundry introduced Ask TFY, a conversational interface built directly into its AI Gateway that acts as an AI agent for platform operations. The interface allows engineering teams to use natural language to query real-time configuration data, analyze traces, troubleshoot bugs, and generate YAML configuration updates.
This feature reduces the time teams spend cross-referencing monitoring dashboards, trace logs, and routing rules across different systems. By allowing operators to diagnose failures and directly apply configuration changes from a single prompt, it lowers the operational overhead of managing complex AI deployments.
Integrations
xAI has launched a native Tavily search plugin within Grok Build, its developer platform. This integration allows developers building on the Grok ecosystem to invoke Tavily's agent-optimized web search and extraction API directly to retrieve structured, LLM-ready results.
For developers building research agents or autonomous systems using Grok models, this native integration eliminates the need to configure and maintain a custom web retrieval layer. It streamlines the architecture by providing direct access to real-time web data from within the Grok developer environment.
MCP / tool calling / API
Synthflow introduced the ability to play background audio while custom actions and individual Model Context Protocol (MCP) tools are running. This feature provides callers with auditory feedback while slower requests are processing.
This improves the end-user experience by preventing dead air during lengthy backend operations. It reduces caller abandonment rates and makes the voice agent feel more responsive.
Agent capability
Sierra acquired the AI startup Takeoff and introduced the Horizon Agent Platform. This new offering focuses on long-horizon AI agents designed to execute end-to-end tasks across sectors like lending and healthcare, rather than being limited to single-turn customer support conversations.
Enterprise buyers can now deploy autonomous agents that handle complex, multi-step workflows operating over extended periods of time. This expands the platform's utility into core operational tasks and introduces outcome-based pricing options tailored to industry-specific solutions.
Observability / auditability
Siena AI introduced Ask Siena, an internal AI agent built for brand operators to query their customer conversation history and business data in plain language. The capability functions as an automated data analyst to surface quantitative trends and qualitative quotes, and it includes an MCP integration allowing teams to connect their own custom agents to Siena Intelligence.
CX leaders and product managers can bypass manual ticket tagging or sample-based analysis to instantly identify retention risks, product gaps, and support trends. The ability to schedule recurring Slack briefs and pipe comprehensive customer intelligence directly into broader company workflows reduces reliance on fragmented reporting tools.
MCP / tool calling / API
Rootly introduced the Rootly Wizard, a guided command-line interface tool accessed via an npx command that automates initial workspace configuration. The tool allows administrators to instantly set up teams, on-call schedules, alert routing, Slack integrations, and status pages, culminating in a live end-to-end test page.
This drastically reduces time-to-value and onboarding friction for new deployments. Engineering teams can validate their entire incident configuration, including paging a real device, before broadly rolling the platform out to users.
Agent capability
Prophet Security announced the general availability of Prophet AI Threat Researcher and Emerging Threats within the Prophet AI Threat Hunter. This new AI agent autonomously scours open-source intelligence for emerging vulnerabilities and attacker activity, synthesizes its findings, and collaborates with reactive hunting agents to automatically build and execute hunting plans.
Security teams can significantly reduce the manual effort required to aggregate scattered threat intelligence and translate it into actionable hunts. By automating OSINT collection and the hunt execution cycle, organizations can proactively detect and defend against emerging campaigns with greater speed and minimal analyst intervention.
Funding / partnership
Pelico secured a strategic investment from AE Ventures, the venture capital platform of AE Industrial Partners. The funding is intended to accelerate the expansion of its AI-powered manufacturing orchestration platform across North America and deepen its integration within the aerospace ecosystem.
This capital injection ensures sustained product development and regional growth. It serves as a strong indicator of corporate stability, providing confidence to enterprise buyers who require a vendor with long-term viability and dedicated North American support.
Integrations
o9 Solutions integrated NVIDIA cuOpt, a GPU-accelerated optimization solver, into its Digital Brain platform. When run on NVIDIA B200 GPU infrastructure, the integration demonstrated over a 10x reduction in solve times for large-scale supply chain linear programming challenges.
The significant reduction in compute time allows enterprises to run complex supply chain optimizations iteratively against real-time data instead of relying on scheduled compute resources. This enhances planning agility and enables faster responses to disruptions without sacrificing solution quality.
Agent capability
Maven AGI introduced Charters, a feature that provides AI agents with structured instructions to trigger revenue actions such as upsells, cross-sells, and retention offers during support conversations. The feature uses Intelligent Fields to evaluate live context, customer intent, and journey stage.
Customer experience leaders can now drive commercial outcomes directly within service interactions without routing customers to a separate sales flow. This enables support teams to generate and measure revenue contribution while operating under defined governance controls.
Deployment / data residency
Jina AI introduced Jina On-Prem, a fully self-contained installation suite that packages all 28 of its embedding and reranking models into ready-to-deploy Docker containers. The containers run entirely offline with zero telemetry, license servers, or outbound connections. They serve models via API schemas compatible with OpenAI, Cohere, Voyage AI, Gemini, and the Elastic Inference Service.
Enterprises in highly regulated, air-gapped, or privacy-sensitive environments can now integrate local embedding models while maintaining strict data sovereignty. The drop-in compatibility with standard industry APIs ensures that organizations can transition their existing workflows and agent infrastructure to self-hosted models without incurring refactoring costs.
Workflow orchestration
HubSpot launched Agent Hub and Agent Builder in public beta for all Professional and Enterprise customers. Agent Hub provides a centralized workspace to monitor, activate, and manage all AI agents across marketing, sales, and service. Agent Builder introduces a no-code canvas that lets teams build custom agents and agentic workflows using natural language prompts, with agent runs now consuming HubSpot Credits.
Revenue teams gain a single console to oversee their AI agents, ensuring coordinated operations with shared customer context rather than siloed interactions. The natural language builder lowers the technical barrier to entry, allowing operations teams to deploy custom AI automations without custom development or separate field mapping.
Integrations
Equixly released its July 2026 product update, introducing a Service Overview dashboard for visualizing the complete API attack surface in a single view. The release also adds a native integration with Qualys VMDR for asset and finding synchronization, alongside Discovery scans that automatically adjust their speed to bypass rate limits.
Security teams can now synchronize vulnerability findings and asset inventories directly with Qualys, improving remediation workflows. The visual attack surface mapping and adjustable discovery scans provide security leaders with reliable test coverage against defended APIs.
Funding / partnership
Cognition acquired The Interaction Company of California, developers of the conversational AI assistant Poke, for a sum in the low nine figures. The company intends to integrate Poke's proactive, personality-driven interaction architecture directly into the Devin coding agent.
Engineering teams evaluating Devin can expect future updates that shift the agent from a rigid, command-based utility to a persistent, communicative colleague. This integration aims to improve developer collaboration by adding contextual continuity and proactive conversational reporting to software engineering workflows.
Observability / auditability
Augment Code released Cosmos Week 30, introducing session forking to branch conversations from completed turns and cost analytics with detailed charts for enterprise organizations. The release also expanded the Model Context Protocol (MCP) catalog with Google Workspace integrations and added Claude Fable 5 to the available model picker.
Engineering leaders gain granular visibility into AI usage costs and can better manage enterprise billing. The addition of session forking prevents context loss during complex problem-solving, improving developer efficiency.
MCP / tool calling / API
Arcade updated its runtime to support the new stateless Model Context Protocol (MCP) specification. The platform simultaneously maintains support for older session-based MCP versions, allowing both stateful and stateless connections to run in parallel.
Platform teams can incrementally migrate their MCP servers to the new protocol version on their own schedule without causing downtime for existing long-running agent workflows.
Integrations
Arcade launched four new agent-optimized toolkits for Power BI, Postman, Fireflies, and Insightly. The release also includes functional upgrades to the existing integrations for Gmail, Excel, Outlook, and Linear.
These native integrations allow agents to securely authenticate and execute complex tasks within familiar enterprise systems, expanding the range of delegated actions without developers needing to build custom connections.
Deployment / data residency
Anthropic launched the public beta of Claude for Government Desktop, bringing Claude Code into a FedRAMP High authorized environment. The government-specific deployment features robust administrative controls, tamper-evident audit logs, and dedicated spend governance.
Public sector agencies and contractors can now utilize AI coding agents without violating stringent federal compliance standards. The application provides necessary oversight features like local session history and detailed audit logs for compliance reviews.
Security / enterprise
Anthropic released the Claude Security plugin for Claude Code in beta. The multi-agent tool runs local vulnerability scans inside a Claude Code session, evaluating code for reachability and impact before generating patch files for developer review.
Engineering teams can integrate automated vulnerability scanning directly into the local terminal workflow rather than waiting for CI/CD pipelines. Because the agent only proposes patches for manual review, developers maintain full oversight of security remediation.
Agent capability
Swimlane launched Swimlane AI SOC for MSSPs, a platform built on Swimlane Turbine that uses agentic AI to automate alert handling and multi-step security investigations. The offering includes a central command center for multi-tenant management and cross-tenant threat intelligence sharing that aggregates enrichment results across all connected client environments.
Managed Security Service Providers (MSSPs) can increase their analyst capacity and scale operations without requiring a proportional increase in headcount. The platform's multi-tenant architecture ensures that MSSPs maintain strict data isolation and retain full control over their customer relationships while standardizing their service delivery.
Integrations
Workday announced the general availability of Workday Learning powered by Sana, integrating Sana's AI learning technology with Workday's Human Capital Management platform. The launch introduces a personal AI tutor that coaches learners through concepts in real time, alongside AI authoring tools that convert static files like PDFs and presentations into interactive courses. This Sana-powered version is now the default learning offering for new Workday customers.
Learning and development teams can reduce course creation time and simplify administration by managing training directly alongside core human resources data. The native connection allows organizations to replace fragmented learning management systems and deliver personalized employee training tied to individual roles, skills, and career paths.
Integrations
H Company integrated Linkup's web search API as the default retrieval layer across its computer-use agent stack. This embedded integration replaces the platform's previous search setup, enabling agents to pull real-time web context mid-task with a median latency of 1.5 seconds.
Enterprise buyers in regulated industries can now deploy these agents while maintaining full EU data residency compliance for web searches. The faster retrieval prevents latency compounding when executing hundreds of parallel sub-agent tasks.
Workflow orchestration
Qualtrics added the ability to select AI Topics when building conditions within a workflow. This enables users to create and trigger workflows automatically based on the presence of specific AI-generated topics detected in the data.
Support teams can now dynamically orchestrate routing and automated actions based on AI-identified customer intent or topics. This significantly reduces manual triage by ensuring the right workflows activate automatically when specific issues are detected in unstructured responses.
Deployment / data residency
Oracle released Base Database Cloud@Customer, a hybrid cloud infrastructure enabling mid-sized workloads to run Oracle AI Database on-premises. The solution includes the Private Agent Factory, a containerized no-code builder for deploying AI agents, and the Private AI Services Container for running private models isolated from third parties.
Allows organizations with strict data residency, regulatory compliance, or low-latency requirements to build and manage private AI agents entirely within their own firewalls. Customers can benefit from cloud operations and autonomous lifecycle management without exposing sensitive enterprise data to the public cloud.
MCP / tool calling / API
OpenRouter introduced a dedicated audio transcription endpoint (/api/v1/audio/transcriptions) that accepts base64-encoded audio. The service routes requests to Whisper-class and newer speech-to-text models across multiple providers, offering automatic load balancing and delivering JSON transcripts alongside usage statistics.
Developers can natively convert speech to text using their existing OpenRouter setup and API keys, completely avoiding the operational overhead of running standalone Whisper servers or integrating secondary SDKs. This centralizes billing and authentication for text and multimodal agent workflows.
Integrations
Novi Connect launched a direct integration with the product content syndication platform Syndigo. This connection allows brands to route AI-optimized product detail page content from Novi directly into Syndigo for publication across retail networks such as Amazon, Target, and Walmart.
This eliminates the need for teams to manually transfer optimized content between platforms. It allows brands to deploy AI-ready product data into their existing operational systems and distribution workflows without manual data entry.
Deployment / data residency
Mastra launched several updates to its hosting platform, including deployment environments for isolating production and staging workloads, EU and US geographic regions for localized data hosting, and managed workspaces that equip agents with persistent filesystems and secure sandboxing.
These infrastructure additions allow organizations to standardize agent release cycles, ensure strict data residency compliance, and easily run stateful agents that execute code without managing external sandbox services.
Human approval / guardrails
Manus launched Plan Mode, an opt-in feature that adds a human-in-the-loop review step before the AI agent executes a build or task. When activated, the agent generates an editable Markdown document detailing its planned goals, constraints, and implementation steps. Execution is paused and nothing is built until the user confirms or modifies the proposed approach.
Autonomous execution often burns through compute credits quickly if an agent misinterprets a prompt. This workflow control directly addresses cost predictability and output quality, making the platform much safer for enterprise users running complex, multi-step generation tasks.
Security / enterprise
Lovable achieved AIUC-1 certification, an independent security, safety, and reliability standard designed specifically for autonomous AI agents. The certification process involves quarterly third-party adversarial testing across secure code generation, secrets protection, and mitigation of hallucinations.
Enterprise buyers and security teams can deploy Lovable with greater confidence. This independent validation provides a clear framework to mitigate the unique risks associated with AI-generated code interacting with production infrastructure and user data.
Security / enterprise
LiteLLM launched an early beta of its AI Gateway completely rewritten in Rust, along with a new AIGatewayBench testing tool. The Rust gateway reduces per-request latency overhead to 0.7ms at p99 and lowers peak memory consumption to 21.8MB.
Enterprise teams processing high-throughput LLM traffic can lower their infrastructure costs and reduce memory-related crash risks. The sub-millisecond overhead specifically accelerates fast-loop applications like autonomous coding agents.
Observability / auditability
LangWatch introduced Langy, an automated AI engineering agent embedded directly inside the platform. Langy analyzes production traces to identify and cluster recurring behavior patterns, automatically writes scenario tests and evaluations, and opens pull requests on connected GitHub repositories to propose code fixes.
Teams can automate the remediation loop between discovering agent failures in production and generating regression tests. By drafting validated code changes directly in a pull request, this feature reduces the manual engineering overhead required to maintain and optimize deployed AI agents.
Human approval / guardrails
Langflow released version 1.11, introducing native Human-in-the-Loop checkpoints, Agent-to-Agent (A2A) protocol support, and AG-UI streaming for the Workflow API. The update also includes the NextPlaid extension bundle, providing out-of-the-box support for multi-vector and visual document retrieval.
Organizations building autonomous agents can now natively enforce mandatory human approvals directly within workflows. The addition of multi-vector and image retrieval significantly improves generation accuracy on complex or visually diverse documents.
Deployment / data residency
LangSmith introduced multiple platform updates, including control plane API support for assigning fixed resource tiers to deployments and IAM authentication support for Clustered Azure Redis. The release also updated LangSmith Model Context Protocol (MCP) tools to accept project UUIDs when querying run or thread histories.
These updates improve infrastructure management for enterprise teams running self-hosted or managed deployments by automating resource adjustments and strengthening database authentication. The MCP enhancements also provide developers with a more robust way to query trace data programmatically.
Observability / auditability
GitHub released a new Copilot usage metrics impact dashboard for enterprise administrators and organization owners. This visual dashboard categorizes users by AI adoption phases, measuring not just seat activation but depth of feature utilization based on a 28-day rolling window.
Buyers and administrators can now track ROI and product maturity beyond basic active-user counts. By surfacing underutilized features and suggesting targeted enablement steps, organizations can drive deeper adoption and realize greater value across their engineering teams.
Browser/computer use
Firecrawl upgraded its /search API endpoint with a custom relevance model that evaluates and scores every paragraph, list, and table on retrieved web pages against the user's query. Rather than returning full page content, the API now extracts and returns only the most relevant structural excerpts.
This pre-filtering step reduces the amount of text sent to downstream models, lowering input token consumption for web research tasks by up to 90 percent. Organizations can improve their AI agents' response accuracy and reduce inference costs without modifying their existing Firecrawl integration code.
MCP / tool calling / API
Dust released a significant update to its integration and agent capabilities, highlighted by making the platform available as a remote MCP server to allow access from any external MCP-capable client. The release also adds personal Zendesk account connections, a Lemlist MCP server for outbound prospecting, and AI-driven duplicate skill warnings during agent creation. Furthermore, Dust improved its Slack functionality to support agent reactions and status updates, and configured email interactions to keep entire threads within a single continuous conversation.
Enterprise teams can now embed their trusted AI agents and workspace knowledge directly into existing external tools via the remote MCP server, drastically reducing context switching. Deepened channel integrations with Slack, Zendesk, and Lemlist also empower organizations to automate more sophisticated workflows while ensuring seamless continuity in agent-led communications.
Workflow orchestration
depthfirst released Workflows, a new automation layer that converts security events into repeatable actions. The feature enables teams to define event-driven automations based on attributes like severity or repository, triggering actions such as Slack notifications, Jira or Linear issue creation, and direct pull request initiation.
This capability allows security teams to significantly reduce the manual overhead of data copying and routing by automating their incident response processes. Integrating depthfirst findings directly into existing DevSecOps pipelines accelerates the time from vulnerability detection to remediation.
Agent capability
Cursor has introduced Cursor Router, an intelligent model-routing layer that powers its Auto mode for Teams and Enterprise plans. It automatically analyzes each coding request and routes it to either a frontier model for complex reasoning or a cost-efficient model for routine tasks. The update includes three optimization modes: Intelligence, Balance, and Cost.
Engineering leaders can maintain code quality while reducing API spend by an estimated 30 to 50 percent. The feature also includes new admin controls, allowing organizations to enforce optimization modes, set default models, and manage allow and block lists for better cost governance.
Funding / partnership
CrowdStrike has formed a strategic partnership with Cerebras Systems to power its Falcon AI Detection and Response (AIDR) platform using Cerebras's high-speed AI inference technology. Under the reciprocal agreement, Cerebras will also standardize on the CrowdStrike Falcon platform to secure its own business operations.
Security leaders and enterprise teams deploying AI agents at scale can benefit from faster threat detection optimized for AI workloads. By leveraging specialized inference hardware, the AIDR platform is better positioned to process behavioral telemetry and defend against rapidly moving cyber threats.
Security / enterprise
ControlUp introduced longer Device Registration Codes with an enhanced security algorithm for tenant configuration and agent deployment. An updated management interface allows administrators to generate and oversee these codes via a dedicated permission.
Administrators must plan to migrate to the new registration codes since all legacy codes expire on August 31, 2026. Transitioning before the deadline ensures continued operations and prevents disruptions when registering new devices.
Observability / auditability
Cleric launched a production verifier that dynamically generates code to continuously poll live production telemetry and determine if an incident was successfully resolved. It evaluates whether the agent's suggested fix was actually applied and tracks if the underlying defect disappeared. These outcomes are then surfaced in a new in-product performance dashboard to account for the value delivered.
Evaluating the real-world effectiveness of an AI operations tool is notoriously difficult. This capability gives engineering teams deterministic proof of how often the system's proposed remediations eliminate root causes instead of just treating symptoms. It builds trust by providing objective success metrics backed by live environment data.
Integrations
Capsule Security launched a native integration with the Claude Platform using Anthropic's Compliance API. This connection enables continuous monitoring of Claude usage across an organization by ingesting supported activity logs directly into Capsule to surface security signals without altering Claude's underlying behavior.
Security and compliance teams gain centralized visibility into Claude Enterprise workflows, allowing them to proactively detect sensitive data exposure and anomalous agent activities. This audit-ready oversight helps safely scale AI adoption by bringing Claude activity directly into existing security operations.
Security / enterprise
Box announced new security and governance capabilities to control both Box-native AI agents and third-party agents connected via the Box MCP Server. The release includes agent guardrails, prompt injection detection, classification-based access policies, and agent activity oversight. It also introduces audit trails and human-in-the-loop controls for sensitive actions.
Enterprise buyers can deploy AI agents with greater confidence by enforcing strict data access policies directly at the content layer. These oversight features ensure agents do not expose or modify sensitive information beyond their approved scope.
MCP / tool calling / API
BAND released its Python SDK, providing the runtime and transport layer required to connect Python-based AI agents to its collaborative communication platform. The release includes framework adapters for LangGraph, Pydantic AI, CrewAI, and Anthropic, alongside the band-python-kit Docker Sandbox for running agents securely in isolated microVMs.
Engineering teams can now easily integrate their existing Python agents into BAND's shared interaction layer without having to build custom transport infrastructure. The included Docker Sandbox accelerates enterprise deployments by providing built-in isolation, safety rules, and strict egress controls out of the box.
Workflow orchestration
Augment Code released Auggie CLI 0.33.0, introducing cloud triggers to enable or disable persistent workflow triggers and a worker sessions view to display child worker processes. The update also improves Model Context Protocol (MCP) reliability with proactive usability probing, automated secret scrubbing, and dynamic tool list refreshing.
Platform teams can more reliably automate workflows and trace agent actions across child processes. Strengthened MCP reliability ensures safer and more stable connections to internal enterprise tools.
Agent capability
Astelia has introduced agentic capabilities to its reachability analysis platform to automate vulnerability management across the entire lifecycle. The new workflow evaluates newly disclosed vulnerabilities, assesses operational impact, and coordinates evidence-based remediation across IT and security teams. The platform integrates with over 100 MCP-enabled systems and requires human approval at key decision points.
Security teams face shrinking exploit windows and a growing volume of vulnerabilities that exceed human capacity to manually investigate. By automating repetitive analysis and orchestration tasks, buyers gain greater operational scale and can rapidly prioritize the vulnerabilities that pose actual risk. This approach allows organizations to execute targeted mitigation, such as network segmentation, without relying solely on software patches.
Integrations
Anchor Browser is now accessible as a Docker container image inside Daytona sandboxes. This integration allows AI agents running in Daytona workspaces to directly utilize Anchor's hardened Chromium browser, VPN, and managed authentication services.
Engineering teams can provision isolated compute environments alongside trusted browser infrastructure simultaneously. This reduces the operational burden of separately managing CAPTCHA evasion, proxies, and credential handoffs for web-operating agents.
Integrations
Ada has released a native integration enabling its AI Agent to directly answer Instagram direct messages. Connected entirely from the dashboard with a Meta login, the integration handles inbound and outbound text and media without requiring third-party middleware.
Organizations can now deploy their Ada agent on Instagram to automatically resolve customer inquiries on a major social channel. Bypassing external middleware lowers technical overhead and simplifies the architecture of an enterprise's social media support operations.
Agent capability
Ada has introduced availability rules for Glossary terms, enabling administrators to restrict when specific definitions are applied based on user variables. A term whose rule does not match the current conversation context is skipped. These rules can also be managed programmatically via MCP.
Support teams can now tailor their agent's vocabulary and definitions strictly to specific customer segments or conversational contexts. This improves the accuracy of the agent's responses and ensures specialized terms are only triggered when relevant.
Agent capability
Tinyfish launched Mako, a web-agent-native AI model built to operate the live web and execute multi-step workflows at scale. Trained on real production data from authenticated enterprise tasks, the model is now the default engine behind the Tinyfish Web Agent. Mako natively encodes web elements and caches page state to maintain context across long workflows without exceeding context limits.
Enterprise teams can deploy faster and more reliable web agents by leveraging a model purpose-built for web automation rather than a general-purpose LLM. Mako significantly reduces token consumption and latency, making complex tasks like authenticated data extraction and workflow automation more cost-effective.
Integrations
Vercel introduced a comprehensive feature rollout for v0, highlighted by general availability for the Shopify integration and a public preview of the Snowflake integration. The agent can now execute terminal commands via user permission prompts, autonomously resolve GitHub pull request merge conflicts, write SQL in DB Studio, and connect to OAuth-authorized MCP servers. Additionally, the v0 Max tier was upgraded to utilize the Claude Opus 4.8 model.
These capabilities elevate v0 beyond frontend UI generation, enabling teams to build enterprise applications and adopt real version-control workflows. The addition of OAuth MCP server support and terminal command execution provides developers advanced flexibility, while the underlying model upgrade improves overall reasoning capabilities.
Integrations
Trae has integrated with the BytePlus ModelArk Coding Plan. Users can configure Trae to use the BytePlus Plan service provider, granting direct access to a suite of foundation models including DeepSeek, Kimi, GLM, and ByteDance's Seed models.
Enterprise teams can manage their AI coding infrastructure through a centralized ModelArk subscription rather than individual third-party API keys. The integration supports automatic model routing based on performance and speed, offering flexibility and centralized cost control for organizations.
Integrations
CedCommerce has introduced an integration with Shopify Sidekick, bringing AI-powered marketplace intelligence directly into the Shopify admin. Merchants can now ask Sidekick questions about their marketplace orders, shipments, and setup, and the assistant will retrieve the relevant information from their connected CedCommerce app.
This integration allows merchants to access critical marketplace data directly through natural language queries in Sidekick without having to switch tabs or dig through dashboards. It streamlines daily operations by surfacing immediate answers about failed orders, delayed shipments, or setup issues within the primary Shopify environment.
Workflow orchestration
RegASK introduced an end-to-end agentic AI label compliance workflow that transitions product labels from draft to market-ready compliance in a single governed system. The workflow replaces manual reviews with an AI first pass and expert-in-the-loop confirmation, checking label elements against target market requirements and routing follow-up actions to specific owners.
Regulatory teams can reduce their manual workload by turning a multi-day review process across disjointed spreadsheets and documents into a five-minute governed workflow. By centralizing collaboration and generating a defensible audit trail, compliance leaders can catch errors before artwork moves to production, preventing costly rework and launch delays.
Security / enterprise
Rasa Pro 3.17.3 was issued to urgently backport the removal of the LangChain dependency cluster, addressing a critical arbitrary local file read vulnerability within the LangSmith TracingMiddleware. It also fixes missing input and output spans in Langfuse tracing.
Security teams managing Rasa Pro deployments on 3.17.2 or older must apply this upgrade immediately to prevent potential data exfiltration. The patch resolves a critical dependency risk without requiring teams to instantly migrate to the 3.18.x track.
Agent capability
Quantro Security released AI-Recon, an external attack surface scanner, and AI-XI, a vulnerability scoring model. Both tools are powered by the company's Exploit Harness, an autonomous system of AI agents that can generate verified exploits for disclosed vulnerabilities without human intervention.
Security teams can use these tools to map their external attack surface and prioritize vulnerabilities based on actual exploitability rather than theoretical risk scores. This helps defenders address flaws that autonomous agents can weaponize, which might otherwise be deprioritized by traditional frameworks.
Agent capability
Poolside released Laguna S 2.1, a 118-billion-parameter Mixture-of-Experts foundation model optimized for agentic coding and extended reasoning. The open-weight model activates 8 billion parameters per token, features a 1-million-token context window, and introduces separate thinking modes while remaining compact enough to run on a single NVIDIA DGX Spark desktop machine.
Organizations can self-host this coding model on hardware they control, keeping proprietary codebases secure and avoiding metered API costs. The sparse architecture delivers strong performance for long-horizon agent workflows efficiently, enabling enterprise-scale deployment without reliance on external cloud providers.
Observability / auditability
LangChain launched native Python tracing integrations in LangSmith for Pipecat voice agents. The integration captures end-to-end voice pipeline telemetry, including conversation audio, speech-to-text and text-to-speech latency, voice activity detection events, interruptions, and tool calls.
Production observability is a critical hurdle for voice AI. This integration enables teams deploying Pipecat to monitor, debug, and evaluate their real-time agents using LangSmith, allowing them to share the same evaluation workflows used for text-based agents.
Agent capability
OpenRouter detailed a new architecture pairing prompt caching with sticky routing. By passing a stable session_id in API requests, multi-turn agent conversations are automatically routed back to the specific provider holding a warm cache of the prompt.
Agents that rely on large, repetitive system prompts or intricate tool schemas can drastically reduce inference costs. Developers simply include a session_id in their payloads to ensure follow-up requests hit cached tokens, which are priced at roughly 10% to 50% of the cost of fresh inputs.
MCP / tool calling / API
n8n released version 2.32, which introduces granular scope selection for Model Context Protocol (MCP) OAuth consent. Users can restrict external MCP clients to specific resources, such as workflows or data tables, using Read-only or Custom presets. The update also adds detailed MCP connection tracking, enables publishing agents directly from the AI Agent builder, and expands node capabilities for Telegram and Google Calendar.
Security-conscious organizations can now safely expose their n8n workflows and data to external AI agents by enforcing least-privilege access. Restricting MCP client permissions ensures that unauthorized tools remain and uncallable, which minimizes the risk of unintended modifications to critical systems.
Memory / state
Mastra released Memory Extractors, a feature that enables agents to automatically pull and persist structured data from natural language conversations.
This automation reduces the manual overhead required to parse chat logs. It reliably captures user preferences and key entities, which improves context retention across multiple agent sessions.
MCP / tool calling / API
JetBrains Air, the agentic development environment, received an update adding support for multiple ACP (Agent Client Protocol) agents, including GitHub Copilot, OpenCode, Pi, and Cline. The release also introduces Java and Kotlin IDE intelligence and the ability to run Windows tasks in Docker.
Developers can now utilize a wider array of third-party coding agents within the JetBrains Air environment without being locked into a single AI provider. The addition of Java and Kotlin intelligence also enhances the contextual accuracy of agents working on JVM-based projects.
Memory / state
JetBrains launched JetBrains Context, a repository intelligence layer that incrementally builds a semantic index of codebases for coding agents like Claude Code, OpenAI Codex, and JetBrains Junie. It enables multi-repo search and semantic retrieval, allowing agents to query related concepts instead of relying on repeated keyword searches.
Engineering teams utilizing AI coding agents can reduce execution latency and API costs by giving their tools direct semantic access to large codebases. By avoiding repetitive file exploration, agents can more efficiently validate APIs and dependencies across an organization's entire repository landscape.
Agent capability
JetBrains open-sourced Mellum2, a 12B parameter model engineered specifically for routing, Q&A, sub-agents, and private AI use in software engineering systems. The model is designed to optimize latency, throughput, and cost for production AI workflows.
Organizations requiring strict data residency or looking to reduce API costs can deploy this open-source model locally for their software engineering agents. It provides a specialized alternative to general-purpose cloud LLMs for powering private development assistants.
MCP / tool calling / API
Itential launched Agentic Builder Skills on the Anthropic Claude Marketplace, bringing spec-driven development to infrastructure automation. The plugin connects Claude Code to the Itential Platform, providing 13 skills and 5 agents that translate natural-language requirements into deterministic, production-ready workflows and FlowAgents.
Infrastructure teams can accelerate automation development by allowing AI to draft complex workflows from plain-text specifications. Because the generated assets execute deterministically through Itential's established governance framework, buyers achieve the speed of AI-assisted coding without compromising operational safety.
Workflow orchestration
incident.io launched an update enhancing workflow automation and the actionability of its AI Agent. The release introduces a secure secrets store for workflows, the ability to sign workflow webhook requests, and triggers based on alert creation or resolution. Furthermore, the AI Agent has been upgraded to independently resolve incidents, direct escalations, and suggest both custom fields and timestamps during an active response.
Engineering and SRE teams can now orchestrate more complex and secure automations without hardcoding credentials in webhooks. The expanded AI Agent capabilities reduce manual overhead during incidents by allowing responders to execute critical workflow steps conversationally.
Security / enterprise
Harvey released its generally available integration with Intapp Walls for AI. The update allows law firms to sync their existing ethical walls and information barriers directly into Harvey. It automatically enforces matter-level restrictions and conflict-of-interest policies across Harvey Assistant, Vault, and Shared Spaces.
Confidentiality risk is a primary barrier to firm-wide AI deployments. By inheriting established Intapp policies rather than relying on manual configuration or attorney discipline, law firms can confidently scale Harvey usage while maintaining strict compliance controls.
Integrations
Gumloop released version 10.13.0, introducing the ability to sync organizational agent skills directly from GitHub repositories. The update also expands the Company Brain feature to ingest knowledge from any connector, adds a Reddit Ads MCP integration, and supports secure RSA key-pair authentication for Snowflake.
Enterprise buyers gain stricter version control for their AI workflows by managing skills via GitHub commits. The broader connector support and RSA authentication for Snowflake provide flexibility and compliance for handling proprietary business data.
Integrations
Block launched Buzz, an open-source collaboration workspace and Git forge that natively integrates Goose. Using the Agent Client Protocol, Goose agents can now be deployed into Buzz channels with their own cryptographic identities to participate in threads, review code, and execute automations alongside human developers.
This integration allows engineering teams to shift from running Goose as an isolated local tool to a multi-player team agent. Buyers gain shared agent context across the team, collaborative human-in-the-loop oversight, and a durable audit trail for every action the agent takes.
Agent capability
Google added new controls for model reasoning effort by introducing a dedicated effort command and launch flag. The platform also introduced stable model slugs for consistent sessions, subagent model configuration options, and a redesigned model picker.
This allows engineering teams to dynamically balance latency and reasoning depth depending on task complexity while ensuring reliable model routing across multi-agent setups.
Agent capability
Glia introduced three configurable Precision AI Response Modes for its voice and digital agent, Glia Banker. This release allows financial institutions to tune the conversational flexibility of their AI agent on a topic-by-topic basis, utilizing Strict, Rephrase, and Compose modes to balance automated guidance with rigid safety constraints.
Financial institutions often hesitate to deploy generative AI due to the compliance risks of unconstrained models. By providing granular controls that mandate approved answers for high-stakes topics while permitting conversational freedom for routine inquiries, this update allows banks to safely increase automation without assuming unacceptable risks.
Agent capability
Deepgram released new Nova-3 monolingual models to expand its language coverage. These models are available for both batch and real-time streaming transcription workloads.
Language expansions for Nova-3 allow global teams to deploy native speech models in additional regions without routing through translation or relying on less performant fallback models.
Workflow orchestration
DBOS released its July 2026 product update, introducing a Vercel AI SDK integration that provides out-of-the-box durability and checkpointing for AI agents. The release includes the production-ready DBOS Java v1.0, queryable workflow attributes via PostgreSQL JSONB, and significantly higher durable stream throughput. Additionally, DBOS added high-availability deployments for self-hosted Conductor and introduced DBOSify, a replacement framework for Temporal.
Engineering teams can now build resilient JavaScript or TypeScript agents using familiar AI SDKs without managing complex external orchestrators. The launch of a stable Java library and high-availability self-hosted deployments makes the platform viable for strict enterprise environments, while Temporal users gain a streamlined migration path.
MCP / tool calling / API
Day AI introduced Model Context Protocol (MCP) support, enabling the platform to function as a context server. This capability allows users to query Day AI's real-time customer graph directly from external MCP clients, such as Claude Desktop and Claude Code, without needing to open the CRM interface. The release also highlighted the expanded functionality of agent Skills, which execute recurring or trigger-based background tasks like drafting meeting follow-ups and sending pipeline summaries.
Exposing customer memory through an open protocol allows technical teams to natively integrate deep CRM intelligence into their preferred external tools and coding environments. Defining automated background skills shifts the agent from a reactive chat interface to a proactive system, reducing the manual administration required for revenue operations.
Funding / partnership
Cresta received an investment from W23 Global, a venture capital fund backed by five leading global grocery retailers including Ahold Delhaize, Tesco, and Woolworths Group. The funding supports Cresta's platform, which combines autonomous AI agents with real-time coaching and conversation intelligence for human teams.
This investment from major retail corporations signals strong enterprise validation and financial stability for Cresta. Buyers evaluating Cresta can expect continued development of AI features tailored for large-scale consumer-facing operations and enhanced support for complex retail contact center environments.
Integrations
ControlUp added end-to-end Azure Virtual Desktop host pool provisioning and management within DaaS IQ. Administrators can now create host pools via a guided wizard, import existing pools from Azure, and manage host pool deletion and filtering directly from the platform.
This centralizes cloud desktop administration by removing the need to switch to the Azure portal for provisioning. IT teams can natively orchestrate Azure Virtual Desktop resources and automate lifecycle operations entirely within ControlUp.
Human approval / guardrails
Confident AI introduced a reviewer-first workspace featuring Custom Annotation Forms and an Auto-Annotate capability. Teams can attach reusable, no-code evaluation forms with customized text, numeric, and choice fields to focused review queues, while Auto-Annotate drafts initial labels, explanations, and expected outputs across thousands of test cases.
This enables non-technical subject matter experts (SMEs) to participate directly in the AI quality and testing loop without requiring engineering configuration. The addition of automated labeling accelerates evaluation dataset creation, allowing teams to align automated evaluation metrics with human judgment much more efficiently.
Human approval / guardrails
Charm released Crush v0.86.0, bringing a new interactive Question Tool to the terminal-based AI coding agent. This capability allows connected models to present developers with single-choice, multiple-choice, and free-form text prompts with mouse support. The release also implements new session-affinity headers to enable provider-side prompt caching.
Structured input mechanisms reduce conversational friction by helping developers explicitly resolve codebase ambiguities when queried by the model. Additionally, utilizing session pinning can lower API inference costs and response latency for context-heavy development workflows.
Pricing / packaging
camelAI introduced a free tier powered by a self-hosted DeepSeek V4 Flash model running on AWS spot instances. The architecture routes requests through a Cloudflare AI Gateway with automatic failover to Azure's hosted DeepSeek service.
Evaluators can now test the platform's data analysis capabilities at no cost. The dedicated routing and failover infrastructure ensures that free-tier users experience consistent performance during their trials.
Agent capability
Blackbox AI introduced a two-model orchestration architecture that combines GPT-5.6 Sol and Claude Opus 4.8. The system uses a sandbox-executing critic agent to evaluate responses and gate the final output, achieving improved accuracy on the Terminal-Bench v2.1 benchmark.
Engineering teams gain access to a highly reliable execution path for complex coding tasks by routing prompts through multiple frontier models simultaneously. Buyers should note that evaluating outputs with a critic agent increases code accuracy but scales token costs to approximately two to three times that of a single-model run.
Observability / auditability
Arize added a User Friction Evaluator to both the Python and TypeScript SDKs. This built-in classification evaluator automatically detects and labels user frustration, corrections, or challenges to unrequested actions during conversational turns.
Product teams can systematically measure and flag negative user interactions. This makes it easier to track application quality and evaluate whether product updates effectively improve the user experience.
Agent capability
Apify released Apify AI in beta, a natural language interface integrated directly into the Apify Console. Users can now describe their data extraction or automation goals in plain English, and the AI automatically discovers, configures, and executes the relevant Actor to return the requested dataset.
This update lowers the technical barrier to entry for utilizing the platform by replacing manual keyword searches and configuration forms with simple text prompts. Buyers can enable non-technical team members to build and run data extraction pipelines without requiring developer assistance.
MCP / tool calling / API
AIsa launched the Bundled Skills System, introducing a unified capability layer that allows developers to discover, install, and reuse modular agent skills. By utilizing a single integration point, developers can empower Hermes Agent and other frameworks with reasoning models, live data retrieval, and real-world action execution without managing multiple API credentials.
Engineering teams can significantly reduce the integration overhead typically required to build autonomous workflows. The ability to deploy pre-configured skill bundles via a single unified API key accelerates development timelines and simplifies access control across multi-tool AI ecosystems.
Security / enterprise
Airia released an integration with the Claude Compliance API to extend AI agent governance and security visibility for Claude Enterprise users. The integration enables out-of-band evaluation of Claude conversations, uploaded files, and projects against configurable guardrails. It also detects sensitive data such as PII, PHI, PCI, and secrets in conversation content.
Security and compliance teams gain a single monitoring feed to correlate activity across Claude Code, Cowork, and Claude. Automated runbooks can route triggered guardrail violations directly to incident management tools like Slack and ServiceNow for rapid remediation.
Observability / auditability
Glossary usage is now surfaced directly in the Ada Conversation View through transcript events that indicate when the AI Agent applies glossary terms. Administrators can review translation modes and definitions used per turn, as well as filter conversations and Analytics reports by specific terms.
This enhanced visibility allows AI trainers to verify that the agent is properly adhering to brand terminology in real conversations. It drastically simplifies debugging and evaluating the effectiveness of the glossary in steering agent responses.
Integrations
Zendesk updated its Copilot auto assist feature to generate suggestions based on external and internal knowledge articles. Alongside this update, Zendesk released new knowledge connectors for Contentful, Jira, and Freshdesk, enabling the AI agent to pull information directly from these third-party systems.
Connecting external repositories directly to Zendesk's AI means support teams can leverage existing documentation without manually duplicating it. This expands the AI's ability to cover more complex topics and handle a wider range of customer queries autonomously.
MCP / tool calling / API
Workable expanded its Model Context Protocol (MCP) Server by adding 37 new tools, bringing the total to 94. The update grants AI assistants read and write access to performance reviews, account administration, and advanced candidate searches.
Recruiting teams can execute platform actions directly from external AI interfaces like Claude and ChatGPT. Organizations on higher-tier plans gain deep query capabilities across HRIS and candidate data to reduce context switching for daily tasks.
Agent capability
Vapi introduced Model Intelligence, allowing users to configure an assistant's transcriber, LLM, and voice models in a single click using pre-tuned presets like Balanced, High Intelligence, Ultra Fast, or Cost Saver. The update surfaces latency, cost, and quality metrics directly in the dashboard. End-of-call reports also now feature short-lived presigned URLs for secure recording downloads and are reliably delivered even when using Zero Data Retention.
The model presets and visible metrics help teams deploy appropriate model combinations that balance cost and performance without running manual benchmarks. The presigned URLs and robust zero data retention support ensure that compliance-focused engineering teams can safely and programmatically access post-call artifacts.
MCP / tool calling / API
Unit21 released a new API suite that allows financial institutions to call its financial crime AI agents directly from their existing case management systems. Instead of migrating to Unit21's full platform, customers can assign investigation tasks to the agent via API and receive the decision, gathered evidence, and a written narrative as structured data.
Compliance teams can now integrate Unit21's AI capabilities into their current workflows without undergoing a complete system replacement. This reduces the engineering time and financial risk associated with adopting AI for alert triage, allowing organizations to start small on specific queues before scaling.
Security / enterprise
Tabnine released version 6.4.4 of its private installation, introducing strict server-side model enforcement. This feature blocks disabled models from inference regardless of any client-side reload behavior. Additionally, the update unifies the default Tabnine model context limit at 200,000 tokens, an increase from the previous 180,000 limit.
Enterprise administrators can now reliably govern AI usage by blocking unapproved models at the server level, preventing developer workarounds. The expanded context window also improves code generation capabilities by allowing users to pass significantly larger codebase segments into their prompts.
MCP / tool calling / API
Sourcegraph released Code Finder in Beta through its Model Context Protocol (MCP) server, an agentic tool that helps AI coding agents rapidly locate relevant files and documentation within single repositories. Additionally, Sourcegraph made its redesigned Compare page generally available, introducing a file tree view and full code navigation for analyzing complex diffs.
Code Finder makes automated agent workflows faster and cheaper by minimizing the context window consumed during codebase discovery. Meanwhile, the upgraded Compare page provides essential oversight tooling for engineering teams to effectively review the massive, multi-file code changes generated by these agents.
Agent capability
Smallest AI launched its real-time voice AI platform, featuring a suite of specialized models built for low-latency conversational agents. The release includes Lightning for text-to-speech with approximately 100ms latency, Pulse for speech-to-text with built-in noise suppression, Hydra for native speech-to-speech full-duplex conversations, and Electron, a small language model optimized for voice reasoning. The platform also provides Atoms, an agent orchestration interface for configuring and deploying voice agents.
Enterprise buyers and developers can access a fully integrated modular voice AI stack without needing to stitch together separate vendors. The specialized models are designed specifically to reduce system-level latency and handle real-world audio noise, which is critical for high-volume enterprise telephony and customer support deployments.
MCP / tool calling / API
Rasa Pro 3.18.0 removes the LangChain dependency cluster to patch a critical file-read vulnerability, migrating vector stores to native SDKs and mandating a model retrain. The release also adds a new Capabilities API endpoint for dynamic conversational flows and a local document retrieval backend for the Rasa copilot.
Buyers must allocate time to retrain existing models and update vector store configurations. The removal of LangChain significantly hardens the platform against security risks, while the new Capabilities API simplifies building agents that can natively explain their available functions to end users.
Agent capability
Ramp launched Ramp Router in closed beta, opening its internal large language model gateway to developers. The OpenAI-compatible API dynamically routes AI requests across a catalog of models and service tiers. It uses Thompson sampling to optimize for latency, failure rates, and cost, allowing teams to switch models without rewriting application code.
Engineering and finance teams can manage AI infrastructure centrally to reduce token spend and prevent vendor lock-in. By offloading model selection to a dynamic router, organizations ensure they use the most efficient model per request while maintaining full visibility into per-request costs.
Workflow orchestration
Pydantic AI released version 2.14.0, introducing TemporalDurability, DBOSDurability, and PrefectDurability capabilities. These native durable execution layers replace the framework's deprecated wrapper agents. The release also adds support for Mistral's reasoning_effort parameter via thinking settings.
Engineering teams building long-running agent workflows can now use Temporal, DBOS, or Prefect for native state management and fault tolerance. This simplifies production deployments by eliminating legacy wrapper agents while giving developers finer control over Mistral models' reasoning capabilities.
Security / enterprise
Push Security released its July 2026 product update, introducing new in-browser controls for clipboard operations, file downloads, and file uploads. The release automatically categorizes applications and enables administrators to enforce policies that prevent users from copying sensitive data. The update specifically allows organizations to block the transfer of files to unapproved destinations, such as unauthorized AI tools.
Security teams gain granular data loss prevention capabilities directly within the browser, mitigating the risk of sensitive data exposure to shadow applications. The addition of telemetry streams and custom webhooks allows buyers to route violation events directly to a downstream security information and event management tool for incident response.
Observability / auditability
Opik added support for evaluating multimodal prompts that combine text and images. Teams can now execute structured experiments on multimodal traces directly from the UI or via the Python SDK, with the system automatically identifying vision-capable models using LiteLLM-style identifiers.
As enterprise AI applications expand beyond text, engineering teams require evaluation tools that can process images to accurately grade multimodal outputs. This native capability eliminates the need for manual workarounds, allowing organizations to enforce quality controls across their vision-enabled agents.
Agent capability
monday.com introduced an AI-generated campaigns feature in beta for its CRM module, enabling users to build complete marketing workflows using plain language prompts. The system acts autonomously to analyze sales call notes, aggregate CRM data, identify the target audience, and construct a preliminary campaign draft.
This automation reduces the manual effort required to segment audiences and draft marketing emails. Revenue teams can use this capability to rapidly launch targeted campaigns based directly on recent customer interactions and existing platform data.
Security / enterprise
Security researchers disclosed a sandbox escape vulnerability where the Antigravity agent could write a file that a trusted external tool would execute to bypass the sandbox entirely. Google acknowledged the risk of these indirect prompt injections and rolled out patches to address the issue.
Security and compliance teams must verify their Antigravity environments are patched to the latest version to prevent malicious workspace files from compromising host systems during autonomous tasks.
Memory / state
Genspark launched AI Workspace 6.0, introducing a persistent memory layer called SecondBrain that consolidates context across emails, documents, and meetings. The release also includes the Super Agent intelligence engine, a GenTeam collaboration layer for human-agent teamwork, and a companion hardware device called SecondBrain Note to capture offline conversations.
Enterprise buyers can leverage this unified memory architecture to reduce repetitive context-setting and allow AI agents to act on cross-platform data natively. The addition of hardware and dedicated team layers provides a structured environment for integrating autonomous agents into daily corporate operations.
Deployment / data residency
Decagon has announced its availability on the AWS Marketplace, enabling enterprises to deploy its conversational AI agents directly through their AWS environments. This move extends existing AWS data residency, compliance, and access controls to Decagon deployments from day one. Additionally, it allows teams using Amazon Connect to plug Decagon directly into their current contact center infrastructure.
This deployment option drastically streamlines procurement for existing AWS customers by allowing them to apply pre-committed cloud spend toward Decagon. It also accelerates time-to-value by eliminating the need for redundant security reviews, as the platform inherits the buyer's established AWS infrastructure controls.
Observability / auditability
Organization administrators can now access conditional rate-limit impact metrics for users and reviews via the Billing and Usage dashboard. The update includes a Most Active Users table and alerts administrators when monthly spend reaches 90 percent of the configured cap.
Equips buyers on metered plans with the transparency needed to track highly active users, anticipate rate limits, and proactively manage their AI agent usage budgets.
Integrations
CodeRabbit's interactive Change Stack interface has been expanded to support Bitbucket Cloud and Bitbucket Data Center. Authenticated users with write access can now launch AI-agent tasks, run review actions, and monitor merge blockers natively within Bitbucket.
Extends CodeRabbit's multi-layered code review and AI-agent interface to Atlassian-centric development teams, closing a critical feature gap for enterprise Bitbucket users.
Observability / auditability
Browser Use updated its Browser Harness v4 API to support client-controlled session recording, allowing API callers to explicitly enable or disable recording via the browserSettings.record parameter. The release also added a new v3:browsers:read OAuth scope that grants external applications read access to existing browser sessions.
This provides enterprise users and developers with granular control over data privacy, allowing them to disable recording during sensitive tasks. The new OAuth scope simplifies the integration of live browser monitoring into third-party dashboards and internal observability tools.
Observability / auditability
Agno released version 2.8.0, introducing the agno.scorer module to quantitatively evaluate and test agent runs. The update adds CodeScorer for typed field comparison, JudgeScorer for LLM-as-a-judge evaluations, and ToolCallScorer for deterministically verifying tool executions, with support for synchronous and asynchronous execution.
Testing non-deterministic AI agents is a significant operational hurdle. The new scoring module gives engineering teams built-in, automated evaluation primitives to rigorously benchmark agent performance and ensure reliability before production deployment.
Funding / partnership
Abridge acquired the founding and engineering team from Altrina, a startup specializing in browser-based AI agents and healthcare workflow automation. The newly integrated team will focus on advancing Abridge's capabilities in deploying AI agents within complex clinical environments.
This acquisition signals an intent to move beyond passive ambient transcription toward active, browser-based workflow automation. Enterprise buyers can anticipate future platform updates that automate broader clinical and administrative tasks.
Security / enterprise
Actian expanded AI Analyst (formerly Wobby) with a Data Steward Agent that synchronizes business definitions and enforces data policies. The update also introduces Data Observability Agents to monitor pipeline health and a Model Context Protocol (MCP) server, allowing third-party AI agents to query governed business logic via open standards.
Data teams can now safely expose their semantic layer to external AI workflows without losing policy controls or context. This prevents external agents from operating on stale or ungoverned data, ensuring enterprise compliance and reducing hallucination risks.
Integrations
NeuBird AI released a new native integration with ControlUp to bring automated root cause analysis to virtual desktop infrastructure (VDI) monitoring. The integration acts as an agentic analyst that continuously analyzes performance metrics, identifies relationships between hosts and sessions, and automates VDI investigations.
This integration enables IT operations teams to shift from reactive ticket resolution to proactive VDI management by automating routine troubleshooting. It reduces the time administrators spend manually correlating desktop performance data and helps prevent infrastructure changes from impacting end-user experience.
Human approval / guardrails
Anthropic released Claude Code version 2.1.215, removing the agent's ability to autonomously trigger the /verify and /code-review skills. Developers must now explicitly invoke these commands to initiate a review cycle.
Engineering leaders can better control token consumption and cost by ensuring the agent only performs expensive verification cycles when requested. This explicit boundary prevents the agent from extending sessions unnecessarily during simple mechanical code changes.
Memory / state
Mem0 introduced a multi-signal retrieval redesign in its open-source memory algorithm. The update replaces external graph store support with built-in entity linking, running semantic similarity, BM25 keyword matching, and entity matching in parallel before fusing them into a single result score.
This architectural shift eliminates the infrastructure overhead of maintaining an external graph database. It also provides measurable performance improvements on temporal queries and multi-hop reasoning, allowing agents to handle complex user histories more effectively without requiring existing pipeline modifications.
Integrations
Consio released an update introducing a native integration with the eDesk customer support platform, alongside various AI voice agent enhancements and bug fixes. The eDesk connector automatically generates support tickets for completed calls, syncing customer data, call dispositions, and AI-generated conversation summaries directly into the eDesk workspace. It also embeds a direct link to the full call recording and transcript within the ticket.
Support teams utilizing eDesk can now manage voice channel interactions in the exact same system as their text-based tickets. This reduces context switching and manual data entry, enabling agents to quickly review call outcomes and determine necessary follow up actions from a unified dashboard.
MCP / tool calling / API
Arize released Phoenix 19.1.0 with a one-command setup via the Phoenix CLI to connect coding agents like Claude Code and Cursor to the Phoenix MCP server. The command automatically registers the server into the agent configuration and defaults to OAuth login.
Developers save time and reduce configuration errors when connecting their preferred AI coding agents to Phoenix observability data. This accelerates the adoption of agent-based workflows within an organization.
MCP / tool calling / API
StackBlitz launched Bolt Slides, an open-source library and in-app capability that enables AI agents to generate interactive presentation decks as responsive React web apps. Each slide functions as a live web component, supporting embedded 3D models, live data visualizations, and interactive UI elements generated from a single text prompt.
This expands the product's utility beyond standard web application development, allowing teams to quickly generate dynamic sales pitches, internal reports, and interactive product demos. Because it is available as an open-source skill library, developers can also integrate this presentation framework with other coding agents like Cursor or Claude Code.
Workflow orchestration
Skyvern launched its Agentic Process Automation (APA) Control Plane to govern complex browser automations across credential-guarded portals. The platform introduces save-draft blocks to persist partial progress during session timeouts, alongside centralized credential management, human-in-the-loop approval gates, exception escalation, and full audit logging.
Operations teams can orchestrate long-running, multi-step workflows across unpredictable external portals without the risk of silent failures or total restarts upon timeout. Features like state-aware authentication and approval gates provide the governance and compliance controls necessary for enterprise-scale robotic process automation replacement.
Agent capability
Shopify has overhauled its Collections merchandising tool, allowing merchants to use Sidekick AI to build or edit product collections through plain-language requests. The update also enables Sidekick to pull live product sources from apps that track bestsellers or trends and keep them synced across collections and workflows.
Merchants can now curate complex product collections dynamically using conversational AI rather than manual configurations. This significantly reduces the time spent on merchandising tasks and allows teams to blend automated conditions with hand-picked products more efficiently.
Agent capability
Resolve AI officially launched its AI Production Engineer, an agentic AI system designed to automate production engineering tasks. The platform autonomously troubleshoots issues, manages operational tasks, and investigates incidents by leveraging context from source code, telemetry, and cloud infrastructure.
Enterprise engineering and DevOps teams can adopt this platform to significantly reduce mean time to resolve (MTTR) for production incidents. By automating repetitive on-call tasks, the product frees engineers to focus on higher-value development work.
Funding / partnership
Resolve AI announced a $125 million Series A funding round led by Lightspeed Venture Partners. This brings its total funding to over $150 million and values the company at $1 billion, providing capital to scale research and development for its production AI agents.
This substantial funding and unicorn valuation provide strong signals of market traction and vendor stability. Enterprise buyers can confidently invest in the platform knowing the company has the resources to sustain long-term enterprise-grade support and innovation.
Integrations
Prime Intellect integrated NVIDIA Vera CPUs, NVIDIA Dynamo inference orchestration, and NeMo Gym open-source reinforcement learning tools into its Lab platform and sandbox infrastructure. This combined hardware and software alignment optimizes continuous reinforcement learning loops for agentic AI models.
Engineering and research teams can execute post-training workflows more efficiently on Prime Intellect infrastructure. The integration enables a reported 30% increase in sandbox throughput per CPU, lowering the overall compute costs required for continuous agent refinement.
Security / enterprise
Paxton AI introduced Secure Matter Collaboration, providing litigation teams with a shared workspace for managing case files and generated work product. The capability allows multiple attorneys to jointly build chronologies, review evidence, and draft documents without overwriting previous work.
Law firms can now collaborate directly on a single matter without resorting to email chains, forwarding attachments, or sharing credentials. By integrating ethical walls and individual account access to a centralized workspace, the platform improves team alignment and meets strict enterprise security requirements.
MCP / tool calling / API
nexos.ai introduced an Anthropic-native Messages API endpoint that accepts requests in Anthropic's wire format and forwards them byte-for-byte to the provider. This update ensures that native Anthropic features, such as cache_control markers for prompt caching, are preserved rather than stripped by intermediate translation layers.
Engineering teams can point existing applications built on the Anthropic SDK directly to the nexos.ai gateway without rewriting code to an OpenAI-compatible format. This unlocks native prompt caching, allowing organizations to drastically reduce token costs and latency for workloads that use repetitive or long-context prompts.
Pricing / packaging
Mastra published its platform pricing model, introducing a free Starter tier with pay-as-you-go allowances alongside a Teams tier. The Teams plan provides increased monthly capacity, reduced usage rates, and additional production controls.
Evaluating buyers gain clear visibility into the cost of scaling Mastra deployments. The structured upgrade path accommodates organizations that require higher throughput and advanced enterprise controls.
Security / enterprise
Make introduced Private spaces, giving each team member an isolated workspace where their scenarios, connections, and module settings are from others. The feature also includes admin controls for Enterprise plans, such as read-only visibility into private spaces and the ability to set individual credit limits.
This allows larger organizations to safely onboard more employees by ensuring that personal or sensitive credentials are not exposed to the entire workspace. It also helps operations teams prevent runaway costs by capping automation usage per user.
Integrations
Make integrated OpenAI's GPT-5.6 model family, including the Sol, Terra, and Luna variants, into its AI Toolkit, OpenAI modules, and AI Agent app. The update allows users to build workflows utilizing the models' 1.05-million context window for complex reasoning or rapid data extraction.
Automation builders can now optimize for cost, speed, and capability within a single workflow. High-volume extraction tasks can be routed to the cheaper Luna model, while complex logic or deep reasoning on massive documents can leverage the Sol model.
Agent capability
LiteLLM introduced Router Plugins, enabling users to configure and chain custom plugin pipelines to determine which model processes a given input.
Platform engineering teams gain granular, programmatic control over model routing logic. This facilitates customized cost optimization, dynamic failover strategies, and compliance checks before requests are sent to an upstream provider.
Agent capability
Kana launched Category Intelligence, an agentic application that continuously ingests external market signals, such as search trends, social signals, and analyst reports, and synthesizes them with proprietary point-of-sale (POS) data. The product offers a conversational interface for category managers to query the combined data in plain language.
Category managers and marketers can now anticipate demand shifts weeks before they appear in standard scan data. By cross-referencing external signals with internal transaction records, teams can make proactive, decisions for brand planning and innovation reviews without needing SQL expertise.
Deployment / data residency
IBM released the 10.0.2607.0 Major Update for Sterling Order Management System and Intelligent Promising containers. This release brings previously introduced agentic AI features, such as the Cancel Orders Flow AI agent and expanded Order Information AI agent search, to containerized environments.
Organizations deploying IBM Sterling in containerized environments can now utilize recent AI-assisted workflow updates. This allows engineering teams to deploy the 10.0.2607.0 version of the AI toolkit without being forced to use the SaaS or on-premises deployment models.
MCP / tool calling / API
GoodTime released a Model Context Protocol (MCP) integration that connects external AI assistants like Claude and ChatGPT directly to GoodTime data. The server integration allows users to query interviews, interviewers, and tags directly from their preferred AI clients.
Talent acquisition teams can integrate their existing AI assistants with GoodTime's scheduling platform without building custom API connectors. This enables direct conversational access to interview data, speeding up the retrieval of candidate and schedule information for recruiters.
Human approval / guardrails
Genspark redesigned its AI Slides canvas into a full presentation workspace, adding a left outline rail, tabbed project views, and comprehensive zoom controls. Users can now watch the AI generate slide decks page by page, verify figures, annotate changes like a printed draft, and generate speaker notes before using the Presenter View.
This interface overhaul gives users greater visibility and editorial control over AI-generated presentations. By enabling real-time human intervention and markup, teams can ensure the final output strictly adheres to their messaging and formatting standards before presenting.
Agent capability
Deepgram updated its conversational Flux models to support automatic numeral formatting. The feature converts written numbers into digits, such as changing 'nine hundred' to '900', and applies to both Flux English and Flux Multilingual models.
Properly formatted numerical data is critical for voice agents capturing phone numbers, addresses, and pricing. Native handling of numerals directly within the speech recognition model prevents developers from needing to build separate formatting logic.
Agent capability
Deel announced the Deel AI Workforce, a suite of ready-made autonomous AI agents embedded directly within its global HR and payroll platform. These agents, which include specific roles like The Hiring Guru and The PTO Fairy, integrate with existing enterprise tools to execute multi-step administrative workflows from start to finish.
HR and operations leaders can automate complex, multi-country tasks like hiring compliance, visa checks, and IT provisioning. By delegating these workflows to specialized agents operating within established enterprise guardrails, organizations can scale global operations efficiently without a proportional increase in administrative headcount.
Integrations
Cursor released enhancements to its Slack integration, enabling the agent to share an execution plan before beginning a task. The update adds support for multi-repo environments directly from Slack and introduces cross-channel workflows, allowing the agent to pull context from other channels or threads.
Development teams using Slack for operations gain better visibility and intervention controls over autonomous agent actions before execution. The multi-repo support enables teams managing complex architectures to execute cross-service code changes without leaving the chat interface.
Memory / state
Cognee released version 1.4.0, featuring an optional dataset-level overview index that groups ingested documents into topic clusters and generates short summaries. The release also upgrades the ingestion pipeline to automatically chunk documents and retry failed uploads, and introduces new API endpoints for dataset management.
For organizations building AI agents, the new overview index improves search relevance for broad queries by providing topic-level context rather than relying solely on granular document matching. The ingestion and API improvements make it more reliable to continuously feed large volumes of data into the product's persistent memory graph.
Human approval / guardrails
Repository administrators can now disable the 'Fix Failing CI' and 'Resolve Merge Conflicts' finishing touch commands for GitHub and GitLab integrations. Disabling these options hides the corresponding checkboxes and prompts the agent to decline direct execution requests.
Grants engineering leaders greater control over how AI remediation tools are used within their repositories, ensuring automated fixes do not circumvent team policies for manual review or conflict resolution.
MCP / tool calling / API
Browser Use released the Browser Harness agent API to the public. The update also introduced new AI models across the platform and lowered the pay-as-you-go top-up minimum.
Developers gain direct programmatic access to build custom web automation applications on the Browser Harness infrastructure. The reduced financial minimums allow smaller teams to adopt the tool and start experimenting with new models at a lower initial cost.
Observability / auditability
Box launched the AI Insights Dashboard within the Admin Console. This new dashboard provides administrators with a comprehensive view of AI consumption across the organization. It displays key metrics such as AI queries over time, AI usage by product, top agents by usage, and total AI Units consumed.
IT administrators and operations teams can leverage these metrics to monitor consumption and identify high-usage individuals or agents. This visibility enables organizations to make budgeting decisions and manage enterprise costs effectively.
Security / enterprise
Arize Phoenix 19.0.0 introduces an embedded OAuth2 authorization server and centralized REST CRUD capabilities for API key management. The release also includes a beta Remote MCP server that allows MCP-compatible clients to directly query and operate on projects, traces, and datasets without requiring a local installation.
Enterprises benefit from robust authentication controls and centralized credential management. The built-in Remote MCP server securely connects coding agents directly to the platform to simplify local development workflows.
Pricing / packaging
11x published transparent starting prices for its Growth tier plans, shifting away from an exclusively quote-based enterprise model. The outbound AI SDR, Alice, now lists plans from $2,000 per month based on lead volume, while the inbound agent, Julian, starts at $5,333 per month for voice capabilities and $2,417 per month for chat on the Growth plan. Pro and Enterprise plans remain custom-quoted.
Providing clear starting costs enables revenue teams to model return on investment and compare the platform against human SDR headcount without entering a lengthy procurement cycle. The bundled pricing also simplifies total cost of ownership calculations by including CRM sync, onboarding, and infrastructure setup in the base fee.
Security / enterprise
Sonar introduced SonarQube Server 2026.4, featuring built-in architecture analysis that automatically visualizes source code dependency graphs across commercial editions. The release also added beta support for Cross-Translation-Unit (CTU) and taint analysis for C and C++, allowing the symbolic-execution engine to track untrusted data flows across multiple files.
The automatic architecture mapping helps engineering teams enforce intended system design without manual configuration. Furthermore, the cross-file security capabilities are crucial for safety-critical applications, detecting complex injection vulnerabilities that standard single-file scanners typically miss.
Workflow orchestration
Oracle introduced SQL Search (NL2SQL) for Enterprise AI Agent workflows within OCI Generative AI, enabling agents to convert natural language requests into validated SQL. The capability uses a semantic enrichment layer mapped to structured data vector stores to safely query data stored in Oracle Autonomous AI Databases.
Developers building AI agents on OCI can now safely grant them access to structured database records using natural language. This eliminates the need to copy data or write custom SQL generation logic, maintaining strict governance while expanding the analytical reach of the agents.
Security / enterprise
GitHub introduced new enterprise security and administration controls for the Copilot code review agent. The agent now runs behind a firewall by default to restrict network access, features independent runner configurations separate from the Copilot cloud agent, and supports custom setup steps via a YAML file.
Enterprise administrators gain granular control over network security and infrastructure isolation during AI-assisted code reviews. Operating behind a default firewall addresses strict compliance requirements for enterprise environments and limits potential data exfiltration.
Workflow orchestration
Anthropic released Claude Code v2.1.212, introducing background session isolation for the /fork command and hard budgets for automated actions. The update limits WebSearch calls and subagent spawns to 200 per session to prevent runaway loops, and enforces strict authorization prompts for file-modifying commands in Plan mode.
Engineering teams get stronger guardrails against unintended token spend and unwanted file mutations during autonomous execution. The ability to fork conversations into true background sessions allows developers to spin off parallel agent tasks without interrupting their primary interactive workflow.
Integrations
Agno released version 2.7.4, adding toolkits for Plivo, The Context Company, and Superserve, which provides a Firecracker-based sandbox for secure agent code execution. The release also updates the CLI with new project deployment starters for Azure, Helm, Modal, and Render.
Engineering teams can now safely execute agent-generated code in isolated cloud environments without maintaining custom sandboxing infrastructure. The newly included deployment starters also reduce the friction of hosting agents on enterprise platforms like Azure and Kubernetes.
Workflow orchestration
Validfor introduced AgentV, a voice-based AI agent designed for GxP digital validation environments. The tool allows validation professionals to retrieve records, navigate documentation, create new entries, and execute authorized workflows using natural voice commands.
This conversational interface reduces the need for manual data entry, accelerating compliance tasks for enterprise operations. The agent strictly adheres to existing role-based permissions, ensuring that organizations maintain the traceability and human oversight necessary for pharmaceutical regulatory standards.
Workflow orchestration
Searchable introduced Skills, a feature that allows users to turn previous interactions with the Searchable Agent into reusable workflows. Users can save the exact instructions, context, and operational steps from a chat session to rerun AI search optimization tasks without rebuilding them.
Growth teams and agencies can use Skills to standardize their operations into a shared playbook. This eliminates repetitive prompting and provides consistent execution for recurring tasks like weekly visibility reporting and content gap analysis.
MCP / tool calling / API
Rogo integrated with Snowflake's managed Model Context Protocol (MCP) server. This connection allows Rogo's AI agent, Felix, to securely discover and reason over governed datasets stored in Snowflake. The integration utilizes OAuth 2.0 authentication to connect the data directly to Rogo's platform without deploying separate middleware.
Financial institutions can analyze proprietary information using Rogo's AI without exporting data outside their established security perimeter. This setup respects all existing Snowflake governance rules, including role-based access and data masking, and eliminates the need to build bespoke data pipelines.
Agent capability
Replit introduced a continual learning system for Replit Agent, powered by an internal agent-based architecture. The system automatically processes user feedback, proposes system improvements, and validates changes using benchmarks and A/B testing.
Buyers benefit from an agent that continuously updates its coding accuracy and problem-solving logic based on real-world usage. This automated feedback loop enables the product to improve over time without relying entirely on discrete, major version releases.
Agent capability
OpenRouter released a unified API endpoint that supports multiple AI modalities, including text, image generation, video, text-to-speech, speech-to-text, and embeddings. Developers can switch between models and modalities by changing the model string and content type in their request to a single OpenAI-compatible base URL.
This update significantly reduces the integration overhead for engineering teams building multi-modal applications. Instead of managing separate SDKs, authentication schemes, and billing for different providers, organizations can centralize their AI access. It also enables consistent routing controls and fallback logic across all modalities.
Integrations
Kilo Code released version 7.0.7 of its JetBrains plugin to add native support for custom OpenAI-compatible model providers. The release also improves setup validation, error dialog handling, and provider disconnection workflows.
Development teams using JetBrains IDEs can now connect to self-hosted models or corporate proxy endpoints that conform to the OpenAI API specification. This ensures feature parity with the VS Code extension and expands local deployment flexibility.
Security / enterprise
Kilo Code added network restriction enforcement for sandboxed agent commands running on Linux environments. The update blocks unauthorized network access across TCP, UDP, IPv4, IPv6, and descendant processes, while giving administrators the ability to whitelist specific external network destinations.
Security and IT teams can confidently deploy the agent in locked-down or isolated environments by tightly controlling its network egress. This mitigates the risk of unauthorized data exfiltration or unintended API calls during autonomous execution.
MCP / tool calling / API
JetBrains introduced new AI agent capabilities as part of its 2026.2 IDE release wave. The update adds an Agent Skills manager in IntelliJ IDEA for handling custom AI instructions, as well as three new database-focused agent skills in DataGrip powered by the Model Context Protocol (MCP). The 2026.2 release also brings native GitHub Copilot integration and AI completion support for third-party model providers.
Development teams can now extend their JetBrains AI agents with reusable domain knowledge and connect them directly to databases for automated querying and schema management. Support for MCP tools and third-party model providers gives organizations greater flexibility to customize their AI stack while maintaining CLI consent controls for database operations.
Agent capability
Demandbase introduced AI Chat in Demandbase One for Sales (DBS), a conversational assistant that allows sellers to ask natural language questions about accounts. Currently in beta, the feature combines first-party CRM data with Demandbase's third-party firmographic, technographic, and intent intelligence. It delivers account insights, buying group assessments, and outreach recommendations directly within the sales platform.
Sales teams can significantly reduce the time spent navigating dashboards and building reports by directly querying their data. The immediate availability of prioritized outreach recommendations accelerates sales cycles and helps representatives capitalize on high-intent buying signals faster.
Human approval / guardrails
CrewAI released version 1.15.3, introducing a generic hook dispatcher with new step and execution-boundary interception points. The update also adds per-call usage metrics to kickoff results, resolves hook divergence bugs, and changes tool-result caching to an explicit opt-in feature.
Engineering teams can now pause, inspect, modify, and meter agent actions mid-execution instead of relying on post-run dashboards. This provides the granular control necessary to implement runtime approval gates and data sanitization policies in regulated environments.
Agent capability
Coupa released version 45.5 of its Coupa Supplier Portal, introducing AI-powered profile prefilling to automate data entry for newly registering suppliers. The update also enhances the CSP Invoices API by adding OAuth2 Bearer token authentication and requiring validation of a supplier's AI credit balance before processing programmatic invoice creations.
Procurement teams can achieve higher supplier adoption rates by removing friction during the initial onboarding process. Furthermore, the new API controls give administrators the ability to secure automated billing endpoints while strictly managing the consumption of AI credits across their vendor network.
Workflow orchestration
Arcade launched SkillBench, an evaluation index that grades the quality of AI agent skills. The platform currently scores approximately 39,000 skills across six dimensions, including safety and tool boundaries. It assigns each skill a letter grade from A to F to help developers assess reliability.
Enterprise engineering teams can use this tool to vet third-party agent skills before deployment. It provides a transparent framework to ensure that integrated workflows meet production safety and quality standards. This reduces the risk of insecure or poorly defined agent actions.
Security / enterprise
Zed shipped versions 1.12.0 (preview) and 1.11.3 (stable), introducing sandboxing for the agent's terminal and fetch tools. The releases enable ACP (Agentic Communication Protocol) elicitations by default to gather structured user input, add support for searching inside terminal threads, and expand AI model support with GPT-5.6 via Amazon Bedrock.
Adding sandboxing for terminal and fetch tools significantly reduces security risks during autonomous agent execution, making the editor safer for enterprise deployment. Broadened integrations via Bedrock and structured input for custom ACP agents give engineering teams more flexibility to govern and deploy AI assistants.
MCP / tool calling / API
Sprinklr launched its Summer '26 Release, introducing Voice AI agents with sub-second response times and agentic AI for autonomous issue resolution. The update includes built-in testing, simulation, and quality scoring tools to validate agent behavior prior to deployment, alongside a beta Model Context Protocol connector.
Contact center leaders can deploy highly responsive voice agents that work across digital channels to ensure continuous customer support. Operations teams gain essential pre-deployment validation tools to safely test and assess agent behavior in a controlled environment before releasing them to live queues.
Agent capability
Rime launched Coda, a new dual-decoder text-to-speech model built for real-time enterprise conversations, announced alongside a Series A funding round. The release also introduces a native integration with Together AI that co-locates STT, LLM, and TTS to streamline voice agent pipelines.
The Coda model offers sub-100ms latency to help contact centers deploy highly responsive voice agents that reduce caller hang-ups. Furthermore, the Together AI integration eliminates multi-vendor network hops to achieve sub-700ms end-to-end latency.
Integrations
Parloa's AI Agent Management Platform achieved SAP Endorsed App premium certification and is now available for purchase on the SAP Store. The platform was tested for security and data quality to connect conversational AI agents directly into SAP Service Cloud.
Organizations using SAP Service Cloud can procure Parloa's agents directly through SAP to handle customer interactions within their existing service workflows. The endorsed status ensures the integration meets strict interoperability standards and reduces friction for IT and procurement teams.
Workflow orchestration
OpenCode version 1.18.2 introduces default execution constraints that prevent subagents from independently launching nested subagents. Administrators and developers can override this safeguard by configuring a specific subagent_depth limit when deeper recursive delegation is necessary.
This safeguard provides critical protection against runaway recursive agent loops and unexpected API costs. Enterprise buyers deploying complex autonomous workflows can ensure strict bounds on how far the system can delegate tasks without manual intervention.
Workflow orchestration
Mastra released version 1.51.0 of its core framework and published the Mastra Code SDK as a public package. The SDK enables third-party developers to build custom user interfaces and surfaces on top of the Mastra Code coding agent. Core framework updates introduced scoped AgentController sessions for running parallel agent workflows over a single resource, file-system-routed configurations, and multi-tenant scoping for evaluators.
Opening the Code SDK allows engineering teams to natively embed autonomous coding capabilities into their own internal developer portals or custom deployment pipelines. Additionally, the ability to run parallel sessions against a shared resource unlocks higher throughput and concurrency for enterprises deploying large-scale agent applications.
MCP / tool calling / API
Lovable introduced agent integrations, allowing published applications to be utilized directly within external AI tools like ChatGPT and Claude. Developers can expose specific application actions via the Model Context Protocol (MCP), turning their app's functionality into callable tools for external agents.
This expands how end-users can interact with built applications by moving beyond a standard browser interface to natural language workflows within their preferred AI assistants. It provides a new distribution and engagement channel for applications built on the platform.
Observability / auditability
Langfuse updated its observations table to default to a focused view of application entry points for projects using the latest Python and JS/TS SDKs. The new automatic root filter temporarily hides nested tool and retrieval calls to surface top-level requests first.
This change reduces noise when initially investigating agent logs and helps users quickly identify critical inputs and outputs. It streamlines the debugging process for nested applications by establishing a clear starting point for drill-down analysis.
Funding / partnership
Anaconda acquired Kilo Code to integrate its open-source agentic engineering capabilities directly into the Anaconda Platform. Kilo Code will remain free and open-source for individual builders, while Anaconda focuses on adding enterprise-grade governance, intelligent agent routing, and centralized monitoring for organizational users.
Enterprise buyers can deploy Kilo Code within Anaconda's secure infrastructure to gain centralized visibility over AI development activity and token costs. Existing users and independent developers face no workflow disruption and retain full model-agnostic flexibility.
Agent capability
Eightfold AI introduced Candidate Agent, a conversational AI agent that manages candidate engagement from discovery through application and interview handoff. The announcement also included Avatar, a digital-human persona for interviews, and 360 Interview, which consolidates multiple interview types into a continuous AI-led session.
Talent acquisition teams can use these multi-agent capabilities to handle high-volume applicant inquiries 24/7 across channels like SMS, WhatsApp, and the web. This automation accelerates the hiring cycle and reduces administrative overhead, allowing recruiters to focus strictly on human-led final evaluations.
MCP / tool calling / API
Creatio launched version 10x Unlimited, a major platform update that integrates human-led and AI agentic workflows into a single environment. The release features AI Twin, a conversational interface enabling business users to build personal AI agents without coding skills. It also introduces Creatio AI Studio for centralized agent lifecycle management and an AI app development toolkit that connects with external coding agents like Claude Code and GitHub Copilot via MCP tools.
Enterprises can consolidate their CRM and AI automation efforts into a unified platform, reducing reliance on fragmented point solutions. The democratization of agent creation empowers business users to automate routine tasks, while centralized governance ensures IT retains control over policies, access, and overall AI spending.
Workflow orchestration
n8n released version 2.31.0, featuring a comprehensive overhaul of its Notion integration. The updated node migrates to a newer Notion API, replacing legacy database queries with Data Sources, and introduces native support for markdown operations, JSON blocks, and database file downloads.
Organizations using Notion as a knowledge repository for their AI agents can now build more reliable workflows. The addition of full markdown and JSON block support allows agents to read, write, and reorder complex page structures natively, reducing friction for documentation use cases.
MCP / tool calling / API
Waniwani released its open-source SDK and command-line interface (CLI) to help developers build and manage Model Context Protocol (MCP) servers. The CLI allows teams to wire local repositories to a Waniwani agent and test against a hosted playground, while the SDK introduces server-side state persistence to avoid serializing data through the model on every turn.
Engineering teams can now accelerate agent development by using native command-line tools for local testing and deployment. The SDK's session bridging optimizes performance by managing state server-side, reducing the context payload and associated costs during model interactions.
Integrations
TidalWave launched a bidirectional integration with the Total Expert CRM platform to automatically sync borrower contacts and application statuses. The integration passes real-time application activity from TidalWave into Total Expert as actionable Insights, and allows loan officers to initiate applications directly from their CRM.
Lenders utilizing Total Expert can eliminate duplicate data entry and manual record reconciliation between their CRM and point-of-sale systems. This connectivity empowers mortgage teams to trigger automated follow-up marketing journeys based on a borrower's real-time application progress without switching platforms.
MCP / tool calling / API
SnapLogic announced the general availability of SnapCode and the SnapLogic MCP Server to extend its platform to AI coding environments. SnapCode enables developers to generate integration pipelines via natural language in tools like Claude Code. The SnapLogic MCP Server acts as a headless runtime that exposes platform operations to external AI agents via secure, Model Context Protocol-compatible tool calls. The release also adds an automated self-correction toggle for downstream tool failures in AgentCreator.
Engineering teams can connect AI coding agents to enterprise systems without bypassing centralized security and governance controls. By leveraging governed MCP tool calls instead of raw API connections, IT can maintain a fully auditable integration layer while developers speed up application delivery.
Workflow orchestration
OpenCode released version 1.18.0, finalizing the Desktop v2 migration. The update introduces a redesigned review panel with persistent file tabs, a new composer menu for adding context without losing draft text, and per-prompt model selection. A transition setting is included to allow users to temporarily toggle between the legacy and new interfaces.
A stabilized native desktop client with dedicated review tools improves developer workflows and task oversight. The addition of per-prompt model switching allows teams to optimize costs by reserving expensive reasoning models for only the specific steps that require them.
Security / enterprise
Ocean has opened direct customer access to Ray, its central autonomous AI investigation engine. While Ray previously operated strictly in the background to automatically triage and remediate emails, security teams can now query the AI directly to run on-demand investigations, reconstruct attacks, and trace threat actors.
This update transforms Ocean from a purely autonomous background tool into an interactive investigative assistant. Security analysts can now leverage the platform's pre-gathered context and machine-speed reasoning to accelerate manual incident response and answer stakeholder questions about specific threats.
Workflow orchestration
Moveworks published its mid-year product updates, highlighting a speed boost for Enterprise Search that includes improved accuracy and multilingual support. The release also showcases new access governance workflows in Agent Studio, enabling automated access approvals and reviews across systems like SailPoint and ServiceNow.
Buyers benefit from faster and more reliable multilingual search capabilities and the ability to automate complex identity and access management directly through the AI assistant. This functionality reduces manual intervention for IT teams and accelerates access provisioning across the enterprise.
Browser/computer use
Make announced the deprecation of its native Google Chrome app and browser extension, effective August 31, 2026. The retirement is driven by the end of support for Google's Manifest V2. After this date, the extension will be removed from the Chrome Web Store, scenarios relying on the app will stop working, and Make will not offer a replacement module.
Teams relying on Make's Chrome extension for browser-based automation workflows must audit and migrate their processes before the deadline. Buyers evaluating Make for browser automation should be aware that native Chrome extension support is being permanently sunset with no direct replacement module provided.
Integrations
Kiro has introduced three tiers of OpenAI's GPT-5.6 models: Sol, Terra, and Luna, across its IDE, CLI, and Web platforms. Sol acts as the flagship model for complex multi-step tasks, Terra serves as the balanced option for routine development, and Luna provides a fast, low-cost option for high-frequency workflows.
Engineering teams now have the flexibility to choose OpenAI's newest models alongside existing Anthropic options. This allows users to directly control the tradeoff between advanced reasoning capability and cost efficiency depending on the complexity of their development tasks.
Agent capability
HubSpot transitioned its Prospecting Agent out of gated early access, making the AI sales development agent generally available for all paid portals. The release introduces new buyer intent signals, custom intent configuration using plain language, outcome-based pricing metered via credits, and a native integration with Seamless.ai for net-new contact sourcing.
Bundling an autonomous SDR agent into the core CRM lowers the barrier to entry for AI-driven outbound sales. Sales operators will shift focus from manual list building to refining targeting guardrails, as they are now billed per sourced lead rather than via a flat software license.
MCP / tool calling / API
GPT Researcher introduced Deep Research, an advanced recursive workflow for exploring topics with agentic depth. The update also moves the platform's Model Context Protocol (MCP) server to a dedicated repository (gptr-mcp). Additionally, it adds multi-agent assistants built with LangGraph and AG2 and an enhanced frontend for real-time progress tracking.
These additions expand the platform's versatility for developers building autonomous AI applications. The dedicated MCP server allows assistants like Claude to natively trigger deep research, while the multi-agent frameworks provide more reliable orchestration. The new frontend also improves the user experience for teams tracking complex tasks.
Agent capability
Findem introduced Simple Shortlist, a redesigned candidate shortlist experience built to help recruiters move candidates into active campaigns faster. The update transforms the shortlist from a static holding area into an actionable interface that clarifies outreach statuses and next steps.
Talent acquisition teams managing multiple roles and high-volume campaigns can operate more efficiently without manually investigating individual candidate records. This enhancement reduces operational friction and accelerates the overall candidate engagement lifecycle.
Security / enterprise
Factory updated its Terms and Conditions for Individual Plans, detailing the licensing agreement for its AI coding platform and development tools. The revised terms explicitly state that customers grant Factory a non-exclusive license to access their connected codebases across platforms like GitHub, GitLab, and Bitbucket.
Legal and procurement teams must evaluate these updated data licensing clauses during procurement. Granting a direct codebase license to the vendor often requires strict review to ensure it does not conflict with internal intellectual property protection or data privacy policies.
Integrations
Exa introduced Agent Skills, a set of pre-built workflows formatted as portable SKILL.md files that integrate directly with compatible coding agents like Claude Code, Cursor, and Codex. This open-source repository allows developers to natively execute Exa's web search, content extraction, list building, and deep research capabilities without writing custom orchestration code.
Engineering teams evaluating Exa can now embed its semantic retrieval functions directly into their existing IDEs and agentic tools. This standardization accelerates prototyping and reduces the effort required to deploy complex search operations across different agent frameworks.
Workflow orchestration
Airia introduced Model Change Management, a capability designed to protect enterprises from AI model deprecations. The tool provides alerts up to 90 days before a model is retired, offers centralized visibility into all affected agents, and includes bulk migration features to update multiple agents simultaneously.
Enterprise teams can proactively prevent broken workflows and compliance blind spots caused by model retirements. The bulk migration features eliminate the need for manual updates, ensuring continuous AI operations and reducing the risk of silent agent failures.
Security / enterprise
Ada released the Bulk End-User Deletion API to programmatically erase personal data associated with specific users across its systems. The API accepts up to 1000 identifiers per request via email, external ID, or stored variables, automatically propagating the erasure downstream across all data stores.
This update allows enterprise compliance and security teams to fully automate data privacy operations in accordance with GDPR and CCPA requirements. Programmatic control removes the manual overhead of processing deletion requests and ensures comprehensive data removal across all connected support records.
Integrations
Netlify partnered with Anthropic to embed deployment capabilities directly inside Claude Design. Users can deploy generated designs to a live Netlify site with a single click and push ongoing iterations without manual exports.
This integration collapses the gap between prototyping and deployment, accelerating the release cycle for AI generated applications. It enables rapid iteration by keeping the design tool and the live environment synchronized automatically.
Observability / auditability
Langfuse rebuilt its trace graph view on a deterministic renderer to include two layout modes: Aggregated and Expanded. The Aggregated mode groups repeated steps into single nodes with counters to show an overall shape, while the Expanded mode unrolls all loops into an execution-ordered directed acyclic graph.
This update makes auditing complex agent executions much easier. Evaluators and developers can quickly switch between a high-level structural overview and a granular step-by-step breakdown without losing their place on large traces.
Integrations
HubSpot introduced a native integration for ChatGPT Ads, allowing marketing teams to connect their ChatGPT advertising accounts directly into the CRM. Users can now build ad campaigns, monitor reporting, track attribution, and measure ROI on OpenAI's platform alongside their traditional search and social ad networks.
As AI search engines increasingly capture buyer discovery traffic, managing ad spend on platforms like ChatGPT becomes a necessity. This integration gives revenue teams a unified interface to treat AI search advertising as a standard, measurable channel without needing external data handoffs.
Deployment / data residency
HubSpot began automatically migrating all existing custom assistants into Breeze projects, converting the legacy standalone assistants to read-only mode. Teams must now build and manage their AI assistants and agents centrally within Breeze Studio, which involves manually porting welcome messages and conversation starters to the new project instructions.
Centralizing AI builds into Breeze Studio forces organizations to adopt a more standardized governance model. Administrators gain better control over instructions and data access, reducing drift across scattered ad hoc bots and simplifying change management for brand or legal compliance.
Agent capability
Botify updated its SpeedWorkers edge network solution to allow enterprise sites to serve content directly in Markdown format to selected AI bots. Administrators can specify which AI agents receive Markdown instead of HTML and use a built-in preview tool to review the bot's reading experience.
Markdown is significantly more efficient for large language models to process compared to raw HTML, which contains extensive structural code. This capability helps marketing teams ensure their critical content is quickly parsed and accurately represented in generative AI search responses.
Security / enterprise
Cybersecurity firm Wiz publicly disclosed "GhostApproval," a technique allowing malicious code repositories to trick AI coding assistants into modifying files outside of their intended workspaces using symbolic links. While several competitors issued patches, Augment Code disputed the vulnerability classification, stating that this behavior aligns with its product design since agents inherently operate under user credentials. The company declined to issue a patch or implement an isolation fix for symlink abuse.
Security and engineering leaders evaluating Augment Code must note that the agent can read and write files outside its immediate workspace if exposed to malicious symlinks. Organizations will need to enforce strict endpoint controls and independently vet third-party code before bringing it into privileged environments, as the vendor considers sandboxing the file system a shared responsibility.
Agent capability
Cresta launched Training Simulator, an agentic training product that replaces scripted role-play with live, adaptive simulations. The product generates simulated customers based on a company's actual historical conversation data, allowing human agents to practice handling complex interactions like retention saves and upselling.
Contact center leaders can use this tool to reduce agent onboarding time and improve readiness before agents interact with real customers. By grading simulated scenarios with the same artificial intelligence quality criteria used on live calls, buyers gain a unified system for training, quality management, and performance measurement.
MCP / tool calling / API
ServiceNow released version 1.0.0 of its Advanced Approval Management AI, which includes Model Context Protocol (MCP) tools. This integration allows agents and approvers to submit, review, and act on quote approval requests directly from external MCP clients, such as Anthropic's Claude desktop app.
Surfacing approval workflows directly within an enterprise's chosen conversational AI assistant removes the need for users to switch context to the ServiceNow platform. This reduces friction for managers and accelerates decision-making by placing administrative approvals in their natural flow of work.
Agent capability
ServiceNow launched the NOW Quote AI Agent within the Quote Management application. The agent automates B2B sales quoting by reading opportunity notes, automatically configuring products, resolving non-standard names, applying discounts, and generating ready-to-review quotes with full audit trails.
Sales operations teams frequently face compliance risks and administrative bottlenecks when handling manual pricing and catalog configurations. This automation enables reps to bypass manual quoting entirely, accelerating deal velocity while ensuring price list accuracy.
Agent capability
ServiceNow integrated Agentic AI into its Telecommunications Network Inventory (TNI) application via version 2.0.2. The release introduces an autonomous infrastructure-allocation agent that processes free-text change requests and converts them into validated, policy-aware network allocations.
Telecom operators deal with complex, manual infrastructure provisioning that requires strict adherence to network policies. This update ensures that operators can provision resources faster with fewer errors, offering full traceability directly within the Network Inventory workspace.
Agent capability
Swimlane introduced Turbine Widgets, enabling security teams to build custom web components that extend the Turbine UI on case records and dashboards. Users can now utilize the 'Build with Hero AI' feature to automatically generate widget code from plain-English descriptions.
This enhancement allows SOC teams to rapidly create tailored, specialized views of their security data without requiring advanced front-end development skills. By leveraging generative AI to build custom interfaces, organizations can streamline analyst workflows and present exact contextual data for faster incident response.
Workflow orchestration
Thread announced the general availability of Super Magic, an AI assistant that automates task execution for managed service providers. Evolving from the read-only Ask Magic feature, it chains multi-step processes like reading tickets, pulling client documentation, and queuing actions in the PSA. The capability includes out-of-the-box connectors for external tools such as Notion, Zapier, Liongard, and Linear.
Service desk leaders can significantly reduce the time technicians spend toggling between disparate systems and knowledge bases. Because the system requires a one-tap confirmation before writing any data, teams can accelerate ticket resolution safely without risking unauthorized changes to client environments.
Funding / partnership
Infobip has acquired SocketLabs, a US-based email infrastructure provider known for high-volume email routing. The transaction integrates SocketLabs' vendor-agnostic observability and intelligent routing analytics into Infobip's platform to enhance its AI-powered Email Deliverability Agent.
Enterprises and ISVs gain a neutral, cross-provider view of deliverability metrics directly within the Infobip platform. This enables organizations to intelligently route email traffic across multiple senders and optimize inbox placement without relying on single-provider tools.
Integrations
TrueFoundry has integrated TrojAI into its AI Gateway to enable native real-time traffic enforcement and monitoring. Users can configure the gateway to route traffic through TrojAI as either a proxy layer or a guardrail service to detect prompt injections, data leakage, and unwanted AI content.
This allows security and AI engineering teams to deploy multi-model generative AI applications faster without building custom defense infrastructure. By layering TrojAI threat detection directly onto the TrueFoundry gateway, organizations can block malicious prompts in production traffic natively.
Agent capability
Cartesia released Ink-2, a streaming speech-to-text (STT) model optimized for real-time voice agents. The model features native turn detection that emits events when a user starts, wraps up, and finishes speaking, eliminating the need for external Voice Activity Detection modules.
Voice agent builders can simplify their application architecture and lower response latency by relying on Ink-2's built-in turn detection. The model improves transcription precision across multiple accents and noisy conditions, ensuring reliable customer-facing interactions without unintended pauses.
Agent capability
Box brought its AI Home and Box Agent capabilities to its native iOS and Android mobile applications. The release features persistent sessions, allowing users to initiate agentic queries on a desktop and seamlessly continue them on a mobile device while securely accessing organizational files.
Mobile access ensures that field workers and traveling executives can leverage Box AI without needing a laptop. The cross-device persistent state reduces friction for enterprise workflows that happen outside traditional office environments.
Agent capability
Warp expanded its payroll product into an AI-native employee management platform. The update introduces autonomous AI agents that open state tax accounts, file payroll forms, and resolve tax notices automatically. It also includes Warp Fabric for automated IT application and device provisioning.
Founders and operations teams can consolidate payroll, IT provisioning, and benefits into a single system. The addition of autonomous agents removes the operational overhead of manually tracking multi-state tax regulations and completing cross-system onboarding steps.
MCP / tool calling / API
SAP released a Joule Studio extension for Visual Studio Code and a new command-line interface (CLI) to support professional developers and DevOps teams. The updates allow developers to access AI-powered project scaffolding, code generation, and CI/CD orchestration directly from their local environment or terminal.
This expansion shifts Joule Studio's capabilities beyond low-code users, enabling pro-code engineering teams to build, manage, and deploy custom enterprise AI agents using standard CI/CD pipelines and familiar IDEs.
Workflow orchestration
Taskade launched TSK-1, a system kernel serving as the primary intelligence layer for all applications built using Taskade Genesis. The kernel coordinates AI models, persistent memory, multi-agent reasoning, and automated workflows within a single workspace, automatically routing tasks across more than 15 different models.
Buyers can deploy AI applications that maintain state and execute tasks continuously rather than functioning as static generated artifacts. The built-in kernel removes the need for teams to manually integrate separate databases, orchestration frameworks, and model routing logic.
Integrations
LangChain and NVIDIA launched the NemoClaw for LangChain Deep Agents blueprint, a reference architecture for enterprise agent systems. The stack combines LangChain Deep Agents Code, NVIDIA's Nemotron 3 Ultra model, and the NVIDIA OpenShell runtime to optimize agent performance and inference efficiency.
Enterprises can use this blueprint to deploy top-performing open agent systems at a significantly lower inference cost compared to closed models. The open architecture also ensures engineering teams retain complete control over their proprietary data, workflows, and deployment environments.
Pricing / packaging
UiPath updated the AI Unit consumption model for Document Understanding, charging per operation rather than a flat rate per page. Under the new model, digitization is free while classification and extraction are billed individually.
Organizations using Document Understanding can better optimize their costs, especially for workflows requiring only partial document processing. Customers can also combine Intelligent Xtraction and Processing licenses with Document Understanding without incurring additional per page costs.
Human approval / guardrails
Quiq introduced Verified Intelligence, a new control layer for its agentic AI platform. The suite adds built-in guardrails, pre-deployment simulation testing, and step-by-step decision visibility to all AI agent deployments on the platform. It includes specific features like Verify Claim, which automatically cross-references AI responses against approved company knowledge prior to delivery.
Customer experience teams can now deploy autonomous agents with stricter governance and significantly reduced brand risk. The combination of multi-turn simulation capabilities and granular audit trails provides enterprise leaders with the necessary oversight to scale AI workflows confidently beyond isolated pilots.
MCP / tool calling / API
UiPath launched a public preview of UiPath for Coding Agents, enabling developers to build automations directly from AI coding agents like Claude Code, Cursor, and Codex CLI. This update introduces a new command-line interface (uip CLI) and UiPath skills that allow coding agents to interact with a user's UiPath organization using natural language. It also introduces a Coding Agents for Test feature that allows these agents to autonomously read test cases, write execution results, and update requirement coverage data.
Engineering and QA teams can now leverage popular AI coding assistants to design, orchestrate, and test enterprise automations without relying solely on a separate visual designer. This allows developers to integrate UiPath automation and testing management natively into their existing AI-assisted development environments.
Security / enterprise
Element451 introduced a dedicated View Payments permission to restrict access to payment information on student profiles. The release also includes underlying security enhancements for accessing this payment data via the API.
Higher education institutions gain more precise control over who can view sensitive financial data. Administrators can implement this restriction seamlessly since existing authorized users automatically retain their access during the transition.
Pricing / packaging
Box updated its documentation outlining how AI Units operate as a standardized metric for chargeable AI features. The update clarifies that units are consumed by at-scale or automated tasks, including Box Automate workflows and the new Expanded Mode Agents used for processing massive datasets.
Enterprise administrators can now predictably forecast and manage their AI budgets by understanding which specific agentic actions consume paid units. This level of transparency is essential as Box enforces daily AI usage limits starting in July 2026.
Integrations
Gallabox introduced platform support for WhatsApp's new username feature and Business-Scoped User IDs (BSUIDs), reducing the historic reliance purely on phone numbers for customer identity. The update adds updated contact-matching logic to prevent duplicate customer records when transitioning from a single identifier to two. Additionally, agents and bots on the Gallabox platform have been updated to recognize and correctly display username-based contacts.
With WhatsApp introducing username capabilities and masking phone numbers via BSUIDs, businesses risk breaking integrations or orphaning CRM records if their systems expect phone numbers as the primary key. This update ensures that customer service teams and automated workflows continue to function correctly as users hide their numbers.
Agent capability
Hyground released Omni, a major 2.0 rearchitecture that transitions the platform from a single-cluster agent model to a unified, multi-cluster AI agent. The update introduces a real-time topology map that allows the agent to reason across up to 20,000 nodes and supports arbitrary workloads such as ECS, virtual machines, and serverless functions.
Platform engineering teams can now manage their entire infrastructure through a centralized AI agent rather than deploying one per cluster, simplifying operations and reducing token costs. The addition of dynamic guardrail agents provides stricter security by validating inputs and verifying actions before they are executed in production.
Human approval / guardrails
Tabnine released version 6.4.0 of its platform, introducing independent agent and chat sessions for developers running multiple IDE instances simultaneously to prevent state conflicts. The update also adds a fast model routing setting for background CLI tasks and a new administrative control to restrict YOLO mode, preventing the automatic execution of CLI tool calls.
Enterprise security teams can now mandate human-in-the-loop approval for autonomous agent actions, mitigating the risk of unauthorized or destructive CLI command execution. Developers working across multiple projects simultaneously will also experience a more reliable coding assistant with independent session states.
Security / enterprise
Cybersecurity researchers at Sand Security disclosed WriteOut, a critical session isolation vulnerability in the managed sandbox of the Writer AI platform. The flaw allowed attackers to exploit the agent preview feature to leak session tokens and gain cross-tenant access. Writer promptly deployed a server-side patch that removes session credentials from sandbox previews and isolates the preview origin.
Security-conscious buyers must account for the platform's shared responsibility model and incident response capabilities. The patched vulnerability underscores the importance of session isolation in AI agent deployments, though Writer confirmed no customer data was compromised.
Agent capability
Pydantic AI released version 2.6.0, adding file support to its CodeExecutionTool for Anthropic and OpenAI models. The update also introduces time-to-first-token tracking for streaming model requests and adds Amazon Bedrock model profiles for Writer, Z.AI, and Moonshot AI.
Teams building agents can now pass files directly into sandboxed code execution environments, expanding the complexity of tasks agents can handle. The new time-to-first-token metric improves observability for streaming applications, and the expanded Bedrock support offers more model routing options.
Agent capability
UiPath upgraded Autopilot in Studio Desktop to function as a fully autonomous coding agent rather than a drafting tool. The agent can now plan, build, run, debug, and refactor automations directly within the Studio environment using official UiPath skills and command line interface tooling.
Developers can streamline their automation building lifecycle without needing separate subscriptions or context switching. By acting as a native coding assistant with full project context, Autopilot reduces development time and simplifies troubleshooting tasks like repairing broken selectors.
Agent capability
Capsule Security released Guardian Agent, an in-product AI companion built on its pi-agent-core client stack. The agent runs its entire loop locally within the browser, retaining direct access to live application state and driving UI changes in real time. It uses a three-layer bottom-up context assembly mechanism to explore GraphQL schemas on demand, while a hardened backend proxy enforces centralized governance, authentication, and auditing.
This release provides an interactive, context-aware AI assistant directly within the Capsule platform to help investigate and manage agent runtime risks. The hybrid local-execution architecture demonstrates how enterprises can deploy UI-driving agents without exposing systems to unchecked prompt injection or unmonitored actions.
Security / enterprise
ThreatModeler updated its Intelligent Threat Engine (ITE) to detect advanced AI-specific threats at the architectural design stage. The release introduces curated threat libraries that map agentic tool compromise and Model Context Protocol (MCP) vulnerabilities to industry frameworks such as MITRE ATLAS and OWASP Agentic AI Threats.
Security and DevSecOps teams can proactively model risks associated with autonomous agents before they are deployed. By visualizing trust boundaries where local MCP clients connect to remote LLMs, architects can prevent critical vulnerabilities like unauthorized database access.
Agent capability
TrueFoundry introduced its Agent Gateway and Agent Harness, providing a dedicated control layer and runtime environment for AI agents in production. The Agent Gateway acts as a data plane for stateful sessions and multi-step executions, while the Agent Harness manages orchestration and context.
Enterprise buyers can now govern agentic traffic, tool calls, and state management across multiple platforms without rebuilding their core applications. This centralizes control over permissions, approvals, and observability, making autonomous agents safer for production deployment.
Integrations
Levelpath launched a native integration with OneTrust to embed third-party risk and compliance assessments directly into procurement workflows. The connector allows teams to trigger, track, and complete assessments within the platform, automatically syncing status updates and completing workflow steps upon assessment approval.
This integration removes the need to switch between procurement and compliance tools, accelerating the purchasing lifecycle. Buyers can enforce governance policies effectively by making risk assessments an automated and auditable step within the broader orchestration process.
Agent capability
Assail released Sidewinder, a complete version 2 redesign of its Ares offensive security platform powered by a new 31 billion-parameter model. The platform shifts to a plan-driven architecture featuring 12 specialized autonomous agents that operate against a persistent knowledge graph. It also introduces vision-grounded analysis to drive real browsers for crawling single-page applications.
Security teams gain a continuous, autonomous penetration testing system capable of auditing its own work and repairing mistakes without human intervention. The ability to handle complex browser interactions and multiple authenticated identities allows organizations to uncover deep authorization flaws that traditional scanners typically miss.
Security / enterprise
Pluvo updated its Terms of Service to explicitly prohibit the use of customer data and model outputs for training or fine-tuning AI models. The revised terms clarify that customers retain full ownership of their data and the generated analysis. Pluvo also formalized its policies around the use of third-party AI sub-processors and cloud infrastructure.
Finance teams dealing with highly sensitive company data require strict privacy guarantees before deploying agentic workflows. By contractually ensuring that data will not bleed into public or aggregate foundation models, the vendor removes a major legal blocker for enterprise adoption.
MCP / tool calling / API
Ramp introduced Ramp for Agents, a feature that enables users to incorporate a business, apply for corporate cards, and establish a finance stack through a single text prompt. This release includes a Model Context Protocol (MCP) integration and a CLI toolset with over 50 finance playbooks, allowing AI agents to perform tasks related to banking, bill pay, travel, expenses, and accounting.
Founders can automate administrative setup tasks without leaving their development environments. The provided MCP toolset allows engineering teams to programmatically embed financial operations and autonomous spending controls directly into their custom internal workflows.
Security / enterprise
Element451 has retired its live ICS calendar feed export for Appointments to improve platform security. Users must now utilize the native calendar integrations for Google Calendar and Microsoft Outlook for ongoing synchronization, although downloadable calendar attachments in confirmation emails remain available.
Administrators need to transition their teams to native integrations to maintain live visibility into booked appointments. The removal of the legacy ICS feed helps institutions protect schedule data from unauthorized access.
Agent capability
Unikraft introduced Sandboxes on Unikraft Cloud, providing on-demand, hardware-isolated microVMs for executing short-lived tasks and AI agent tool calls. The environments boot in milliseconds, feature automatic scale-to-zero, and charge users solely for active execution time.
Teams building autonomous agents can safely detonate untrusted code or run third-party plugins without managing complex underlying virtual machine infrastructure. The microVM approach removes the traditional trade-off between the fast startup times of containers and the strict isolation boundaries required for secure execution.
Security / enterprise
Fabrix.ai introduced 'vX' (Vibe Coded User Xperience), an enterprise-grade execution foundation for AI-generated applications. The update delivers built-in single sign-on (SSO), high availability, geographic disaster recovery, full observability, and token spend optimization for deployed agents.
Organizations can safely deploy custom-built AI applications knowing they instantly inherit required security, scalability, and governance controls. This capability removes the operational burden of manually integrating enterprise identity and monitoring into standalone agentic tools.
Agent capability
Itential announced the general availability of FlowAI, an agentic harness for building and deploying AI agents alongside deterministic infrastructure automation. Released with Itential Platform 6.5, the update introduces a Work Center for centralized human-in-the-loop approvals, Itential Gateway 5.5 for dynamic external credential retrieval, and Itential Builder Skills for Claude Code to enable spec-driven development via delivery agents.
Infrastructure teams can securely deploy AI agents into production by maintaining their existing operational policies, role-based access controls, and audit trails. The integration of deterministic execution with AI reasoning allows organizations to accelerate automation while enforcing strict governance and compliance standards.
Security / enterprise
Stacklok added support for Enterprise-Managed Authorization (EMA) to ToolHive, its open-source Model Context Protocol (MCP) project. This allows identity providers to centrally grant and manage server access, ensuring every tool call carries the real identity of the user who initiated the action.
This update allows enterprise security teams to govern AI agent access via existing identity providers like Okta or Microsoft Entra ID. By tying every agent action to a specific named user rather than a shared service account, organizations can maintain strict least-privilege access controls and meet critical compliance requirements.
Agent capability
Spring Labs launched Zanko version 0.27.0, featuring dedicated workspaces for managing fraud and regulatory complaints with integrated case management. The release introduces the 'Ask Zanko Report Agent' to generate statistical executive reports from plain-English prompts, as well as native email capabilities for reviewing AI-drafted responses directly in an inbox. Administrators also gain granular custom roles and permissions along with expanded QA analytics and workflow automation.
Compliance teams at financial institutions can reduce manual reporting overhead and spreadsheet dependencies by utilizing the new conversational reporting agent. Furthermore, the ability to approve drafted communications within existing email clients minimizes context switching, while the enhanced access controls help enforce strict data governance policies.
Workflow orchestration
UiPath released its July 2026 updates for Automation Cloud and Test Cloud, introducing a preview of Test with Coding Agents for managing test data via AI assistants. The update also brings Azure Private Link support for Test Cloud in the US, new public API endpoints for video recordings, and bulk-edit capabilities for queue items in Orchestrator.
The addition of Azure Private Link enhances network security for US organizations with strict compliance requirements for their testing environments. Furthermore, the Orchestrator interface improvements and testing APIs streamline administrative overhead for managing complex automation queues.
Agent capability
Unit21 launched the Agentic Task Builder, a no-code capability that enables fraud and compliance teams to create custom AI investigation tasks using plain English. Currently supporting online search and data analysis, the builder allows agents to autonomously search external data, analyze internal records, and produce structured answers with citations inside existing alert workflows.
This feature eliminates the bottleneck of relying on engineering resources or vendor updates to modify investigation logic. Risk officers can quickly deploy custom rules to address emerging financial crime tactics, ensuring their compliance operations remain highly responsive to new threats.
Security / enterprise
Daytona released the SecretService API for organization-scoped secrets in agent sandboxes. Secrets are injected as opaque placeholders in environment variables, and the real plaintext is substituted only at the network egress layer, and only for allowed hosts.
A secure way to hand credentials to agents: untrusted generated code running inside the sandbox cannot read or leak the underlying plaintext API keys, closing one of the most common exfiltration paths in agent execution.
MCP / tool calling / API
Kilo Code added dependency gating for experimental agents: the CLI blocks execution if declared skill, MCP, or VS Code extension prerequisites are missing, and the extension shows requirement groups with Marketplace shortcuts. ACP integrations moved to acp-next.
For complex agents, gating prevents mid-task failures by enforcing dependencies before a run starts. The expanded Agent Context Protocol support deepens custom tooling and telemetry integration.
Observability / auditability
GitHub opened a public preview of Copilot agent session streaming for Enterprise Cloud with managed users. Prompts, responses, and tool calls across all Copilot clients can stream to SIEM platforms like Microsoft Purview or be pulled via REST API.
As agents act more autonomously, security and compliance teams need granular visibility. Streaming session activity into existing SIEM tooling lets them audit agent behavior and satisfy governance requirements without new tooling.
Agent capability
Poolside released Laguna XS 2.1, a 33B-parameter Mixture-of-Experts model for agentic coding and local deployment, with improved SWE-bench Multilingual and terminal-task performance, multiple quantized checkpoints, and a move to the permissive OpenMDW-1.1 license.
Teams can self-host an open-weights agentic model with lower compute overhead, addressing data residency and compliance directly. Permissive licensing simplifies corporate adoption.
Agent capability
Agent Zero v2.2 added in-chat model setup, unified provider onboarding, and explicit utility-model selection, set Claude Sonnet 5 as default, added safer streaming fallbacks, improved tool parsing, and shipped security dependency upgrades for LiteLLM and Starlette.
Admins can configure and switch providers right in the chat interface, simplifying multi-model setup. Better fallbacks and patched dependencies make autonomous execution more reliable and secure.
Integrations
Make added a native Databricks app for Enterprise plans, letting users run SQL queries, manage jobs and pipelines, and upload files to Unity Catalog volumes directly from Make scenarios.
Enterprise data teams can connect their Databricks lakehouse to Make without custom API work, accelerating data orchestration and simplifying large-scale admin workflows.
Memory / state
Qodo 2.4 removed most of its repo-wide RAG indexing layer for code review, replacing it with an agent-driven fetch-on-demand loop plus long-term memory over PR history. The agent now rediscovers code context itself via git and grep, while retrieval is reserved for team-specific review decisions, repeated feedback, and conventions that cannot be recovered from the working tree. Qodo reports the leaner architecture leads on its review benchmarks with a fraction of the indexing infrastructure.
A concrete architectural signal of where code-review agents are heading: less index maintenance and reprocessing cost, more team-specific judgment carried across reviews. For buyers, PR-history memory means the reviewer applies this team's standards rather than generic best guesses, and the smaller retrieval footprint lowers total cost of ownership.
Agent capability
Copilot Vision reached general availability across Free, Pro, Business, and Enterprise. Users can attach images and PDFs to chat prompts so Copilot reasons about visual context alongside code. It is enabled by default with no admin policy change required.
Developers can feed visual bugs, architecture diagrams, and UI mockups straight to the agent instead of translating them to text. Available on every tier by default, so teams get multimodal help immediately.
Browser/computer use
Browser tools for Copilot agents in VS Code are generally available and on by default. Agents can open pages, navigate, click, type, capture console errors, and screenshot inside the editor, with consent controls for private tabs and admin domain toggles.
Lets Copilot verify UI changes and debug frontend issues without leaving the editor. Domain allowlists and explicit consent keep the agentic access governable in secure environments.
Agent capability
Kimi K2.7 Code, an open-weight model from Moonshot AI, is now generally available in the Copilot model picker, hosted on Azure and billed by usage. It is disabled by default for Business and Enterprise, requiring admin enablement.
A cheaper open-weight option for teams optimizing inference spend. Because it adds a new provider, enterprise admins opt in and govern access under their own compliance policies.
Pricing / packaging
AI credit session limits are now settable in the Copilot CLI (v1.0.66+) and SDK (v1.0.5+), letting developers and admins cap what an agent can spend in a single execution session.
Prevents runaway autonomous sessions from generating surprise overages under usage-based billing. Engineering leaders get hard budget control and cost predictability when shipping agentic features.
Security / enterprise
Tavily achieved ISO/IEC 27001:2022 certification, independently validating its security program, infrastructure, and internal controls against recognized information-security standards.
Enterprises require rigorous security frameworks before adopting an infrastructure provider for production AI. This certification unblocks procurement at organizations with strict vendor compliance requirements.
Browser/computer use
Skyvern's June changelog renames Workflows to Agents across the UI and public API, and adds multi-tab browser control, a Workflow Studio editor beta, native Google Cloud deployment, and stronger one-time-password handling.
The rename signals a focus on autonomous execution over rigid scripting. Multi-tab control and GCP self-hosting target enterprise teams automating complex, secure web portals that traditional RPA struggles with.
Agent capability
Eleven Agno releases (v2.6.10 to 2.6.20) added tool-batch checkpointing and run forking, full CRUD on agent learnings, a StudioTool for dynamic agent composition, ClickHouse trace backends, custom scoped MCP tools, and five new model providers including Cloudflare AI Gateway.
Teams can pause, fork, and resume long-running tasks without starting over, cutting compute cost on complex workflows. ClickHouse backends scale observability, and CRUD controls let teams manage and sanitize agent memory directly.
Security / enterprise
Cognition launched Devin Security Swarm, a multi-agent system on an Agentic MapReduce architecture that scans whole codebases to find, validate, and fix vulnerabilities, reproducing findings in isolated sandboxes and opening patch PRs.
Security teams can automate remediation of vulnerability backlogs with runtime-verified proof of exploitability. This cuts false-positive noise from traditional scanners and delivers actionable fixes at scale.
Security / enterprise
Factory released Droid Shield 2.0 with two fine-tuned models for secret detection during autonomous commits, improving accuracy on real credentials versus false alarms while cutting latency and cost. Model weights were released for security research.
Teams running autonomous coding agents need protection against credential leaks. Fewer false positives and fewer false negatives let security teams scale autonomous commits with more confidence.
Security / enterprise
monday.com added an AI permissions tab and an Agent directory for Enterprise admins. Permissions assign role-based access to specific AI features and agent types (monday and third-party); the directory centralizes monitoring, activation, and deactivation of every agent on the account.
Visibility and access control are critical when deploying agents at scale. Admins can govern who interacts with or deploys agents, reducing unauthorized use and excess credit consumption.
Agent capability
Juicebox introduced Research Signals, letting recruiters find and evaluate specialized talent by publications, patents, citations, and research metrics. Natural-language queries filter by citation count, H-index, recent publications, and specific conferences or institutions, with a Research section on profiles linking to papers and patents.
Evaluating deep technical talent usually means manually cross-referencing papers and patents. Aggregating research output and impact into profiles lets hiring teams identify active, high-impact researchers without leaving the platform.
Deployment / data residency
IBM expanded watsonx Orchestrate to the Paris (eu-fr2) region on IBM Cloud.
A localized deployment option for European buyers meeting data sovereignty, performance, and compliance requirements in France and the EU.
Deployment / data residency
NICE was named a launch partner for the AWS European Sovereign Cloud, making its agentic customer-experience platform natively available on infrastructure located entirely within the EU, keeping customer data and metadata within European borders.
Lets highly regulated buyers in finance, healthcare, and the public sector satisfy digital-sovereignty mandates while running AI agents, real-time copilots, and workflow automation.
Memory / state
Harvey shipped Org Context, which lets firms give the platform standing knowledge of their organization (practice structure, precedents, preferences) so outputs reflect the firm rather than a generic baseline, alongside Internal Spaces for firm-internal collaboration and the ability to attach documents directly to review tables.
Standing organizational context reduces the per-matter setup work of re-explaining firm conventions, and Internal Spaces keeps collaborative legal work inside the governed platform instead of scattering across email and shared drives.
MCP / tool calling / API
E2B deprecated E2B_ACCESS_TOKEN authentication in favor of E2B_API_KEY. This is a breaking change for existing integrations: legacy access tokens stop working on August 1, 2026, and all SDK and API calls must migrate to API-key auth before then.
A hard auth cutover with a fixed deadline is directly material to every team running E2B sandboxes in production. Integrations that are not migrated by August 1 will fail, so this belongs on any E2B customer's near-term maintenance list.
MCP / tool calling / API
ZoomInfo shipped an integration with v0 that lets generated apps read verified GTM data via ZoomInfo's GTM.AI Context Graph over an MCP endpoint, so v0 builds on live, identity-resolved company and contact data instead of static exports.
RevOps, sales, and marketing teams can build internal tools, scoring dashboards, and routing workflows in v0 without separate pipelines, while live data access reduces hallucination and stale results.
Browser/computer use
Browserbase launched Browserbase Agents: fully managed web agents invoked with a single API call. Instead of provisioning browser infrastructure and wiring an agent loop, developers describe the task and Browserbase runs the browsing, navigation, and extraction end to end on its hosted stack.
Collapsing browser infrastructure, agent orchestration, and task execution into one API call removes the largest integration burden in web automation. For teams evaluating browser agents, this shifts Browserbase from infrastructure provider to a turnkey agent option.
Agent capability
Harvey made Claude Sonnet 5 available across its platform, reporting 91.3% on BigLaw Bench, its benchmark of real legal tasks. The upgrade applies to Assistant, Vault review, and Harvey's agent workflows.
A measurable jump on a legal-specific benchmark, not just a model swap. For legal buyers, published task-level scores give a concrete basis to compare platform quality as the underlying models change.
Integrations
Shopify opened Sidekick App Extensions to all developers in its Spring '26 Edition, letting third-party apps expose data and functionality inside Sidekick in the Shopify admin. Merchants can now reach partner app capabilities through the same conversational agent they use for native Shopify tasks.
Extending the merchant agent to third-party apps turns Sidekick into a platform surface rather than a closed assistant. For merchants, more of the stack becomes reachable in one conversation; for app vendors, agent integration becomes a distribution channel.
Agent capability
Optimizely's June 30 Opal release introduced the Opal Agent Library with 45+ prebuilt marketing agents, a reporting dashboard for tracking agent activity and outcomes, and multi-model support for routing work across different LLMs.
A large prebuilt agent catalog shortens time to value for marketing teams that do not want to build agents from scratch, while the reporting dashboard gives leaders visibility into what agents are actually doing. Multi-model support reduces lock-in to a single LLM provider.
Agent capability
LlamaIndex announced Retrieval Harness, giving agents filesystem-style primitives (grep, file read, directory listing) over document collections so they can navigate and fetch context on demand rather than relying solely on a prebuilt vector index.
Aligns the retrieval stack with how modern agents actually work: fetching what they need step by step. Teams get agentic document access without maintaining heavy indexing infrastructure for cases where on-demand navigation is enough.
Agent capability
Sourcegraph launched Agentic Batch Changes in public beta: an agent harness for applying code changes across thousands of repositories or the largest monorepos. Built on Batch Changes and Deep Search, the agent plans a change, chooses between writing deterministic scripts or delegating per-repo work to Claude Code or Codex, executes across all targets, and reacts to downstream events like CI failures and merge conflicts until changes are ready to merge. Beta access is by request.
Large-scale migrations, security patching, and breaking-change upgrades are among the most expensive maintenance work in big codebases. Orchestrating a single coding agent's capability across every repository at once, with human review checkpoints and CI feedback loops, is a distinct capability tier from single-repo coding agents.
Security / enterprise
Sprinklr added platform-wide AI governance controls: a global kill switch to bulk-disable generative AI features, a default-off posture for newly released AI capabilities, and AI+ reporting for visibility into where AI is active across the platform.
Enterprises adopting embedded AI need the ability to say no, centrally and immediately. A kill switch, opt-in defaults for new AI, and usage reporting are exactly the governance controls compliance teams ask for before approving AI features in customer-facing workflows.
Integrations
LlamaIndex shipped v5 and v6 of the LlamaParse Platform community node for n8n, now an officially verified n8n community node. The node brings LlamaParse's document parsing, extraction, classification, and splitting into n8n workflows.
Verified n8n distribution puts LlamaParse's document processing inside a widely used workflow automation tool, letting ops teams parse complex documents in existing pipelines without custom integration code.
Integrations
Kore.ai v11.26.0 added agent transfer summaries that hand conversation context to human agents, expanded prompt context to the last 50 messages, integrated Deepgram Flux for speech recognition, and added Five9 and NICE contact-center integrations.
Transfer summaries and a longer prompt context reduce the context loss that frustrates customers during bot-to-human handoffs, and the Deepgram, Five9, and NICE integrations widen the contact-center stacks the platform plugs into without custom work.
Memory / state
Kore.ai added a personalization layer to AI for Work: Persona, Memory, Custom Instructions, and an Enterprise Search capability. The system builds a persistent per-employee profile that captures role, team structure, aliases, and working preferences, so the assistant carries that context across sessions instead of being re-told each time.
Persistent, per-user memory cuts the repeated context-setting that drags on daily enterprise AI use and pushes responses closer to each employee's actual role and priorities. For buyers, it is a concrete signal of how mature the platform's memory and personalization story is.
Agent capability
Thomson Reuters rebuilt CoCounsel Legal on agentic infrastructure, moving it off discrete prompt-driven tasks. Users describe a legal matter in plain language and the agent plans the workflow, pulls authoritative content from Practical Law and Westlaw, and runs multi-step research and drafting in one continuous conversation.
An autonomous plan-and-execute loop grounded in verified legal databases removes the need to chain individual tools or master prompting. For legal buyers, the draw is reduced manual oversight without giving up sourcing discipline.
Agent capability
JetBrains made Codex (on GPT-5.4-mini) the recommended default agent in JetBrains AI Chat across JVM, .NET, and Python, after internal benchmarking against Junie and Claude Agent on solve rate, cost, and latency.
A vendor-chosen default removes the friction of manual model selection and sets a predictable performance and cost baseline for teams standardizing on JetBrains agents.
Integrations
Rippling launched Data Cloud, a BI layer that ties third-party operational data to worker identities through new Data Connectors and Snowflake Zero Copy. Rippling AI turns natural-language prompts into inspectable SQL, charts, and workflows on top of that unified data.
Linking usage and operational data to employee records lets leaders attribute AI and tooling spend to actual output, while the unified data layer removes the manual pipeline work that usually precedes this kind of analytics.
Memory / state
OpenCode 1.17.11 added session snapshots and rollback, letting users revert a session to an earlier message and undo file changes. The release also brought file-based agent loading, skill discovery, and expanded OpenAI model support via AWS Bedrock (building on the MCP resource tooling and minimal CLI mode shipped the day before in 1.17.10).
Snapshot-and-revert gives teams a safety net to experiment without risking irreversible AI edits, and Bedrock routing keeps prompts and billing inside existing cloud infrastructure.
Agent capability
SAP introduced Joule Work, a unified dashboard across desktop, web, and mobile where users describe a task in natural language and Joule orchestrates the agents, workflows, and data fetches to do it across SAP and non-SAP systems. It builds on Joule Studio, the low-code/no-code environment SAP launched two days earlier for building custom agents.
A single natural-language command surface over SAP's transaction-heavy UI lowers the barrier to routine enterprise tasks, and Joule Studio gives IT and partners a path to build context-aware agents without deep coding.
MCP / tool calling / API
Dify 1.15.0 shipped difyctl, a CLI for running apps and workflows from the terminal or CI/CD, alongside a redesigned role-based access control system, improved human-in-the-loop forms, and a patch for the CVE-2026-41948 path traversal flaw.
A CLI lets automated systems trigger Dify workflows without the web UI, while the RBAC overhaul and CVE fix matter for teams running Dify in production under governance requirements.
Feature release
Mastra shipped a built-in Event System (publish/subscribe via Redis Streams and Google Cloud Pub/Sub) on June 25, a day after adding an agent Inbox with priority notifications (June 24) that lets external sources like GitHub and Slack trigger and wake agents.
Adds native event-driven triggering and channel-based wake-up to the TypeScript agent framework, so teams can build agents that react to external systems without bolting on a separate queue or trigger layer.
Agent capability
Kiro shipped IDE 1.0. The release adds a capability-based permissions system that evaluates every file read, command execution, and MCP call against user-defined rules, prompting for consent on anything not pre-approved and persisting decisions as workspace- or global-scoped rules. Custom agents are now defined in a single Markdown file with read/write/shell/web tool-access tags, inline MCP servers, and permission rules that are shareable via version control. Also ships: experimental Agent Focus mode for directing multiple agents working in parallel, natural-language hook creation in a structured JSON format, dockable chat tabs, and session export.
The 1.0 release pairs per-action consent controls with version-controlled custom-agent definitions and a parallel-agent mode. For buyers evaluating agentic IDEs, the permission model is a concrete governance control; shareable agent definitions give teams a path to standardize behavior across a codebase.
Agent capability
Unify shipped a conversational interface that builds end-to-end outbound campaigns from a single prompt: it navigates 40-plus data sources, searches 1.1 billion contacts, runs enrichment across 7 waterfalls, and drafts personalized sequences. The release adds reusable sequence rulesets for audience guardrails and a native PostHog integration for product-led signals.
Natural-language campaign building puts complex outbound orchestration in reps' hands and consolidates contact data, signal monitoring, and sequencing into one workflow, while rulesets keep reps inside corporate guardrails.
Human approval / guardrails
Kore.ai introduced Agent Blueprint Language (ABL), a compiled declarative language for defining and operating enterprise agents. ABL enforces governance, tool usage, and orchestration at runtime through a compiled intermediate representation, treating agent workflows as structured software rather than loose prompt collections.
Deterministic, compiled agent definitions give engineering and compliance teams a handle on prompt-injection risk and behavior drift, and the compiled path also trims multi-agent token consumption.
Agent capability
Writer launched WRITER Agent, a single interface merging the conversational Ask WRITER with the task automation of Action Agent. It adds file uploads up to 500MB and new AI Studio admin controls to toggle agent capabilities like browser automation, web search scope, and website hosting.
Folding chat and action into one surface simplifies complex task execution, and granular capability toggles let IT enforce internet-usage policy before scaling agents across the org.
Security / enterprise
Qodo announced three governance capabilities for AI-generated code at enterprise scale. Cross-Repo Code Review (beta) extends its Git plugin so that when a PR modifies a shared dependency, the agent reads registered consumer repos and surfaces cross-system impact findings (function-signature violations, API-contract breaks, schema evolution, and infrastructure drift) directly on the PR before merge. Rules Miner automatically discovers coding standards from a team's existing codebase rather than requiring them to be defined up front. Skill Review Standards adds centralized governance for agent skills, discovering them across repos and surfacing them in a portal with impact measurement.
Targets governance gaps that emerge when AI agents drive large cross-repository changes: catching breaking changes at repo intersections, making tribal coding standards machine-enforceable, and tracking which agent skills are shaping reviews. For buyers standardizing on agentic code review, these are auditability and governance controls, not raw generation features.
MCP / tool calling / API
HubSpot launched the HubSpot Agent CLI in public beta, a command-line interface built for coding agents like Claude Code and OpenAI Codex to work against CRM data. External agents can read records, run searches, update objects, and manage pipelines.
A dedicated CLI lets ops and dev teams run bulk, scheduled, or headless CRM automations more token-efficiently than conversational MCP integrations, with no human in the loop.
Agent capability
The GitHub Copilot CLI reached general availability, turning the command line into a terminal-native coding agent that can autonomously run multi-step workflows inside the developer's own environment.
Extending Copilot from the IDE into the terminal opens autonomous code generation and infrastructure tasks at the command line, where much real engineering work already happens.
Observability / auditability
Microsoft took the Azure Copilot Observability Agent to general availability. Built into Azure Monitor, it correlates logs, metrics, traces, and operational context to autonomously investigate cloud incidents and surface root causes.
Autonomous incident investigation moves SRE and IT teams off manual dashboard triage toward contextualized root-cause analysis, directly targeting mean time to recovery.
Security / enterprise
Lovable introduced Workspace Insights, an admin dashboard in its Security Center that consolidates enterprise projects into one searchable table, flagging externally published apps, projects holding PII, and unresolved vulnerabilities via a native Wiz integration. It also adds automated PII detection and redaction across chat inputs, uploads, and cloud storage.
It targets the shadow-IT risk of rapid AI-generated app building, giving security teams the visibility and central controls to deploy the tool at scale without ceding compliance.
Observability / auditability
Rasa Pro 3.17.0 added native Langfuse integration for LLM observability, auto-instrumenting paths like message processing and MCP tool calls. The release also brought token-by-token streaming for custom actions and a new API endpoint for dynamic capability discovery.
Teams get model performance monitoring and workflow tracing without writing custom instrumentation, and custom-action streaming cuts perceived latency during generative responses.
Pricing / packaging
Legora launched Agent Pro, a new tier of its AI execution engine, on a consumption-based model where customers pay for the work the AI delivers rather than per seat. A real-time dashboard with spend controls caps overages.
Outcome-based pricing lets legal teams attribute software cost per client or matter and align spend to actual usage, a notable shift from seat licensing in the category.
Observability / auditability
Lyzr released the Agent Improvement Engine, which analyzes live production traces to detect quality issues in deployed agents. Its Agent Hardening layer spots failure patterns and auto-generates recommendations to revise an agent's instructions or goals.
It moves teams from reactive troubleshooting to continuous, automated agent improvement, helping reliability scale without a large QA function as production agents drift.
Agent capability
Pydantic AI V2 reached stable release after seven betas, moving to a harness-first architecture built around 'capabilities' - a single composable primitive that bundles an agent's tools, hooks, instructions, and model settings - plus a first-party Pydantic AI Harness providing memory, guardrails, context management, file-system access, and code mode.
First stable major version of a widely used type-safe Python agent framework. The capabilities primitive standardizes how tools, hooks, instructions, and model settings compose, and the batteries-included Harness (memory, guardrails, context management, code mode) cuts the glue code teams need to ship production agents.
Security / enterprise
Snyk announced Evo Agentic Development Security (ADS), an addition to its Evo AI platform that embeds security controls inside the agent's workflow rather than scanning output after the fact. It governs three layers: agent supply-chain security (discovering and assessing MCP servers, skills, and external tools before the agent interacts with them, flagging prompt injection and malicious code patterns); real-time runtime policy enforcement (blocking destructive actions before they execute); and CI/CD validation of agent-generated code. GA is set for June 29, 2026.
Frames agent governance as a security control plane sitting inside the development loop, vetting tool supply chains, stopping unsafe runtime actions, and validating generated code. For buyers letting coding agents run with autonomy, this is a direct guardrails-and-governance capability from an established security vendor, and a signal that agentic development security is forming as its own category.
Agent capability
Anthropic launched Claude Tag (beta): a persistent @Claude teammate embedded in Slack channels. It builds channel-level memory over time, supports ambient mode (acts without being tagged), runs under a scoped per-channel identity for tool and data access, and lets admins set token-spend limits per channel and org. Runs on Claude Opus 4.8; replaces the prior Claude in Slack app, which retires August 3, 2026.
Shifts Anthropic's agent from a 1:1 assistant to an always-on multiplayer teammate inside the collaboration layer. Scoped identity, audit logs, and spend ceilings are the governance controls buyers will weigh against Copilot in Teams and Salesforce's Slackbot.
Agent capability
Copilot for JetBrains IDEs added Claude as an agent provider (public preview), letting developers run the Claude Code CLI as the agent inside Copilot Chat. The same release made org- and enterprise-defined custom agents generally available, giving admins a governed set of published agents for their teams.
Provider flexibility lets teams route agent work to Claude from inside Copilot. Admin-published org agents give a path to standardize and govern agent behavior, though the interim bypass-permissions default is a caveat for regulated environments.
Human approval / guardrails
Codex CLI v0.142.0 added configurable rollout token budgets that track usage across agent threads and abort a turn when the budget is exhausted, plus an indexed web-search mode that restricts direct page access to server-approved URLs. The release also reorganized the plugin system into Curated, Workspace, and Shared sections.
Hard token budgets and an allowlisted web-access mode give operators concrete spend and controls over autonomous runs. These are the guardrails enterprises typically require before letting coding agents run unattended.
MCP / tool calling / API
Cursor 3.9 added a Customize page for centrally managing plugins, MCPs, subagents, rules, commands, and hooks at user, team, or workspace scope, including bring-your-own MCPs. Also ships a team marketplace with a popularity leaderboard and plugin-repo imports from GitLab, Bitbucket, and Azure DevOps.
Centralized, governed MCP and agent-extension management across an org is an extensibility and control concern, relevant for teams standardizing tooling and restricting which tools agents can call.
Browser/computer use
Codex added Record and Replay: the agent observes a demonstrated screen workflow on macOS and turns the recorded sequence of browser or file tasks into a reusable executable skill.
Shifts automation from heavy prompting to visual demonstration, lowering the barrier to automating repetitive work. Teams build reliable automations for routine operational tasks without writing scripts.
Feature release
Langfuse shipped Monitors and Alerts (proactive alerting on metrics and thresholds) and the Langfuse Assistant (natural-language trace filtering, powered by AWS Bedrock with zero data retention) on June 19, followed by Multi-modal datasets on June 23.
Adds proactive production monitoring and alerting plus natural-language trace querying to the observability platform. For buyers evaluating LLM observability and eval tooling, these reduce reliance on manual dashboard checks.
Major release / new capabilities
Vellum moved to wide release with plugins as a first-class marketplace, a new Advisor capability that pulls in a second more powerful model on hard problems, Memory v3 at section grain for sharper recall, and upgrades to subagents, workflows, Slack, and the Activity page.
A marketplace-ready plugin ecosystem, multi-model Advisor routing, and higher-precision Memory v3 together signal a maturing extensibility and answer-quality story for teams evaluating Vellum at scale.
Human approval / guardrails
OpenAI Agents SDK v0.17.6 added pre-approval tool input guardrails.
Direct guardrail control on tool execution — one of the clearest SDK-level agent-governance changes in the window.
Observability / auditability
OpenAI Agents SDK v0.17.6 added SDK-only custom data for tool outputs.
Improves instrumentation and downstream metadata handling around tool execution, relevant for auditability and richer action traces.
Browser/computer use
Stack AI's E2B partnership now powers sandboxed Custom Code, Computer Use/RPA replay, and Browser Use on the StackAI platform.
Concrete browser/computer-use infrastructure change with enterprise security implications, especially for regulated-industry deployments backed by isolated E2B VMs.
Workflow orchestration
LangGraph 1.2.6 fixed nested subgraph checkpoint namespace inheritance and cancellation of running subgraphs on v3 stream abort.
Nested subgraphs and stream interruption behavior are core orchestration primitives for multi-step and multi-agent systems.
Agent capability
Haystack 2.30.2 prevents Agent from exiting prematurely when an invalid tool call is discarded; the model can continue looping and recover.
Improves real agent behavior in multi-step/tool-calling runs and improves recovery from malformed tool outputs.
Workflow orchestration
Cursor Automations gained /automate, Slack emoji triggers, five new GitHub triggers, default computer-use support for automation-run cloud agents, default PR opening, and memory-file deletion controls.
Materially changes the trigger surface, creation path, and action surface for always-on agents — relevant for teams building automated coding workflows.
Workflow orchestration
[Pre-release / alpha] CrewAI 1.14.8a / 1.14.8a1 introduced JSON-first crews and expanded FlowDefinition: script/code block actions, crew actions, each composite actions, expressions, DMN mode, no-Python flow run tools, human feedback from flow definitions, config/persistence wiring, experimental `crewai run --definition`, ZIP deployment fallback, and optional `if` expressions for `each.do`.
A large workflow-definition and deployment-surface expansion. This is an official alpha/pre-release — buyers should not rely on these capabilities until stable release confirmation.
MCP / tool calling / API
Tray.ai made MCP dynamic authentication generally available. MCP tools can now run using end-user credentials instead of shared service accounts, with allowlists, authentication management, and a usage monitor.
Materially improves MCP governance, identity propagation, and enterprise traceability for tool calls.
Observability / auditability
Voiceflow added transcript filtering by workflow, playbook, and tool usage.
Improves post-hoc analysis and operational debugging by linking conversations to the workflow and tool paths actually used.
Pricing / packaging
Voiceflow added organization-level usage breakdowns showing spend by workspaces, projects, date range, and by Runtime, Measure, Development, and Generation features.
Changes billing/usage visibility available to operators and improves packaging transparency for larger organizations.
Workflow orchestration
Pipecat v1.4.0 added realtime_service_mode on LLMContextAggregatorPair, the on_user_turn_message_added event, and RealtimeServiceMetadataFrame, changing how realtime speech-to-speech services write context and expose turn behavior.
Significant orchestration/runtime change for realtime agents — affects turn timing, latency, server-vs-local turn strategy selection, and pipeline metadata.
Workflow orchestration
Pipecat v1.4.0 added the pipecat create project-scaffolding CLI via an optional cli extra, including optional Pipecat Cloud enablement.
Lowers the barrier to packaging and standing up agent projects, including cloud-oriented workflows.
Funding / partnership
Legora and Ironclad announced a phased AI-to-AI integration that will let mutual customers bring Ironclad contract intelligence into Legora and bring Legora legal intelligence into Ironclad.
Materially expands enterprise system-of-record connectivity and cross-platform legal-agent workflow coverage for in-house legal teams.
Observability / auditability
deepagents-code 0.1.19 added dual-write of agent traces to extra LangSmith projects.
Concrete observability/routing improvement for teams managing multi-project LangSmith operations.
Integrations
Adomik became a default Dust integration via OAuth with read-only access to monetization analytics and the Adomik knowledge base.
Relevant for specialized vertical analytics workflows where monetization data should be agent-accessible.
Integrations
Lemlist MCP Server became available across all Dust workspaces for lead retrieval, contact enrichment, and sequence management.
Expands outbound-sales workflow coverage for teams using Dust as a workspace AI platform.
Deployment / data residency
Cursor added cloud environment setup in the Agents Window, reusable snapshots via .cursor/environment.json, /in-cloud cloud subagents on isolated VMs/branches, and local-to-cloud handoff.
Materially expands remote execution, reproducibility, parallel agent execution, and local/cloud handoff for teams running Cursor in CI or multi-agent workflows.
Workflow orchestration
Devin Playbooks now support a structured output schema so Devin can return defined JSON results.
Improves deterministic workflow outputs and downstream automation compatibility for teams integrating Devin Playbooks into automated pipelines.
MCP / tool calling / API
Devin added an official Axiom MCP server and improved MCP error cards with direct View logs links and detailed plugin status.
Adds another official MCP server and improves operational visibility for MCP failures — relevant for teams debugging agent tool-use in production.
Workflow orchestration
Devin added Slack / automation / Linear workflow improvements: preceding Slack thread context, `!agent` routing into full agent sessions, default Slack sync, automation-result posting to Slack channels, and Linear project filters for automation triggers.
Materially improves collaborative agent handoff, automation routing, and workflow context inside Slack and Linear for enterprise teams.
Security / enterprise
Devin Review now includes a Security Findings section, and the security reviewer respects the repository's SECURITY.md.
Adds explicit security-review output and repo-specific security policy awareness — directly relevant for teams using Devin in security-sensitive codebases.
MCP / tool calling / API
June Cline CLI releases expanded the agent/plugin/MCP surface: plugin commands can submit prompts to the agent, plugins gained MCP server support with OAuth during install, Cline added `cline skill`, and the CLI added a prefilled MCP install wizard. The same releases added output caps and ranged reads for large files.
Materially improves tool/plugin extensibility, MCP server setup, and safer handling of large tool outputs — relevant for teams building on or extending Cline.
MCP / tool calling / API
Activepieces release 0.85.4 added AI-ready metadata across hundreds of pieces, tagged non-agent actions with audience='human', and updated builder docs to require AI metadata on new actions and triggers.
Improves tool discoverability, tool suitability labeling, and the quality of AI-facing action and trigger metadata — relevant for teams integrating Activepieces pieces into agentic workflows.
Observability / auditability
Activepieces release 0.85.4 shipped broader runtime visibility improvements, including evlog-wide events replacing OpenTelemetry + pino, date ranges for queue health metrics, and worker/sandbox state and RAM visibility.
Strengthens operator visibility into worker, sandbox, event, and queue health — relevant for teams running Activepieces in production environments that require operational observability.
Workflow orchestration
Activepieces release 0.85.4 added error handlers to steps in flows.
Improves failure handling and recovery behavior inside multi-step automations — relevant for teams building resilient workflows where individual step failures should not abort the entire flow.
MCP / tool calling / API
langgraph-cli 0.4.30 added support for compatible API version ranges.
Version-range compatibility reduces operational friction between CLI tooling and server/API deployments.
Observability / auditability
Agent Chat Evaluations let teams define criteria and automatically grade agent chats.
Adds native QA/evaluation instrumentation for production agents instead of relying only on external eval tooling.
Human approval / guardrails
Human-in-the-loop for agents lets agents pause mid-task to request approval before running a tool, or ask a question with options, then resume where they left off in agent chats and Slack.
A meaningful governance/control upgrade for agent deployments where autonomous tool use needs checkpoints.
Integrations
Clari Copilot integration became generally available to all Dust workspaces via MCP server configuration.
Adds enterprise sales-call intelligence as an agent-usable source/action surface.
Memory / state
Email replies now stay in a single Dust conversation instead of creating a new conversation for each thread reply.
Material state and continuity improvement for email agents — preserves conversation history across thread replies.
Funding / partnership
Bland announced an additional $50M in funding, bringing total raised to more than $100M. Named backers include Scale, Emergence, HubSpot, Dell Technologies Capital, and Upfront.
A material company-level signal for durability, hiring capacity, and enterprise go-to-market momentum.
Deployment / data residency
Tray Headless MCP added region-specific endpoints for US, EU, and APAC workspaces.
Regional MCP endpoints affect deployment geography and handling for MCP-compatible clients.
Memory / state
Tool responses can now be saved to properties, not just variables.
Expands how tool output can persist in project state and metadata — relevant for downstream tagging, tracking, and memory-like state handling.
MCP / tool calling / API
Retell removed legacy list endpoints and pushed users to versioned v2/v3 list APIs with unified pagination.
Material API compatibility and integration-maturity change affecting list/query behavior for teams building on Retell's API.
Funding / partnership
Salesforce signed a definitive agreement to acquire Fin for about $3.6B, with closing expected in Salesforce fiscal Q4 2027.
Major ecosystem and go-to-market event with likely downstream impact on enterprise reach, partner strategy, and roadmap signaling.
Workflow orchestration
ElevenLabs JS and Python SDK v2.53.0 added conversation evaluation rerun support, workflow entry_behavior and end-procedure tools, telephony filters, and expanded agent test batch limits.
Material for SDK-driven testing, evaluation, and workflow controls in voice agents.
MCP / tool calling / API
Dust became connectable as a remote MCP server to any MCP-capable client. Admins can disable MCP access and restrict allowed redirect URIs.
Major interoperability change — Dust agents can now be invoked from any MCP-compatible client, with meaningful admin-control implications for enterprise deployments.
Funding / partnership
Decagon announced an accredited integration with Five9 and availability on the Five9 CX Marketplace.
Material enterprise ecosystem and distribution change for contact-center deployment — relevant for buyers evaluating Decagon in CCaaS environments.
Security / enterprise
Custom Roles API added an `org_id` parameter for enterprise customers using workspace-scoped API tokens.
Improves enterprise API governance and organization-specific access control — relevant for large teams managing multiple repositories or workspaces.
Agent capability
CodeRabbit now runs oasdiff during PR reviews to detect breaking changes between OpenAPI specifications.
Adds API-aware review capability for detecting breaking contract changes — relevant for teams with public or partner-facing APIs.
Human approval / guardrails
Claude Code v2.1.178 added parameter-level permission rules using Tool(param:value) syntax and changed auto mode so sub-agent spawns are classifier-checked before launch.
This strengthens tool governance and sub-agent launch controls — relevant for teams that need fine-grained control over what tools agents can invoke and how sub-agents are authorized before execution.
MCP / tool calling / API
Agno v2.6.15 made the AgentOS MCP server a configurable extension point via MCPServerConfig, with custom tool registration, built-in tool scoping, caller-identity injection, authorize hooks, and allowed_hosts / allowed_origins protection.
A major MCP and tool-governance update. It expands agent interoperability while adding stronger controls around tool exposure and caller identity — relevant for enterprise teams requiring scoped, auditable MCP integrations.
Workflow orchestration
Databricks announced Omnigent, an open-source meta-harness above existing agents that adds multi-agent composition, contextual policies, real-time collaboration, cloud execution, and multi-harness authoring.
Direct index-relevant shift toward cross-agent coordination, policy control, collaboration, and agent operations rather than a single-agent capability update.
Deployment / data residency
deepagents-code 0.1.16 added a Vercel Sandbox provider and ChatGPT OAuth sign-in for Codex models.
Expands execution-environment options for agent code and reduces authentication friction for Codex-based workflows.
Agent capability
Workspace-level Agent customization added reusable workspace instructions plus on-demand skills.
Changes how teams standardize agent behavior across projects, improving consistency and repeatability.
Security / enterprise
Package Firewall now blocks malicious or compromised packages by default, and Agent adapts its plans to safer alternatives.
Material governance and security improvement for production coding agents — reduces supply-chain risk in agent-assisted builds.
Memory / state
LangGraph 1.2.5 fixed DeltaChannel state updates on empty threads and merged lc_versions config metadata.
These fixes affect correctness of persisted state and version metadata in long-running, stateful agent graphs.
Security / enterprise
Devin added custom OAuth configuration for Marketplace MCP servers and an enterprise setting to enable or disable web search.
Strengthens enterprise control over MCP authentication and external web-search behavior — key governance controls for regulated or security-conscious deployments.
Deployment / data residency
Braintrust released terraform-aws-braintrust-data-plane v5.5.0, updating Braintrust Services/API ECS to v2.2.1, changing the default Redis instance type from cache.t4g.medium to cache.r7g.large, and enabling Brainstore fast readers by default.
Materially affects self-hosted AWS deployment defaults, capacity assumptions, and upgrade behavior for teams running Braintrust on their own AWS infrastructure.
Observability / auditability
Claude Code v2.1.174 added usage attribution in /usage showing cache misses, long context, subagents, and per-skill/agent/plugin/MCP breakdowns over 24h and 7d windows in VS Code.
Improves visibility into where agent usage is coming from across tools, plugins, subagents, and MCP — relevant for teams managing cost attribution and operational oversight in multi-agent deployments.
Memory / state
Agno v2.6.14 added create, read, update, and delete endpoints for AgentOS learnings.
CRUD access to learnings is a concrete memory and state management capability for agent persistence and operational memory — relevant for teams building agents that need to accumulate and manage knowledge over time.
MCP / tool calling / API
Agent Bricks Supervisor Agent added support for Unity Catalog volumes as subagent tools.
Expands the tool surface available to supervised multi-agent systems and is a concrete tool-access orchestration update.
MCP / tool calling / API
Outreach released Outreach MCP Client for Omni, letting reps pull knowledge and actions from third-party MCP connectors, browse/install marketplace connectors, and use private MCP connectors restricted to their instance.
Adds MCP consumption, marketplace connectors, and private connector controls inside the agent surface — a clear MCP/index-relevant capability change.
Memory / state
Outreach added Knowledge Enhancements surfaced in Omni, allowing admins to upload approved documents, index them, assign channels, validate them in Playground, and expose cited source documents in Omni responses.
Materially changes Omni's grounding layer: enterprise memory becomes admin-curated, channel-scoped, testable, and source-visible.
Browser/computer use
Outreach added Web Search in Personalization Agent, enabling real-time external web context, prompt-controlled retrieval behavior, and visible source citations in preview.
Direct web-augmented agentic capability: externally grounded personalization with reviewable provenance before send.
Workflow orchestration
Outreach consolidated Smart Account/Deal Assist into Omni on Account, Prospect, and Opportunity, routing account and deal queries through Omni across multiple entry points.
Product-surface orchestration change: the AI experience becomes centralized rather than split across separate assistants.
Observability / auditability
Outreach released Individual Revenue Agent Reporting, adding direct performance-report links from Revenue Agent list and overview views to agent-scoped Sequence Performance Reports.
Strengthens measurement of autonomous/semi-autonomous revenue-agent outcomes with agent-level performance reporting.
Security / enterprise
langgraph-cli 0.4.29 added support for passing a certfile and cert key to run the dev server under HTTPS.
Improves secure local/staging deployment ergonomics for agent apps and agent-server testing.
Integrations
CrewAI 1.14.7 added native Snowflake Cortex LLM provider support and Databricks / Snowflake integration guides.
Expands enterprise data and AI platform integration surface for teams running CrewAI alongside Snowflake or Databricks.
Observability / auditability
CrewAI 1.14.7 surfaced real `finish_reason`, sampling parameters, and `response.id` on LLM events.
Improves observability into LLM calls and execution outcomes — useful for teams debugging agent behavior or auditing model usage.
Workflow orchestration
CrewAI 1.14.7 added chat API for conversational flows, route-aware DSL triggers, and FlowDefinition construction from Flow DSL metadata.
Expands CrewAI's workflow orchestration and conversational flow surface, enabling more dynamic agent routing and flow construction.
Memory / state
CrewAI 1.14.7 added pluggable default backends for memory, knowledge, RAG, and flow.
Materially improves CrewAI's configurability around memory/state and orchestration infrastructure — teams can now swap storage/retrieval backends without forking.
Integrations
Finishing Touches expanded platform coverage: Autofix on Azure DevOps PRs, plus Generate unit tests and Custom recipes on GitLab merge requests.
Broadens CodeRabbit's agentic review and remediation capabilities across additional code-hosting surfaces.
Memory / state
CodeRabbit added Learning approvals so admins can review newly created learnings before they enter the knowledge base.
Adds governance over agent memory and knowledge-base mutation — relevant for teams where incorrect learnings could affect review quality.
Security / enterprise
CodeRabbit added SkillSpector scanning for AI agent skills and MCP configuration files during PR reviews, inspecting SKILL.md, MCP configs, Claude desktop config, Cursor rules, and Codex YAML for risks.
Directly relevant to agent security — provides automated risk detection for AI agent skill definitions and MCP configurations in code reviews.
Integrations
CodeRabbit IDE extension added direct Devin handoff links.
Strengthens cross-agent workflow support by allowing CodeRabbit findings or plans to move directly into Devin for implementation.
Memory / state
Braintrust released bt CLI v0.12.0, adding dataset snapshot create/list/restore/delete support plus bt topics btmap download.
Dataset snapshots improve reproducibility and state/version control for eval datasets. Topic-map artifact export improves analysis workflows.
Security / enterprise
Sierra announced FedRAMP High certification, making its agent platform available for U.S. federal-agency deployment.
FedRAMP High materially changes Sierra's enterprise/government eligibility and security/compliance posture for regulated deployments.
Memory / state
Voiceflow added Personas, letting teams test agents with predefined starting variable values for specific customer types or account states.
Adds reusable stateful testing context and improves repeatability for agent evaluation.
MCP / tool calling / API
Kapa launched Kapa for Agent Context, a retrieval layer for agents available via Hosted MCP Server or Retrieval API, grounded in docs, code, PDFs, tickets, Slack, and 30+ sources, with citation-backed fallback behavior.
Clear shift from 'AI assistant for documentation' toward agent-context infrastructure and MCP-native grounding — relevant for teams building agents that need authoritative, cited knowledge retrieval.
Human approval / guardrails
Fin over Email added pre-live reply preview/testing, a dedicated Spam view with customizable spam guidance, and channel-specific guidance/escalation controls.
Meaningful safe-deployment controls — preview/testing before go-live, spam governance, and per-channel behavior separation are relevant for teams deploying Fin in regulated or high-volume email environments.
Workflow orchestration
Fin over Email now sends configurable follow-up emails to customers who go quiet, with delays from 1 hour to 30 days, across Email, Zendesk, Salesforce, Freshdesk, and HubSpot.
Changes case-closure and re-engagement automation behavior in agent-driven support workflows.
Funding / partnership
E2B joined the Stripe Projects developer preview so agents can discover, provision, and authenticate E2B sandboxes without manual API-key setup.
Material ecosystem change for agent runtime provisioning and credential flow — lowers the integration barrier for teams building agentic systems with E2B sandboxes.
Observability / auditability
Bugbot became roughly 3x faster, about 22% cheaper, and found about 10% more bugs per review. /review can now run Bugbot and Security Review before push, with diff deduplication across local reviews and GitHub/GitLab PRs.
Material change to agentic code-review performance, cost, and pre-push workflow coverage — relevant for teams using Cursor's automated review capabilities.
Observability / auditability
Devin Review now cancels pending PR reviews when new commits arrive, and the PR Review API exposes a new cancelled status.
Improves auditability and state tracking for automated PR reviews — ensures stale reviews don't complete on outdated code.
MCP / tool calling / API
Devin added the official Figma MCP integration and deactivated the previous unofficial integration.
Adds an official MCP-based design-tool integration, replacing an unofficial path — relevant for teams using Devin in design-to-code workflows.
Workflow orchestration
CodeRabbit Plan became available directly in the VS Code extension, with a Plans tab for creating agent-ready Coding Plans and handing phases to configured AI coding agents.
Moves planning/orchestration into the IDE and makes CodeRabbit more directly useful in agentic coding workflows.
Workflow orchestration
Cline CLI v3.0.23 added support for configured agents as subagent tools and centralized OAuth management into the SDK.
A concrete subagent/tool orchestration expansion — configured agents can now be invoked as tools, improving multi-agent workflow composition.
MCP / tool calling / API
browser-use released 0.13.1, adding support for Claude Fable 5 and changing Anthropic behavior so tool choice is auto when using thinking. This is a dated compatibility release; current Fable 5 availability has not been separately verified.
Post-cutoff compatibility change relevant to model support and tool-calling behavior for browser-use deployments using Anthropic models. Buyers should verify current Fable 5 availability directly.
Workflow orchestration
Claude Code v2.1.172 enabled sub-agents to spawn their own sub-agents, up to five levels deep.
This materially expands recursive and delegated agent orchestration depth, relevant for teams building complex multi-agent coding pipelines or automated software-engineering workflows.
Human approval / guardrails
Agno v2.6.13 added socket support for human-in-the-loop workflows.
Directly relevant to approval loops and interactive workflow orchestration — socket-based HITL enables real-time human intervention in running agent workflows.
Integrations
Tray.ai updated its OpenAI connector to improve token management and model-selection logic across model families.
Real integration update for agent builders using Tray's OpenAI connector with model-family-aware token handling.
Observability / auditability
Voiceflow introduced Tests, which simulate conversations and check whether responses, routing, and tool calls match expected behavior before production exposure.
Material verification and QA capability for agent builders, directly affecting how reliably agent behavior can be validated before go-live.
Observability / auditability
Salesloft added three AI-usage metrics in Analytics: Account researched, Person researched, and Agent tasks completed.
Relevant to measurable agent adoption and observability — teams can now track AI task completion alongside traditional activity metrics.
MCP / tool calling / API
Salesloft added direct connection to Claude through the MCP Connectors Directory for customers with the Agentic add-on.
Expands the agentic surface area beyond the native UI — clear MCP/connectors change gated by the Agentic add-on.
Workflow orchestration
Cadence Collections introduced a new parent container for grouping cadences and inheriting settings, permissions, and CRM sync rules.
Changes how teams govern and orchestrate repeatable engagement workflows at scale.
MCP / tool calling / API
Agentforce Vibes 2.0 added an Abilities & Skills framework that dynamically activates domain-relevant context including MCP tools, plus additional frontier models and React app generation.
Changes how Agentforce-based builders access tool/context scaffolding and expands supported model and UI outputs.
Workflow orchestration
NICE announced that agentic AI is native at the core of the CX platform, naming NICE AI Agents, the Agentic Engagement Plane, Guardian AI, and Agentic Analytics as key layers of the operating model.
Formalizes orchestration, governance, and analytics layers as first-class platform components rather than ancillary messaging.
Observability / auditability
NICE launched the CXone Workforce Empowerment Suite for managing human and AI agents under one operating model, with shared dashboards, quality/compliance controls, and AI operations visibility.
Relevant as hybrid human/AI workforce governance and monitoring, especially for enterprise CX operations.
Observability / auditability
LlamaParse added opt-in granular bounding boxes at line, word, and cell level, with beta availability across paid tiers.
Granular provenance materially improves auditability, citation precision, and human verification in document-grounded agent workflows.
Integrations
Fin can run as a service agent on top of HubSpot and Freshdesk without requiring migration off the current helpdesk.
Major platform-openness and deployment-surface expansion — buyers on HubSpot or Freshdesk can now deploy Fin without switching helpdesks.
MCP / tool calling / API
Haystack 2.30.1 lets AzureOpenAIChatGenerator accept Secret values for azure_endpoint and api_version, enabling environment-variable-based runtime switching.
Production-configuration improvement for teams moving the same serialized pipeline across environments without hardcoding credentials.
MCP / tool calling / API
Large MCP toolsets now lazy-load tool details only when needed, improving context efficiency and reducing unnecessary token usage.
Materially improves tool-heavy agent scalability and lowers context/cost pressure in large MCP deployments.
Integrations
Agents can now be connected to Microsoft Teams as an external channel.
Broadens deployment surfaces into a major enterprise collaboration environment.
Agent capability
Decagon announced Duet Autopilot, which turns production signals into validated agent updates ready for human review. Decagon also announced DuetBench alongside it.
Material change toward self-improving agents with explicit human-review gating — relevant for buyers evaluating autonomous quality improvement in customer-support agents.
Security / enterprise
Composio CLI 0.2.31 included security dependency updates: authlib bumped to 1.7.2 for GHSA-wvwj-cvrp-7pv5 and protobufjs pinned to 7.5.5 for a critical Socket.dev CVE.
Security maintenance relevant for developer tooling and API/MCP infrastructure. Low-impact public index update but important for security-conscious teams.
Integrations
Automatic Repository Linking can now discover and link related repositories across an organization so cross-repo context appears in reviews.
Improves cross-repo context and repository-level memory for code review workflows in multi-repo organizations.
Agent capability
Cline v3.89.0 added Claude Fable 5 model support. This is a dated release event; current Fable 5 availability has not been separately verified.
Model availability affects the coding agent's supported backend capability set — buyers should verify current Fable 5 availability directly with Cline/Anthropic.
Security / enterprise
Stagehand server-v3 v3.7.2 added Azure Entra model authentication support.
Improves enterprise authentication posture for Browserbase / Stagehand model-backed browser-agent infrastructure — relevant for enterprise teams requiring Entra-based auth.
Pricing / packaging
API access through API keys became available on Beam's Free plan.
Changes programmatic-access gating and lowers the threshold for experimentation, evaluation, and integration.
Human approval / guardrails
Beam added "Test Before You Publish," a pre-publish testing mode that runs flows against historical real data in a safe environment, sandboxes integrations, retains test execution review, and warns if live connections remain active before publish.
A meaningful governance and control upgrade for agentic workflows, especially when integrations can mutate external systems.
Agent capability
Augment Code added Claude Fable 5 to the model picker, positioning it for long, multi-step, deep-reasoning work. This is a dated release event; current Fable 5 availability has not been separately verified.
Expands the model surface available to Augment's agent workflows — buyers should verify current Fable 5 availability directly with Augment.
Agent capability
Claude Code v2.1.170 added access to Claude Fable 5 as of this release date. This is a dated release event; current Fable 5 availability has not been separately verified.
Model/backend availability is index-relevant when it materially changes the agent's available reasoning or coding capability — buyers should verify current availability directly with Anthropic.
Workflow orchestration
Activepieces release 0.85.2 added data manipulation triggers.
Expands trigger-layer flow input and transformation capability, which is relevant to workflow flexibility and pre-processing in agentic automation pipelines.
Agent capability
Salesforce added Google Gemini as a selectable model for Agentforce / the Atlas reasoning engine.
Material model-choice expansion for Agentforce deployments — buyers can now select Gemini alongside existing model options.
Workflow orchestration
Business Details headers became configurable with up to 12 fields and org-wide/default layouts.
Changes investigator workflow design and layout governance for business-review operations.
Observability / auditability
Device Intelligence now supports up to three custom charts, and the prior map visualization was removed.
Improves risk-team monitoring surfaces with configurable dashboard visualizations.
MCP / tool calling / API
Sardine added multi-party AML screening for wire transactions so one API call screens sender, sender bank, receiver, receiver bank, and memo.
Material API/screening capability expansion — reduces integration overhead for compliance teams requiring multi-party wire coverage.
Observability / auditability
Screening widgets now support batch entity resolutions across sanctions, PEP, and adverse media lists.
Reduces manual queue handling and improves audit consistency in high-volume review operations.
MCP / tool calling / API
ElevenLabs changed the default conversational ASR provider from elevenlabs to scribe_realtime and added conversation_product to conversation details responses.
Concrete API/runtime changes for conversational-agent speech processing and metadata — relevant for teams building on or integrating ElevenLabs Conversational AI.
Integrations
Gamma became available as a Remote MCP Server in Dust for generating and reading Gamma decks, docs, and webpages.
Adds a new MCP-based artifact generation and output surface for teams using Dust in content-creation workflows.
Workflow orchestration
Dust introduced skill composition, letting skills reference other skills and tools inline in instructions. The separate tools section was removed.
Material orchestration change for reusable multi-step capabilities — enables modular, composable agent skill design.
Agent capability
Builders can ask @dust to create or edit skills directly from a conversation.
Materially lowers friction for agent-skill authoring and iteration — no need to leave the conversation to update agent capabilities.
Browser/computer use
browser-use released 0.13.0, "Rebuilt in Rust [beta]," introducing a new Rust-backed beta agent with a more direct browser-control loop. The existing Python agent remained unchanged.
A major browser-agent architecture change — the Rust-backed beta agent directly affects how browser/computer-use tasks are executed.
Agent capability
CrewAI published 1.14.7a1 (June 3) and 1.14.7a2 (June 5) with trained-agents file support, native Snowflake Cortex LLM provider, Databricks and Snowflake integration guides, conversational flow traces, a chat API for conversational flows, route-aware flow DSL work, more detailed LLM event surfacing, and flow/lock-store refactoring.
Meaningful OSS movement for CrewAI's agent orchestration and conversational-flow direction. These are pre-release/alpha builds, not stable GA releases — buyers should evaluate accordingly before relying on these capabilities in production.
Pricing / packaging
Cognition introduced an AI Productivity Guarantee for eligible enterprise Devin customers: if Devin delivers less engineering value than the customer pays for, Cognition will issue usage credits up to $10M. Value is measured in equivalent engineering hours using an estimator that reviews completed Devin sessions.
This is a rare commercial guarantee tied to measurable agent output, and it materially affects enterprise buyer conversations around ROI, accountability, and spend governance.
Agent capability
Activepieces v0.85.0 added Mistral AI as a platform AI provider, formula/data-manipulation functions in the flow builder, platform-admin visibility into run success rates/live queue depth/30-day health history, clearer failed-step errors with one-click AI help, AI chat reliability improvements, model tiering with native Anthropic thinking, improved tool-calling UX, pure tool-calling parsing, multi-item action execution with live progress, display tools with agentic rendering, and AI badges for MCP-created flows.
A meaningful platform release touching agent execution, AI chat, model/provider support, workflow building, MCP labeling, and operations visibility.
Workflow orchestration
Haystack v2.30.0 added PythonCodeSplitter for syntax-aware Python code splitting in code-RAG/code-search pipelines; allowed ChatGenerator components to accept plain strings as messages; updated DALL-E image generation defaults for OpenAI's newer image models; added async retriever support; and fixed an Agent bug where multiple tool calls could prevent the configured exit condition from stopping the loop.
The most material framework-level update in this review window: touches code-centric agent/RAG workflows, developer ergonomics, async execution, model migration, and agent control-flow reliability.
Agent capability
Augment announced Cosmos, a platform for AI-native engineering teams. Cosmos is positioned around an agentic SDLC where agents operate across triage, spec, implementation, review, testing, deployment, and feedback; it supports teams of specialized agents that coordinate, delegate, and share memory. Available to all plans and every team plan.
This shifts Augment's positioning from coding assistance toward a broader software-engineering agent platform, with workflow orchestration, shared memory, MCP/webhook integrations, and multi-agent coordination.
MCP / tool calling / API
Gumloop 9.11.0 expanded the Gumloop MCP server with tools for agents, sessions, skills, files, teams, connected MCP servers, resources, prompts, and MCP tool calls. The release also added Looker MCP support and BigQuery Workload Identity Federation.
Improves Gumloop's value as an agent/workflow platform that can expose and govern more tools, enterprise data systems, and MCP-accessible capabilities.
Security / enterprise
Gumloop 9.11.0 added the ability to add skills from Shared With Me and Organization tabs directly to agents, and introduced skill permission roles — Editor, Viewer, and Use Only — so teams can govern who can update, inspect, or use skills inside agents.
An enterprise-relevant governance update for reusable agent skills, permissions, and multi-user workspace control.
Agent capability
Langflow 1.9.6 fixed Agent LLM streaming so tokens stream incrementally, and included a remote component error-handling fix for Docling.
Improves runtime behavior and user experience for agent execution and streaming responses.
Agent capability
Cognition announced Devin Desktop, repositioning Windsurf around the Agent Command Center as the default IDE surface for managing local and cloud agents, PRs, and context. The release adds Spaces so related agents can share context, and adds Agent Client Protocol (ACP) support so ACP-compatible agents can run inside Devin Desktop alongside Devin.
A major product-surface shift from IDE/coding assistant toward multi-agent management for software engineering teams.
MCP / tool calling / API
Langflow 1.9.5 shipped a fix for model handling for tool calling in agents and updated IBM models. The release also included documentation updates around AGENTS.md and project philosophy docs.
Incremental but material for framework users relying on agent tool-calling and IBM model support.
Integrations
Activepieces released v0.84.0 with new integrations including Resend, Plausible Analytics, PostHog, UptimeRobot, Azure DevOps, Frill, Beebole, YouTrack, Hootsuite, PDF4me, Sendr, and Streak CRM, plus GitHub App authentication and safer handling for disabled subflows.
A broad integration push across dev-ops, analytics, CRM, and productivity tools extends Activepieces' automation reach without custom connectors.
Agent capability
Devin added Auto-Triage for autonomous issue prioritization, native Windows VM support for cross-platform development tasks, and broader deployment and financing announcements signaling expanded enterprise ambition.
Auto-Triage and Windows VM support expand Devin's autonomous engineering footprint into issue management and cross-platform builds — material for enterprise teams evaluating end-to-end coding agent coverage.
Human approval / guardrails
Cursor shipped Composer 2.5, Jira integration, shared canvases, /loop, and Auto-review Run Mode for safer long-running Shell, MCP, and Fetch execution — using allowlists, sandboxing, and classifier-based approval logic to control what agents can run autonomously.
Auto-review Run Mode with classifier-based approval is a significant guardrails addition for teams concerned about autonomous agent actions in production codebases and CI pipelines.
MCP / tool calling / API
Gumloop's late-May releases added MCP Artifacts, queued mid-run steering messages, hosted agent pages, richer HTML artifacts, Freshdesk and Freshsales MCP, Apify MCP, team-level secrets, a notification center, and Claude Opus 4.8 support.
MCP Artifacts and mid-run steering messages strengthen Gumloop's case for complex, interactive agentic workflows where runtime human guidance and structured tool outputs matter.
Security / enterprise
CrewAI released v1.14.6 with StdioTransport protections against environment-variable leakage, improved planning and observation handling, and fixes for structured output leakage, checkpoint restore, and executor resume behavior.
Environment-variable and structured-output leakage protections directly address security concerns for teams running CrewAI agents with sensitive credentials or proprietary data.
MCP / tool calling / API
n8n shipped reliability fixes for binary-data temp-directory cleanup and webhook close-function draining to prevent MCP connection leaks.
MCP connection leak prevention is directly relevant for production deployments where long-running n8n agent workflows depend on stable MCP server connections.
Funding / partnership
Cognition announced it raised over $1B at a $26B valuation, led by Lux Capital, General Catalyst, and 8VC. Enterprise usage is up more than 10x since the start of the year and run-rate revenue is at $492M.
Materially changes the vendor's market position, enterprise traction profile, and funding/scale narrative — relevant context for any enterprise evaluation or vendor-longevity assessment.
Observability / auditability
Mastra added stored-entity HTTP APIs, browser session probing, screencast support, delta-polling observability, stricter authorization defaults, Convex native vector search, Google Cloud Spanner storage, upgraded agent channels, client-side tool observability, and a FilesSDK-backed workspace filesystem provider.
Delta-polling observability, client-side tool visibility, and stricter auth defaults are enterprise-readiness signals for JS/TS teams evaluating Mastra for production agent deployments.
Workflow orchestration
Augment Code's Cosmos releases added service-account attribution, scoped webhooks, hosted artifacts, broader Slack controls, file attachments in web conversations, and simplified custom environment requirements (bash and git only).
Service-account attribution and scoped webhooks improve traceability and governance for enterprise teams running Augment in CI/CD and collaborative coding workflows.
Browser/computer use
browser-use shipped 0.12.7–0.12.9 with a major CLI update, tighter daemon socket permissions, restricted browser-profile handling, cached-content and prompt-history improvements, and new-tab screenshot reliability fixes.
Tighter socket permissions and restricted profile handling address security hygiene for teams running browser-use in shared or production environments.
Agent capability
Agno added support for Google Antigravity and Gemini managed agents for Deep Research, plus approval metadata in post-hooks, PgVector prefix-match fixes, and a Gemini tool-call path correction.
Native Gemini Deep Research and Antigravity integration expands Agno's model-agnostic positioning and strengthens its case for teams building research-grade multi-modal agents.
Browser/computer use
Browserbase updated its Fetch API to return markdown and structured JSON, raised the response cap from 1 MB to 5 MB, and positioned Fetch as a lower-cost read path than launching a full browser session.
A structured-output Fetch API at lower cost than full browser sessions meaningfully reduces cost and latency for agents doing web data extraction without needing full browser automation.
Security / enterprise
Dify released v1.14.2 with security hardening, workflow and HITL reliability fixes, RAG and document-processing improvements, knowledge-base stability updates, tracing reliability, and self-hosted deployment and runtime changes.
Continued hardening across security, HITL, and knowledge-base stability is relevant for self-hosted production deployments where reliability and data governance are evaluation criteria.
Security / enterprise
Sourcegraph added granular admin permissions via RBAC, allowing teams to delegate specific admin capabilities without granting full administrator access.
Granular RBAC is a foundational enterprise governance requirement — directly relevant for large teams and regulated industries where least-privilege access is a procurement criterion.
Pricing / packaging
MindStudio sharpened its agent-pricing position around token-based usage, free and Individual tiers, 200+ model access, and Business-tier controls such as SSO, audit logs, usage alerts, and unlimited collaborators.
Useful for buyers comparing agent platforms against SaaS-native "second meter" pricing, lock-in risk, and enterprise governance requirements.
Workflow orchestration
Gumloop's 9.5.0 changelog added Shared With Me and Organization views across agents, skills, files, and workflows, plus consolidated agent activity in the Home tab.
Helps teams move from individual automation projects to shared departmental workflows with better discoverability, handoff, and operational visibility.
Workflow orchestration
Lyzr published a finance-ops case study on payment-receipt verification automation, including portal, email, and WhatsApp intake, payment-method guardrails, API connectivity, and manual exception routing.
Useful customer evidence for buyers evaluating agentic automation in finance operations, especially where exception handling and system integrations matter.
Funding / partnership
11x published a post claiming Ramp spend data ranked it first in enterprise adoption for AI-first GTM, with 40 percent of enterprises buying in that category choosing 11x.
Useful as a market-validation signal, but should be treated as a vendor-stated procurement-data claim rather than a fully transparent public benchmark.
Browser/computer use
Cognition added Android emulator support for Devin, allowing it to run Android Virtual Devices, build and inspect Android apps, reproduce bugs, and verify changes locally.
Expands Devin's autonomous verification surface for mobile engineering teams and improves its credibility for Android testing, debugging, and app-development workflows.
Agent capability
Decagon introduced Guided Discovery, a capability for exploratory product-discovery, retention, and expansion conversations without relying entirely on predefined scripts.
Signals Decagon is expanding beyond rigid support-ticket automation into concierge-style customer conversations, commerce, retention, and expansion use cases.
Security / enterprise
Dify released v1.14.1, a security and stability patch covering self-hosted SECRET_KEY hardening, internal metrics protection, an IDOR fix, tenant-scoped credential cleanup, LiteLLM CVE remediation, and workflow, HITL, and knowledge-base fixes.
Meaningful for self-hosted and production users because it addresses concrete security, reliability, workflow, and knowledge-base risks.
API / Governance
Devin added a Review API for programmatic pull-request review, per-PR auto-review toggles, service-user permission management, and MCP secret scoping.
These updates strengthen Devin's enterprise story around CI/CD integration, review automation, auditability, permissioning, and secure engineering-agent operations.
Decagon's resources page added a customer clip showing how GlossGenius uses Duet to scale concierge customer experiences.
This is useful buyer evidence for Decagon's Duet adoption story, but should be treated as customer validation rather than a net-new platform capability.
Feature / Orchestration
Gumloop announced Subagents, allowing agents to clone themselves for parallel work and call other specialized agents.
This strengthens Gumloop's position for buyers comparing multi-agent orchestration, delegated workflow execution, and more complex agentic automation patterns.
Product / Positioning
CrewAI published Discovery, a new engine designed to surface the best automation use cases for a business. May 4–8 open-source releases included LLM listing refreshes, a status endpoint fix, gitpython security compliance, and CLI packaging cleanup.
This moves CrewAI's positioning upstream from agent orchestration into automation prioritization, which matters for buyers still trying to identify where agentic workflows are most likely to produce ROI.
Integration / Governance
Gumloop's 9.2.0 changelog added Gmail triggers for any label, audit-log filters, Google Analytics MCP, and Outlook Calendar MCP.
The update broadens Gumloop's operational integration surface while improving audit/log filtering for teams that need better workflow visibility and governance.
Agent capability
Cursor launched Background Agents, enabling long-running autonomous coding tasks that execute in the cloud without requiring the IDE to remain open — agents can clone repos, run code, and open PRs independently.
Background Agents represent a meaningful shift from copilot-style assistance to fully autonomous software engineering, directly relevant for teams evaluating agentic coding platforms.
Integrations
Relevance AI added Confluence Knowledge Sync, project-level public-agent controls, and an Android app, improving enterprise knowledge connectivity, governance, and mobile workforce access.
Relevant for buyers evaluating RAG/knowledge grounding, governance over public-agent sharing, and operational access for mobile-first teams.
Funding / partnership
Decagon announced a new APAC office to support increasing regional demand for its AI concierge platform and enterprise customers in the region.
Signals commercial maturity, expanded customer support footprint, and growing regional enterprise traction — relevant for APAC buyers evaluating vendor longevity.
Workflow orchestration
Gumloop's late-April releases added subagents, agent email inboxes, an organization analytics agent, MCP hosting/proxying, enterprise data drains to S3 and BigQuery, Google Search Console MCP, and higher Pro credit limits.
A major release cycle across multi-agent coordination, observability, MCP infrastructure, enterprise export, and packaging — relevant for buyers evaluating Gumloop for production-scale deployments.
Human approval / guardrails
Dify v1.14.0 added real-time collaborative workflow editing, a service API for human-in-the-loop, multiple MCP and OAuth fixes, optional Langfuse TTFT reporting, billing/quota improvements, stronger tenant checks in knowledge APIs, and security hardening.
A substantial release for self-hosted and enterprise evaluators covering collaboration, HITL, MCP reliability, observability, billing, data governance, and security.
Integrations
StackAI introduced a Box integration for knowledge bases, workflows, conversational assistants, triggers, and write-back to Box. Also published a deployment-options guide covering multi-tenant cloud, dedicated cloud, private cloud/BYOC, and on-premise.
Improves enterprise content integration and gives buyers clearer deployment, residency, and compliance context when evaluating StackAI.
Workflow orchestration
Beam AI added a Code Execution Node that lets builders run JavaScript or Python directly inside flows without LLM overhead. Also added zip-file trigger support and Outlook attachment public-URL handling.
Expands Beam from LLM orchestration toward deterministic workflow execution inside agentic applications — relevant for buyers needing reliable, cost-efficient compute steps.
Security / enterprise
Cognition launched Devin for Terminal, enabling local-to-cloud session handoff, alongside new enterprise controls including network policy restrictions and more granular PR review permissions.
Improves local-to-cloud coding-agent workflows and strengthens enterprise governance and security evaluation criteria.
Observability / auditability
Decagon introduced automatic optimization and Root Cause Analysis through Duet.
Automated quality improvement and RCA tooling reduces manual QA overhead in production support deployments.
Human approval / guardrails
Gumstack added App Policies for enterprise control over how agents use apps and data.
App Policies give enterprise buyers programmatic guardrails over agent behaviour, a key requirement for regulated-industry deployments.
Agent capability
Ava 2.0 is publicly marketed by Artisan, but official pages conflict on whether it is fully live or still launching. This item requires manual review before relying on capability claims.
If confirmed live, Ava 2.0 represents a significant upgrade to Artisan's outbound sales agent. Buyers should verify directly with the vendor before making decisions based on this update.
Workflow orchestration
LangGraph Platform reached general availability, providing a managed cloud runtime for deploying, scaling, and monitoring stateful multi-agent LangGraph workflows, with support for persistence, streaming, and human-in-the-loop interrupts.
GA status means production-grade SLAs and managed infrastructure for teams that want LangGraph's orchestration model without self-hosting the backend.
Human approval / guardrails
CrewAI changelog and docs verify Flow human-in-the-loop (HITL), user input handling, and HITL loop behaviour across recent releases.
Native HITL in CrewAI Flows enables supervised agentic workflows without external orchestration tooling.
Workflow orchestration
Devin can now manage parallel child Devins and schedule recurring sessions with persistent state across runs.
Multi-agent orchestration with scheduling and state persistence is a critical capability signal for teams evaluating autonomous coding workflows at scale.
Funding / partnership
Gumloop raised $50M Series B and positioned Gumstack as a security and observability layer for organizational agents.
Series B funding and a security-positioning shift signals enterprise ambition and product longevity.
Agent capability
Decagon introduced proactive agents, user memory, and outbound campaigns.
Proactive outreach and persistent user memory represent a significant leap beyond reactive support agents, relevant for teams evaluating autonomous customer engagement.
Deployment / data residency
AGI Inc. announced a Qualcomm collaboration to bring agentic AI to Snapdragon-powered devices.
On-device agentic AI via Snapdragon chips reduces latency and enables offline-capable agent deployments.
Deployment / data residency
All Hands AI launched OpenHands Cloud, a managed SaaS offering of the open-source OpenHands coding agent (formerly OpenDevin), providing a hosted runtime without requiring self-hosted infrastructure.
Cloud availability removes the main adoption barrier for teams that want OpenHands' open-source flexibility without managing the execution environment themselves.
Workflow orchestration
StackAI January release added human-in-the-loop, custom RBAC, SDLC environments, node run results, OpenAI web search, and multiple integrations.
Human-in-the-loop and RBAC in a single release indicate serious enterprise workflow maturity.
Funding / partnership
Decagon announced Series D / $250M commitment.
Significant funding validates category leadership in AI customer-support agents and signals long-term vendor stability.
Product / Positioning
Mastra launched as a TypeScript-first agent framework with built-in support for workflows, RAG, evals, memory, and an agent development Studio UI — positioning as the TypeScript equivalent of LangChain for teams building production agents in Node.js environments.
TypeScript-native agent frameworks are underserved; Mastra's built-in eval and observability tooling makes it relevant for JS/TS engineering teams evaluating agent platforms.
MCP / tool calling / API
Zapier launched an MCP server exposing its 7,000+ app integrations as tool-callable actions for AI agents, allowing external agents (Claude, Cursor, etc.) to trigger Zapier automations directly via the Model Context Protocol.
An MCP server covering 7,000+ apps dramatically expands the integration surface available to any MCP-compatible agent — a major interoperability development for buyers evaluating tool-calling breadth.
Funding / partnership
Lyzr announced $8M Series A and Accenture investment and collaboration.
Accenture involvement signals enterprise-channel distribution and adds credibility for regulated-industry deployments.
Verification status update
AGI Inc. publicly claimed AGI-0 reached 97.4% task success on the AndroidWorld benchmark.
Benchmark performance claims at this level, if verified, would represent a significant capability advance for computer-use agents.
Integrations
Julian operates across voice, chat, SMS, and WhatsApp in IBM Agent Connect materials.
Multi-channel coverage (voice, chat, SMS, WhatsApp) broadens Julian's deployment surface for inbound sales workflows.
Product / Positioning
Phidata rebranded to Agno and released v1.0 of its agent framework, repositioning as a high-performance, model-agnostic Python runtime for building multi-modal agents with native support for memory, tools, and multi-agent coordination.
The rebrand and v1.0 milestone signal production readiness and a sharper positioning around performance and multi-modal agent architectures.
Memory / state
Voice 2.0 added faster voice processing, self-serve controls, and cross-channel memory.
Cross-channel memory enables continuity across voice and text support interactions, improving customer experience quality.
Integrations
Dify 1.7.0 added OAuth support for tools and multi-credential management.
OAuth for tools and multi-credential management enables enterprise-grade secure integration patterns.
Pricing / packaging
Relevance AI split pricing into Actions and Vendor Credits, enabled BYO API keys, removed markup on vendor costs, and sunset the Business plan.
BYO API keys and removal of vendor cost markup can substantially reduce total cost for high-volume workloads.
Agent capability
n8n added native AI Agent nodes with memory, tool-calling, and MCP client/server support to its workflow builder, enabling agentic loops to run inside automation workflows without requiring separate orchestration infrastructure.
Native agent nodes with MCP support position n8n as a hybrid automation-plus-agent platform, relevant for buyers wanting agentic capabilities within existing workflow tooling.
Funding / partnership
AGI Inc. announced a Visa partnership for agentic commerce and secure agent-led transactions.
A Visa partnership signals production-grade agentic payment workflows, relevant for commerce and fintech use cases.
Deployment / data residency
Dify Enterprise became available in Microsoft Azure Marketplace.
Azure Marketplace availability simplifies enterprise procurement and cloud-native deployment.
MCP / tool calling / API
Dify 1.6.0 added built-in two-way MCP support, allowing users to call MCP servers from Dify and expose Dify apps as MCP servers.
Two-way MCP support is a major interoperability signal for builders evaluating agent platforms with tool-calling and MCP requirements.
Browser/computer use
Simular introduced Simular Pro as its first production-ready computer-use agent.
A production-ready computer-use agent tier is a significant maturity milestone, relevant for teams evaluating browser and desktop automation.
Observability / auditability
Dify 1.5.0 added real-time workflow debugging with saved node outputs and live variables.
Real-time debugging with persistent node outputs significantly reduces cycle time when building and maintaining complex workflows.
Observability / auditability
Decagon launched A/B experimentation and Watchtower always-on QA monitoring.
Built-in experimentation and continuous QA reduce the engineering overhead of operating production support agents safely.
Integrations
Added Hunter, ContactOut, Prospeo integrations, and an Outlook Calendar trigger.
Sales data enrichment integrations and Outlook support expand Lindy's coverage for B2B outreach and productivity workflows.
Workflow orchestration
Relevance AI launched Invent, Workforce, and AgentOS — a suite positioning the platform as a full agent operating system.
An AgentOS framing with Workforce management signals enterprise readiness for managing fleets of agents across organizational workflows.
Funding / partnership
11x announced an IBM partnership to distribute Alice and Julian through IBM Agent Connect.
IBM distribution significantly expands 11x's enterprise reach and adds procurement credibility for large-org buyers.
Agent capability
11x launched Julian, an inbound sales digital worker, alongside the existing outbound agent Alice.
Adding an inbound-focused agent completes the full sales coverage story and broadens the platform's addressable use cases.
MCP / tool calling / API
Lindy added a Code action that runs custom Python or JavaScript directly within workflows.
Custom code execution within workflows enables advanced automation without external API calls or separate compute infrastructure.
Pricing / packaging
Devin 2.0 introduced a lower-cost entry point. Current pricing shows Core, Team, and Enterprise packaging.
A clearer tier structure with a lower entry price makes Cognition more accessible to smaller engineering teams.
Agent capability
Devin 2.0 introduced an agent-native IDE experience.
An agent-native IDE changes the developer workflow model, relevant for engineering teams evaluating AI coding agents.
Verification status update
Lyzr's Responsible-AI materials document explainability, auditability, bias control, and compliance posture.
Formal responsible-AI documentation is a prerequisite for enterprise procurement in regulated industries.
Human approval / guardrails
Relevance AI added Slack-based approval for agent tool requests.
Slack-native approval flows make human-in-the-loop controls accessible without switching tools, reducing friction for enterprise adoption.
Product / Positioning
deepset released Haystack 2.0, a complete framework rewrite introducing a component-based pipeline architecture, native support for agentic loops with tool use, and improved composability for RAG, extraction, and multi-agent workflows.
The 2.0 rewrite is the relevant baseline for any current Haystack evaluation — the architecture, APIs, and agent capabilities differ substantially from v1.
Current rankings
These changes feed the index. See where each category stands: the best vendors of 2026, ranked by verified capability coverage.
Last updated: August 2026 · 561 tracked changes across 235 vendors































































