Hamming AI
Also known as: Hamming
Voice and chat agent testing, red-teaming and production monitoring platform that simulates thousands of calls, scores them on 50+ metrics and turns production failures into regression tests.
Hamming is a testing, red-teaming and monitoring platform for voice and chat agents. It auto-generates hundreds of test scenarios from an agent's system prompt and runs up to 50,000 concurrent simulated calls with accents, interruptions, background noise, DTMF and IVR trees in more than 65 languages, scoring each with more than 50 built-in metrics and customer-defined expected outcomes. A curated red-team suite probes prompt injection, jailbreaks, PII leakage and policy violations, and every finding becomes a reproducible failing test rerun on each release. Tests run in CI/CD so a failing build does not ship.
In production, Hamming scores every call in real time, shows which tools an agent called or skipped alongside the transcript, ingests OpenTelemetry traces, spans and logs, replays a golden set of calls every few minutes to catch drift, and alerts through email, Slack and PagerDuty, turning any failed call into a permanent regression test. It connects to agents on Vapi, Retell, ElevenLabs, LiveKit, Pipecat and Twilio over SIP or WebRTC, and tests chat agents with the same framework.
Hamming is SOC 2 Type II certified with a HIPAA BAA available, and offers SSO with Okta, Azure AD and Google Workspace, role-based access per workspace, audit trails with SIEM export, configurable retention, US, EU and UK data residency, and single-tenant deployment with customer-managed keys. Teams can start testing free; Startup, Agency and Enterprise plans are arranged by call.
Vendor details
Canonical URL
https://hamming.ai
Category
Agent infrastructure
Subcategory
Voice agent testing and evaluation
Funding status
Founded in 2024 by Sumanyu Sharma (CEO), previously Head of Data at Citizen and a Senior Staff Data Scientist at Tesla. A Y Combinator Summer 2024 company, with offices in San Francisco, Austin, and London (legal entity Forward Inc.). Reports testing more than four million calls across over ten thousand agents. Customers include Podium, CallRail, Synthflow, and 11x. Independent.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
API first with REST APIs and webhooks for CI/CD integration through GitHub Actions, Jenkins, and other pipelines. Connects to voice stacks including LiveKit, Pipecat, Retell, Vapi, Twilio, and Webex via SIP or WebRTC, and tests agents across 65 plus languages and regional accents.
In practice
Your voice agent passes transcript tests but real callers hit failures from interruptions and accents. Hamming places thousands of simulated calls with realistic caller behavior and audio metrics to catch the failures text testing misses.
You change a prompt and worry it will regress your contact center agent. You wire Hamming into CI so every deploy runs a regression suite and blocks bad prompts from reaching production.
Your voice agents handle regulated calls and you need proof they stay on script. Hamming monitors every production call, flags compliance issues in real time, and generates the call quality and compliance reports auditors expect.
Sources & related URLs
Related / legacy domains
Research sources
Agentic Index coverage score
6.5 / 14 capabilities · 46%
| Integrations & Tool Calling | Partial |
|---|---|
|
Agents built on Vapi, Retell, ElevenLabs, LiveKit, Pipecat and Twilio are reached over SIP or WebRTC so Hamming can call them; it also ingests OpenTelemetry data and alerts through Slack, PagerDuty and email. These let Hamming reach the customer's agent under test and send results out; no integration through which an agent reads or writes business systems through Hamming is described. SourceHamming, hamming.ai homepage and enterprise pageread 2026-09-21 |
|
| Workflow Orchestration | Not documented |
|
Voice and chat agents that customers build on other platforms are what Hamming tests and monitors. No sequencing, branching, retries or routing of an agent's steps run by Hamming is described. SourceHamming, hamming.ai homepageread 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
What Hamming tests is whether an agent completes expected outcomes and follows its instructions. No document ingestion, index, retrieval layer or knowledge API that grounds an agent in company data is described. SourceHamming, hamming.ai homepageread 2026-09-21 |
|
| Human Oversight & Guardrails | Partial |
|
Every call is audited against customer-defined compliance guardrails for safety violations, prompt injection attempts and policy breaches, and Hamming pages a human through Slack or PagerDuty when a live call goes wrong. Review happens as or after a call runs; no approval step, hold or block that Hamming applies to the agent's action is described. SourceHamming, hamming.ai homepage FAQ and enterprise pageread 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
Hamming is SOC 2 Type II certified and offers a HIPAA Business Associate Agreement, and customers get SSO with Okta, Azure AD and Google Workspace, role-based access control per workspace, customer-managed encryption keys for single-tenant deployments, audit trails for every test and production call, and configurable retention with PII and PHI redaction. SourceHamming, hamming.ai enterprise page and homepageread 2026-09-21 |
|
| Observability & Auditability | Full |
|
Every production call is monitored and scored in real time. Hamming shows the tool calls an agent made or skipped alongside the transcript, ingests OpenTelemetry traces, spans and logs, keeps complete audit trails for every test and production call with export to the customer's SIEM, and lets retention be configured per workspace from 7 days to unlimited. SourceHamming, hamming.ai enterprise page and homepageread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Call recordings, transcripts, traces, test results and regression suites are kept for testing and monitoring. No session, workflow or long term memory that an agent reads and writes through Hamming is described. SourceHamming, hamming.ai homepage and enterprise pageread 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
Data residency is offered in the US, EU and UK, with EU clusters for GDPR, and a single-tenant deployment with dedicated infrastructure, customer-managed encryption keys and complete data isolation can be provisioned in the customer's preferred region. SourceHamming, hamming.ai enterprise pageread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Not documented |
|
Hamming generates test scenarios from a customer's prompt and ships a curated adversarial suite and simulated caller voices in 65+ languages. No ready-made agents, templates or packaged workflows a buyer adopts for its own work are described. SourceHamming, hamming.ai homepageread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Health checks replay a golden set of calls every few minutes to detect drift or outages and send email and Slack alerts; monitoring pages a human through Slack or PagerDuty when a live call goes wrong, and tests run in CI/CD on every pull request. Work starts on a schedule, on a live call or on a code change, without a person asking. SourceHamming, hamming.ai homepage FAQ and enterprise pageread 2026-09-21 |
|
| Model Flexibility & Routing | Not documented |
|
Unresolved. Agents built on whatever models the customer uses can be tested, and Hamming scores calls with its own evaluators. Whether customers can choose or bring the models behind those evaluators is not published. SourceHamming, hamming.ai homepage; docs.hamming.airead 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Partial |
|
Teams can connect an agent via API, and Hamming ingests OpenTelemetry traces, spans and logs and runs its tests in CI/CD on every pull request. The API and SDK reference sits behind a login, so no API or SDK for calling Hamming from outside is publicly documented. SourceHamming, hamming.ai homepage and enterprise pageread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
From a voice or chat agent's system prompt, Hamming auto-generates hundreds of test scenarios, runs up to 50,000 concurrent simulated calls with accents, interruptions, background noise and IVR trees, scores them with more than 50 built-in metrics and customer-defined expected outcomes, red-teams agents for prompt injection, jailbreaks and PII leakage, runs in CI/CD so a failing build never ships, and turns any failed production call into a permanent regression test. SourceHamming, hamming.ai homepage and enterprise pageread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
Hamming places simulated voice calls and chat sessions against a customer's agent over SIP, WebRTC and platform connectors, and no browser, desktop or computer control by an agent is described. SourceHamming, hamming.ai homepageread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Pricing
All plans contact sales · start testing free
Usage based, scaling with test and monitored call volume; exact rates not publicly disclosed
Included quota
Three plans, Startup, Agency, and Enterprise, all quoted on request. Pricing scales with usage; exact included volumes are not listed.
What is public
Hamming publishes its plan structure (Startup, Agency, Enterprise) and what each includes, but not dollar figures; all plans are quoted on request.
Billing mechanics
Usage based pricing that scales with test and monitored call volume, quoted per customer. Exact per call rates, included volumes, and seat costs are not disclosed.
Cost watchouts
Test call volume and monitored production call volume are the main cost drivers and grow with usage; exact rates are not published.
Variable cost rationale
Pricing scales with the volume of test calls and monitored production calls, so cost grows directly with how much you test and how many live calls you monitor, but exact rates are not public.
Additional watchouts
Cost scales with test and production call volume, so heavy continuous testing and high call monitoring volumes raise the bill. Plans are sales quoted, so compare total cost by test volume, monitored call volume, and seats.
Overage / add-ons
Pricing scales with usage so you pay for what you test; exact per call rates and overage terms are not disclosed.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free sign-up to start testing, per the homepage
Lowest paid plan
Startup plan; pricing quoted on request
Commercial notes
Teams can sign up and start testing free, while Startup, Agency and Enterprise plans are arranged by booking a call; Enterprise adds compliance, support SLAs and a dedicated engineer.
Key ambiguities
What the free start includes and how each plan is priced are not published.
Cancellation / refund
Not publicly disclosed; arranged with the vendor.
Support SLA / resale
Founder and 7 day support on Startup, priority support on Agency, and support SLAs with a dedicated support engineer on Enterprise.
Missing data
All dollar pricing, included call volumes, per call rates, and seat costs are quoted on request and not public.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Hamming AI
The closest documented capability profiles to Hamming AI among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Freeplay8.0 / 14Adds documented Model Flexibility & Routing
- Monte Carlo8.0 / 14Adds documented Prebuilt Agents, Templates & Packs
- Patronus AI6.0 / 14Fuller documented coverage on APIs, SDKs & MCP Extensibility
- Voker6.0 / 14Fuller documented coverage on APIs, SDKs & MCP Extensibility
- Confident AI8.5 / 14Adds documented Prebuilt Agents, Templates & Packs and Model Flexibility & Routing
- Coral5.5 / 14Adds documented Knowledge Grounding & RAG
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded