Speechmatics
Speech to text and text to speech APIs for voice agents and transcription, including Agent STT for turn based agent input, with cloud, on premise, virtual appliance and on device deployment and SOC 2 Type II and ISO 27001 certification.
Speechmatics builds speech to text, text to speech and speech intelligence APIs. Its models transcribe audio in real time or in batch across 55+ languages, with speaker diarization, custom dictionaries, translation, summaries and other add ons. Speechmatics is the trading name of Cantab Research Ltd of Cambridge, UK.
For voice agents it offers Agent STT, running on its Linden 1 model: a WebSocket API that returns speaker labeled, turn based transcription with turn detection, ready to pass to an LLM. Speechmatics does not run the agent itself. The agent's reasoning, tools and call handling live in the customer's code or in voice agent platforms such as Vapi, LiveKit and Pipecat, which ship Speechmatics integrations. A turn based Voice Agent API is in early access preview and not for production traffic. Flow, the call automation platform Speechmatics previously offered, is no longer on its site.
The cloud service runs in EU, US and Australia regions, and Enterprise customers can run speech to text in containers, on Kubernetes, in a virtual appliance or on device. Security documentation covers SSO over SAML or OIDC, Admin and Member roles, ISO/IEC 27001:2022, SOC 2 Type II, GDPR and HIPAA. Realtime audio is never stored and batch data is deleted after seven days. Customers can opt into a model training discount that lets Speechmatics use anonymized data to improve its models.
Vendor details
Canonical URL
https://www.speechmatics.com
Category
Agent infrastructure
Subcategory
Voice infrastructure
Funding status
Private. Speechmatics is the trading name of Cantab Research Ltd, Cambridge, UK.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
APIs and SDKs in Python, JavaScript, .NET and Rust. Voice agent integrations for Vapi, LiveKit and Pipecat, and a Zapier integration for no code transcription. Custom dictionaries tune recognition to domain vocabulary without retraining.
In practice
A team building a voice agent on Vapi, LiveKit or Pipecat uses Speechmatics Agent STT so the agent receives speaker labeled turns and knows who said what in a multi party call.
A regulated organization runs Speechmatics transcription in containers on its own hardware so audio never leaves its infrastructure.
A contact center transcribes recorded calls in batch across dozens of languages and adds summaries, sentiment and topics to each transcript.
Sources & related URLs
Research sources
Agentic Index coverage score
3.5 / 14 capabilities · 25%
| Integrations & Tool Calling | Not documented |
|---|---|
|
No tool calling or action in outside systems. Speechmatics transcribes and synthesizes speech that voice agent stacks call; its documented integrations (Vapi, LiveKit, Pipecat, Zapier) plug Speechmatics into those platforms, which is distribution of the models, not an agent acting. SourceSpeechmatics, docs.speechmatics.com (Integrations and SDKs overview, Agent STT)read 2026-09-27 |
|
| Workflow Orchestration | Not documented |
|
No workflow model. Agent STT returns speaker labeled, turn based transcription ready to pass to an LLM, and the listen, think and speak loop runs in the customer's code or in Vapi, LiveKit or Pipecat, so orchestration lives in the platform that calls Speechmatics. The speechmatics.com/flow page for Flow, a call automation platform, now redirects to the Agent STT page, and Flow is not in Speechmatics' product list (speechmatics.com/llms.txt). SourceSpeechmatics, docs.speechmatics.com (Agent STT overview) and speechmatics.com/llms.txtread 2026-09-27 |
|
| Knowledge Grounding & RAG | Not documented |
|
No retrieval over the customer's content. The custom dictionary adds words, with optional sounds like spellings, so recognition gets names and domain terms right, which shapes what is heard, not what an agent knows; no index, knowledge base or ingestion is documented. SourceSpeechmatics, docs.speechmatics.com (Custom dictionary, llms.txt)read 2026-09-27 |
|
| Human Oversight & Guardrails | Not documented |
|
No oversight of agent actions. Speechmatics takes no actions to approve or constrain; turn detection (the service deciding from voice activity, or the customer's application ending each turn) governs when transcription is finalized, not what an agent does. Flow's call routing is no longer on Speechmatics' site. SourceSpeechmatics, docs.speechmatics.com (Turn detection)read 2026-09-27 |
|
| Security, Identity & Governance | Full |
|
Identity controls and compliance are both documented. SSO runs over SAML 2.0 or OIDC through the customer's identity provider (an add on subscription), alongside domain verification, Admin and Member workspace roles, and API keys owned by the workspace rather than the user. Compliance coverage spans ISO/IEC 27001:2022, SOC 2 Type II, GDPR and HIPAA, with certificates and reports in the SafeBase Trust Center under NDA. SourceSpeechmatics, docs.speechmatics.com (Single sign-on, Manage members, Security and compliance)read 2026-09-27 |
|
| Observability & Auditability | Partial |
|
Aggregate reporting, no per run trace. The portal charts Realtime and Batch usage by model and project with request counts, a Usage API reports batch usage, app analytics aggregate usage for applications built on customers' keys, and each batch job has a log file; realtime content is never stored, so there is nothing per session to inspect. SourceSpeechmatics, docs.speechmatics.com (Usage, App analytics, Get the log file for a job)read 2026-09-27 |
|
| Memory & State Persistence | Not documented |
|
Stateless by design. Realtime audio is never recorded or stored and transcripts are discarded once streamed; batch audio, transcripts and job configuration are kept seven days for retrieval, then deleted. No conversation state or memory is held for an agent. SourceSpeechmatics, docs.speechmatics.com (Security and compliance: Data handling)read 2026-09-27 |
|
| Deployment & Data Residency | Full |
|
Customer environment options, generally available. Beside SaaS in EU1, US1 and AU1 regions open to all customers, Enterprise customers run CPU or GPU speech to text containers, Kubernetes or a virtual appliance on their own hardware or chosen cloud, and on device deployment; Agent STT is SaaS only. SourceSpeechmatics, docs.speechmatics.com (Deployments overview, Regions, Plans)read 2026-09-27 |
|
| Prebuilt Agents, Templates & Packs | Not documented |
|
No packaged agents. The docs offer SDK quickstarts and framework integration guides, and the deprecated Voice SDK's presets (conversation, note taking, captions) are transcription settings, not agents. SourceSpeechmatics, docs.speechmatics.com (Integrations and SDKs, Voice SDK)read 2026-09-27 |
|
| Triggers & Channel Coverage | Not documented |
|
Caller invoked. Transcription and synthesis run when the customer's application opens a WebSocket session or submits a batch job; Speechmatics owns no phone numbers, channels, schedules or events that start an agent, and the Zapier trigger is Zapier's. Flow's inbound call handling is no longer on Speechmatics' site. SourceSpeechmatics, docs.speechmatics.com (Agent STT, Integrations and SDKs)read 2026-09-27 |
|
| Model Flexibility & Routing | Not documented |
|
A single provider: Speechmatics itself. The customer picks among Speechmatics' own models (Standard, Enhanced, the multilingual Melia 1 and the Enhanced Medical domain for Batch and Realtime; Linden 1 for Agent STT), all Speechmatics'; no other provider or bring your own model is documented. SourceSpeechmatics, docs.speechmatics.com (Models)read 2026-09-27 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
A documented API and SDKs for Speechmatics' own platform. A Batch REST API, Realtime and Agent STT WebSocket APIs, a Management API for projects and API keys, and SDKs in Python (Agent STT, Realtime, Batch, TTS), JavaScript,.NET and Rust. No MCP server is documented. SourceSpeechmatics, docs.speechmatics.com (API reference, SDKs)read 2026-09-27 |
|
| Testing, Debugging & Optimization | Not documented |
|
No evaluation of agent behavior. The closest thing is the accuracy benchmarking guide, which shows how to compute word error rate against human reference transcripts of the customer's own audio; that tests the speech model, not what an agent did. SourceSpeechmatics, docs.speechmatics.com (Accuracy benchmarking)read 2026-09-27 |
|
| Browser & Computer Use | Not documented |
|
No browser, desktop or computer control; Speechmatics is a speech recognition and synthesis API. SourceSpeechmatics, docs.speechmatics.com (llms.txt)read 2026-09-27 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Speechmatics released Oak 1, its own medical speech to text model, through version 1.1.0 of its batch Python SDK. Oak 1 is multilingual, tuned for healthcare audio such as ambient scribing and dictation, and handles speakers who switch language mid conversation. The pricing page lists it at $0.15 per hour for batch transcription in the EU and US.
Bears on: Agent capability
View sourcePricing
Free $100 credit · Pro pay as you go from $0.129/hr (Batch Melia 1); Agent STT $0.16/hr
usage based on audio processed, with volume tiers
Included quota
$100 in credit on signup (one credit is $1); accounts created before 1 August 2026 received a one time $25 transition credit. Realtime audio is never stored and batch data is deleted after seven days.
What is public
Per hour rates for every speech to text model and add on, the text to speech rate, both discounts and their sizes, and the Free, Pro and Enterprise plan terms are published on the pricing page and in the docs. Enterprise rates are by contract.
Billing mechanics
Billing is in credits worth $1. Pro draws down granted credits first, then charges the card on the first of each month for the previous month's usage, calculated to the second from each product's hourly rate. Volume and model training discounts can apply together. Enterprise is billed by contract with no card or portal billing.
Cost watchouts
The pricing page has a toggle that applies the model training discount, which lets Speechmatics use anonymized customer data to improve its models; confirm which side of it a quoted rate assumes. Add ons such as translation or summaries bill per hour on top of transcription, and on premise and on device deployment is Enterprise only.
Variable cost rationale
Usage based on audio hours and characters, a unit a buyer can forecast from call minutes or file hours; the effective rate depends on the model training discount election.
Additional watchouts
Confirm whether a quoted rate includes the model training discount (33% off speech to text) before comparing it with another vendor's price.
Overage / add-ons
Usage based, so consumption bills as incurred against credit rather than triggering plan overage.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free plan: $100 in credit, no payment card required
Lowest paid plan
Pro, pay as you go (from $0.129 an hour for Batch Melia 1)
Commercial notes
Self serve Free and Pro plans with published per hour rates sit beside contract based Enterprise. The model training discount prices the customer's permission to use anonymized data for model improvement.
Key ambiguities
Effective cost depends on the model, the processing mode, add ons and whether the model training discount applies, so one headline rate does not describe a workload.
Missing data
Enterprise rates and the on premise license terms are unpublished.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Speechmatics
The closest documented capability profiles to Speechmatics among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Rime3.5 / 14Matches Speechmatics across all 14 documented capabilities
- AgentOps5.0 / 14Adds documented Integrations & Tool Calling and Model Flexibility & Routing, among others
- Mem06.0 / 14Adds documented Integrations & Tool Calling and Memory & State Persistence, among others
- Unikraft6.0 / 14Adds documented Integrations & Tool Calling and Memory & State Persistence, among others
- Veecle3.0 / 14Adds documented Prebuilt Agents, Templates & Packs and Testing, Debugging & Optimization
- Coral5.5 / 14Adds documented Integrations & Tool Calling and Knowledge Grounding & RAG, among others
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded