Back to vendors
D

Deepgram

Visit site
Entry priceFree $200 credit; Pay as you go; Growth from $4,000/year; EnterpriseFull pricing detail

Voice AI infrastructure: real time speech to text (Flux, Nova-3), text to speech (Flux TTS, Aura) and a Voice Agent API that joins listening, thinking and speaking, with function calling and a choice of LLM.

Deepgram is voice AI infrastructure sold as APIs: speech to text, text to speech, audio and text intelligence, and a Voice Agent API, in real time and batch, hosted or self hosted. Flux is its conversational speech recognition model with built in turn detection for voice agents, Nova-3 is its general speech recognition family across 45+ languages, and Flux TTS and Aura produce speech. The Voice Agent API runs one WebSocket session that listens, calls an LLM and speaks, with function calling that runs in the customer's application or against an endpoint the customer provides.

The customer picks the agent's LLM per agent from OpenAI, Anthropic, Google and NVIDIA models Deepgram manages, or brings its own endpoint for Groq, Amazon Bedrock or any OpenAI compatible API, with an ordered fallback chain across providers. Agent configurations can be stored and reused through an API. Deepgram keeps conversation state within a call and can load prior turns the customer supplies; it does not keep a memory of its own across calls. Per call turn by turn records come from events the customer captures on the socket, beside Console request logs kept for 90 days.

Access is managed through account permissions and owner, admin and member project roles with scoped API keys. Deepgram states SOC 2 Type 1 and Type 2, signs a BAA for Enterprise customers under HIPAA, and states GDPR, CCPA and PCI compliance. Regional endpoints keep processing in the EU or Australia, and Premium customers can self host in their own cloud or data center.

Pricing is usage based: a $200 free credit, then Pay As You Go per minute of speech to text, per thousand characters of text to speech and per minute of Voice Agent connection time by tier, or a Growth plan from $4,000 a year in prepaid credits. Enterprise is sold through sales.

Vendor details

Canonical URL

https://deepgram.com

Category

Agent infrastructure

Subcategory

Voice infrastructure

Funding status

Independent voice AI company.

Company status

independent

Use cases & customers

Primary use cases

speech to text transcriptiontext to speechvoice agentsconversational audio intelligence

Target customers

developersenterpriseregulated industries

Deployment options

SaaSself-hostedon-premVPC

Integrations

REST and WebSocket APIs for speech to text, text to speech, text intelligence and the Voice Agent, with Python, JavaScript, Go and .NET SDKs, a Browser Agent SDK and embeddable widget, and the dg CLI with a built in MCP server. Voice Agent function calling runs client side or against a customer endpoint. Telephony guides cover Twilio and Amazon Connect. EU and Australian regional endpoints and self hosting on AWS, GCP, SageMaker, Docker or Kubernetes.

In practice

Your voice agent needs accurate transcription on noisy call center audio. Deepgram Nova-3 leads on challenging audio with sub 300 millisecond streaming, so the agent hears callers correctly in real time.

You want to build a voice to voice agent without stitching components. The Voice Agent API unifies speech to text, LLM orchestration, and text to speech in one route with function calling and bring your own LLM.

A healthcare deployment needs compliant, on premise voice. Deepgram self hosts with HIPAA and a business associate agreement, a Nova-3 Medical route for terminology, and redaction for sensitive data.

Agentic Index coverage score

7.5 / 14 capabilities · 54%

Integrations & Tool Calling Full

Documented custom tool support. Voice Agent function calling lets the customer define functions in Settings; each runs client side in the customer's application or server side, where Deepgram calls a web endpoint the customer provides, so an agent can book appointments, send emails or update records in the customer's CRM, ERP or internal APIs during a live call.

SourceDeepgram, developers.deepgram.com (Function Calling)read 2026-09-27

Workflow Orchestration Partial

A fixed pipeline that Deepgram runs end to end. Each Voice Agent session chains listen, think and speak, with an ordered LLM fallback chain per request and mid call Update messages for the prompt and providers; the buyer configures the pieces, not the control flow. The multi agent architecture guide sequences qualifier, advisor and closer agents through a CallOrchestrator in a reference repository the customer runs, so that orchestration is the customer's code.

SourceDeepgram, developers.deepgram.com (LLM Models, Build a Multi-Agent Architecture)read 2026-09-27

Knowledge Grounding & RAG Not documented

No retrieval structure over the customer's content. The documentation index lists speech to text, text to speech, audio and text intelligence, and Voice Agent pages, and none describes ingesting, indexing or retrieving the customer's documents; context reaches the agent through the prompt, loaded history and function results the customer supplies. Keyterm prompting tunes transcription vocabulary, not knowledge.

SourceDeepgram, developers.deepgram.com (llms.txt, Maintaining Context)read 2026-09-27

Human Oversight & Guardrails Partial

A customer controlled constraint on agent actions, with no approval step. Per function, defer_until_eot holds a call whose side effect cannot be undone (ending a call, spending money, sending a message) until speech to text confirms the caller has finished the turn, while read only functions still dispatch early; Deepgram's docs tell builders to set it on any irreversible action. That is a configurable constraint on what the agent may do; no person approves an action inside Deepgram.

SourceDeepgram, developers.deepgram.com (Function Calling: Irreversible Actions and Turn Confirmation)read 2026-09-27

Security, Identity & Governance Full

Access and compliance are both documented. On access, there is a named access model of account permissions plus owner, admin and member project roles, each implying a listed set of scopes, with API keys created per project and scoped to those permissions. On compliance, Deepgram lists SOC 2 Type 1 and Type 2 by an independent auditor, certificates on request, HIPAA business associate with a BAA for Enterprise, GDPR, CCPA and PCI with a yearly review.

SourceDeepgram, developers.deepgram.com (Working With Roles & API Scopes, Data Privacy Compliance) and deepgram.com/pricingread 2026-09-27

Observability & Auditability Partial

Request logs and usage, with the per session trace left to the customer. The Console shows usage and up to 90 days of request logs, and a Usage API exports them to tools such as Grafana or Datadog. For the Voice Agent, Deepgram's own guide says the dashboard does not give per session, turn by turn observability and there is no separate logging API: the customer taps the WebSocket and persists the transcript, function call, latency and error events it emits. Events a vendor emits for the customer to store are not a trace the buyer can inspect on the vendor's side.

SourceDeepgram, developers.deepgram.com (Session Observability, Logs & Usage Data)read 2026-09-27

Memory & State Persistence Partial

Conversation state lives inside a session. The agent keeps the prompt, turns, injected messages and function results as working memory for the call, and History (on by default) lets a new session load prior turns and function calls through agent.context.messages. The prior turns are stored and supplied by the customer's application; Deepgram documents no memory layer of its own with a stated scope and lifetime.

SourceDeepgram, developers.deepgram.com (Maintaining Context, History)read 2026-09-27

Deployment & Data Residency Full

Named regions offered as an option and a customer environment option. Regional endpoints at api.eu.deepgram.com (EU, never routed outside it) and api.au.deepgram.com (Australian infrastructure for storage and inference) serve speech to text, text to speech and the Voice Agent with the same keys; self hosting runs in customer requisitioned cloud instances such as AWS or GCP, or the customer's data center, for Premium customers. A third party LLM in the think step runs on that provider's infrastructure.

SourceDeepgram, developers.deepgram.com (Regional Endpoints, Deployment Options, Data Privacy Compliance)read 2026-09-27

Prebuilt Agents, Templates & Packs Not documented

No packaged agents a buyer adopts. The Voice Agent template apps are one starter ported to twelve languages and frameworks on GitHub, which is sample code, evidence of a framework rather than of a ready solution; Reusable Agent Configurations store the customer's own agent blocks; industry tuned models such as Nova-3 Medical and Pharma are models, not agents.

SourceDeepgram, developers.deepgram.com (Template Apps, Reusable Agent Configurations)read 2026-09-27

Triggers & Channel Coverage Partial

Voice channel coverage without an autonomous wake. The Voice Agent works on phone calls through documented inbound and outbound telephony builds (Twilio, Amazon Connect) and in web pages through the Browser Agent SDK and an embeddable widget. Every session starts when a caller dials in or the customer's own server opens the socket or places the call (the outbound reference takes a POST from the customer's CRM or CLI), so Deepgram owns no event, schedule or webhook wake.

SourceDeepgram, developers.deepgram.com (Build an Outbound Telephony Agent, Browser Agent SDK)read 2026-09-27

Model Flexibility & Routing Full

The customer chooses the model. Each Voice Agent's Settings name the think provider and model from OpenAI, Anthropic, Google and NVIDIA models Deepgram manages, or the customer's own endpoint for Groq, Amazon Bedrock or any OpenAI compatible API, and an ordered array of providers acts as a per request fallback chain that can mix providers; listen and speak models are chosen the same way, including BYO TTS.

SourceDeepgram, developers.deepgram.com (LLM Models)read 2026-09-27

APIs, SDKs & MCP Extensibility Full

A documented API and SDKs for Deepgram's own platform. REST and WebSocket APIs for speech to text, text to speech, text intelligence and the Voice Agent, a management API for projects, keys and usage, an Agent Configuration API, Python, JavaScript, Go and.NET SDKs with a feature matrix, a Browser Agent SDK, and the dg CLI, whose built in MCP server proxies the developer API's tools.

SourceDeepgram, developers.deepgram.com (llms.txt, Reusable Agent Configurations, MCP Server)read 2026-09-27

Testing, Debugging & Optimization Not documented

No documented way to evaluate an agent change. The documentation index carries no evaluation harness, scored test cases or quality gate for voice agents; the testing pages validate integrations (a streaming starter kit, a SageMaker endpoint check). Reusable Agent Configurations name A/B testing voices or prompts as a use case, but the conversion or CSAT measurement is the customer's own. Custom model training improves transcription, which is model work, not agent evaluation.

SourceDeepgram, developers.deepgram.com (llms.txt, Reusable Agent Configurations)read 2026-09-27

Browser & Computer Use Not documented

The Browser Agent SDK embeds a voice agent in a web page, which is the product running in a browser, not an agent operating one, and telephony is a voice channel. No browser, desktop or computer control is documented.

SourceDeepgram, developers.deepgram.com (llms.txt, Browser Agent SDK)read 2026-09-27

The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Recent platform changes

2026-10-02·Agent capabilityVerified

Nova-3 streaming sessions in Deepgram can now swap their list of key terms mid call without reconnecting, on every Nova-3 model including medical. It is live on the global endpoint, not yet on the regional ones.

Bears on: MCP / tool calling / API

View source
2026-09-17·Agent capabilityVerified

Deepgram released Nova-3 Pharma, a speech to text model trained on pharmaceutical vocabulary for accurate drug name recognition in healthcare voice agents.

Bears on: Agent capability

View source
2026-09-14·MCP / tool calling / APIVerified

Deepgram updated its Python, JavaScript and Java SDKs with new Voice Agent controls, including ForceEndTurn and Flux text to speech expressivity controls.

Bears on: MCP / tool calling / API

View source
View all 9 changes for Deepgram →Tracked since Jul 2026 · Verified from public vendor sources

Pricing

Free $200 credit; Pay as you go; Growth from $4,000/year; Enterprise

usage (minutes, characters, agent minutes)

Free tierTrial available

Included quota

A $200 free credit opens all endpoints (speech to text, text to speech, Voice Agent API, Audio Intelligence) with no card. Pay As You Go has no minimums. Growth (from $4,000 a year) gives prepaid credits at up to 20% lower rates and higher concurrency. Enterprise adds negotiated rates, self hosted deployment and dedicated support.

What is public

Detailed per minute and per character rates, add on rates, free credit, and Growth minimum are published; Enterprise is custom.

Billing mechanics

Separate usage meters per product: speech to text per minute, Aura text to speech per 1,000 characters, and the Voice Agent API per minute of WebSocket connection time, with bring your own LLM or TTS lowering the agent rate. Add ons like diarization and redaction meter separately.

Cost watchouts

Streaming speech to text rates on the pricing page are limited time promotions, with the regular rates shown beside them. Add ons (diarization, redaction, summarization) are billed separately. Multichannel audio multiplies base cost by channel count. Growth is a prepaid annual commitment from $4,000.

Variable cost rationale

Cost is entirely usage across minutes, characters, and agent session time, amplified by separately metered add ons and multichannel multipliers, so spend tracks call and audio volume directly.

Overage / add-ons

Pay As You Go bills per minute (speech to text), per 1,000 characters (text to speech) and per minute of WebSocket connection time for the Voice Agent API by tier (Standard, Advanced, or lower with your own LLM or TTS), with add ons metered separately.

Sales call required

Mixed (some tiers require a call)

Free / trial

$200 free credit, never expires, no credit card, across all endpoints

Lowest paid plan

Pay As You Go usage after the free credit; Growth from $4,000 a year

Commercial notes

Independent. SOC 2 Type 1 and Type 2, HIPAA with a BAA for Enterprise customers, GDPR, CCPA and PCI. EU and Australian regional endpoints and self hosted deployment for Premium customers.

Key ambiguities

Current streaming rates are promotional and will revert to the regular rates shown; effective cost depends on model, Voice Agent tier and enabled add ons.

Agentic Index verified 2026-09-27

Alternatives to Deepgram

The closest documented capability profiles to Deepgram among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • Composio8.5 / 14Fuller documented coverage on Observability & Auditability and Triggers & Channel Coverage
  • Inworld AI8.5 / 14Adds documented Testing, Debugging & Optimization
  • Vocode6.0 / 14A lighter documented profile than Deepgram
  • Gentoro6.5 / 14Adds documented Testing, Debugging & Optimization
  • TrueFoundry9.5 / 14Adds documented Testing, Debugging & Optimization
  • Hyperbrowser8.0 / 14Adds documented Prebuilt Agents, Templates & Packs and Browser & Computer Use

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.