Back to vendors
C

Cartesia

Also known as: Cartesia Sonic

Visit site
Entry priceFree plan; Pro $5/mo; Startup $49/mo; Scale $299/mo; Enterprise; usage on credits and agent minutesFull pricing detail

Real time voice models (Sonic text to speech, Ink speech to text) and Managed Agents, a hosted voice agent platform with tools, a knowledge base, batch calling and call evaluations.

Cartesia builds real time voice models on state space model research and a hosted voice agent platform on top of them. Sonic is its text to speech family, now at Sonic 3.6 across 44 languages with voice cloning, multilingual voices and pronunciation dictionaries; Ink-2 is its streaming speech to text model with built in turn detection and keyterm prompting.

Managed Agents combine Ink-2, an LLM the customer picks from Cartesia's list and Sonic into a realtime voice agent that handles turn taking and interruptions. Agents call webhook tools on the customer's server or client tools in the customer's app, transfer calls to a phone number, and search a knowledge base of documents the customer uploads, which Cartesia indexes on upload. They answer on Cartesia or Twilio numbers, SIP trunks and a WebSocket for web and mobile apps, and call batches dial up to 5,000 recipients, at a scheduled time if one is set. Every configuration change is saved as a version.

Each call keeps a transcript with turn level timestamps, call logs and custom events, and webhooks stream call and turn events including tool calls. Custom LLM judged metrics score completed calls. Enterprise organizations get SSO over SAML or OIDC, dedicated regional deployments in the US, EU, UK, India and Australia, and zero data retention; Cartesia states it is SOC 2 Type II certified.

Pricing is usage based on monthly plans: Free, Pro at $5, Startup at $49 and Scale at $299 a month, each with included credits and prepaid agent usage, and Enterprise through sales. Text to speech bills by the character, speech to text by the second and agents at $0.06 a minute.

Vendor details

Canonical URL

https://www.cartesia.ai

Category

Agent infrastructure

Subcategory

Voice infrastructure

Funding status

Independent.

Company status

independent

Use cases & customers

Primary use cases

text to speech for voice agentsspeech to text and turn detectionvoice agent platformvoice cloning and localization

Target customers

developersvoice agent buildersenterprise

Deployment options

SaaSon-premon-device

Integrations

REST and WebSocket APIs with OpenAPI and AsyncAPI specs, JavaScript/TypeScript and Python SDKs, a CLI and an MCP server. Managed Agents use webhook, client and system tools, a document knowledge base, Cartesia or Twilio numbers and SIP trunks, and call event webhooks. Model plugins for LiveKit, Pipecat, Vapi and other voice stacks.

In practice

Your voice agent feels robotic because of lag. You swap in Cartesia Sonic, which streams first audio in around ninety milliseconds, so responses land before the caller registers a pause.

You need one brand voice across languages. Cartesia clones a voice from ten seconds of audio and localizes it into dozens of languages while keeping speaker identity intact.

A regulated deployment cannot send audio to the cloud. Cartesia runs the same Sonic and Ink models on premise or on device, with inference in region for data residency and compliance.

Agentic Index coverage score

10.5 / 14 capabilities · 75%

Integrations & Tool Calling Full

Documented custom tool support. Managed Agents take webhook tools, where Cartesia sends the HTTPS request to an endpoint the customer exposes, client tools that run in the customer's app over the WebSocket, and system tools for ending a call, sending DTMF and transferring to a phone number, up to 32 tools per agent, so an agent can look up an order or trigger an action during a call. Plugins for LiveKit, Pipecat and other voice stacks distribute Cartesia's models and are not agent integrations.

SourceCartesia, docs.cartesia.ai (Tools, Webhook tools, Integrations)read 2026-09-27

Workflow Orchestration Partial

A fixed agent loop that Cartesia runs. A Managed Agent combines Ink-2 listening, the chosen LLM and Sonic speech with turn taking, interruptions and tool calls, and each configuration change is saved as a version; no workflow graph, branching model or hand off between agents is documented for Managed Agents, and the older Line SDK is being migrated away from.

SourceCartesia, docs.cartesia.ai (Managed Agents, Agent configuration, Migrate from the Line SDK)read 2026-09-27

Knowledge Grounding & RAG Full

A maintained index over the customer's documents. Documents uploaded through the Playground or the documents API (single or bulk, up to 100 a request) are indexed automatically on upload, organized into nested folders with metadata for filtering, and an agent's queries search only the folders attached to it; documents can be updated or deleted through the API.

SourceCartesia, docs.cartesia.ai (Knowledge Base, Create Document API)read 2026-09-27

Human Oversight & Guardrails Partial

A human handoff in one place. The transfer system tool hands the caller to a phone number, and turn settings end or check in on silent calls; the configuration's Guardrails are rules written into the agent's instructions, which is prompting, not an enforced control, and no approval step holds a tool call for a person.

SourceCartesia, docs.cartesia.ai (Agent configuration, System tools)read 2026-09-27

Security, Identity & Governance Full

Access and compliance are both documented. On access, Enterprise organizations get self serve SSO over SAML or OIDC, configured by an organization admin, and Admin and Member roles, where only admins manage members, invitations and settings. On compliance, the pricing page FAQ states Cartesia is SOC 2 Type II certified, and a Vanta trust center sits at trust.cartesia.ai; Zero Data Retention is available on Enterprise.

SourceCartesia, docs.cartesia.ai (Set up SSO, Set up an organization) and cartesia.ai/pricingread 2026-09-27

Observability & Auditability Full

Run level records of what each agent did. Every call has a transcript with turn level timestamps and call logs with the agent's logging statements, custom events can be recorded, runtime logs are retrievable through the API, and webhooks deliver call and turn events carrying each turn's tool calls, variable updates and speech latency, plus a post call analysis.

SourceCartesia, docs.cartesia.ai (Observability, Get Call Runtime Logs)read 2026-09-27

Memory & State Persistence Partial

Conversation state lives inside a call. Within a call the agent keeps the conversation, and dynamic variables set at the start or changed by tools are read by the agent each time it responds; nothing carries into the next call except what the customer passes in as variables, and no memory layer with a scope and lifetime is documented.

SourceCartesia, docs.cartesia.ai (Agent configuration, Dynamic variables)read 2026-09-27

Deployment & Data Residency Full

Named regions offered as an option, plus customer environments by arrangement. Enterprise customers can use dedicated regional deployments in the United States, European Union, United Kingdom, India and Australia that keep inference traffic in region; the homepage states the same models and agents run cloud, on premise and on device, and the pricing FAQ routes on premise and VPC deployment through sales.

SourceCartesia, docs.cartesia.ai (Regional endpoints) and cartesia.ai (home, pricing FAQ)read 2026-09-27

Prebuilt Agents, Templates & Packs Partial

Starting templates exist; what they contain is not documented. The agents API lists public, Cartesia provided agent templates to help a customer get started, but no docs page names a template or what job it does, and nothing shows a whole agent a buyer adopts; the voice library, multilingual voices and pinned model snapshots are model assets, not agents.

SourceCartesia, docs.cartesia.ai (List Templates API reference)read 2026-09-27

Triggers & Channel Coverage Full

A scheduled wake plus phone and app channels. A call batch stores up to 5,000 recipients and Cartesia dials them in the background, at a scheduled_at time when one is set, up to a concurrency limit, with retries for failed calls; agents also take inbound and outbound calls on Cartesia or Twilio numbers and SIP trunks, and stream over a WebSocket from web and mobile apps.

SourceCartesia, docs.cartesia.ai (Batch calling, Phone numbers, WebSocket API)read 2026-09-27

Model Flexibility & Routing Full

The customer chooses the model. Each agent's config sets the LLM by ID from the list GET /v1/agents/models returns, with provider, average latency and token prices shown, plus temperature and output limits; changing it creates a new version so candidates can be compared. Cartesia manages the provider accounts, and speech in and out uses Cartesia's own Ink and Sonic models.

SourceCartesia, docs.cartesia.ai (LLMs, Agent configuration)read 2026-09-27

APIs, SDKs & MCP Extensibility Full

A documented API and SDKs for Cartesia's own platform. A versioned REST API (214 reference pages, published OpenAPI and AsyncAPI specs) covers text to speech, speech to text, voices and agents, including agents, tools, knowledge base documents, call batches, metrics and webhooks; official JavaScript and TypeScript and Python SDKs, a CLI, and an MCP server that runs TTS, STT, voices and pronunciation dictionaries sit beside it.

SourceCartesia, docs.cartesia.ai (llms.txt, Client Libraries, MCP)read 2026-09-27

Testing, Debugging & Optimization Partial

Scoring of live calls, without a test harness or gate. Cartesia scores completed calls with built in metrics (call success, speech latency) and custom LLM as a judge metrics the customer writes and assigns to an agent, with results exportable, so output quality can be tracked over time. No fixture or dataset testing before production, quality gate or tuning loop is documented.

SourceCartesia, docs.cartesia.ai (Metrics, Results)read 2026-09-27

Browser & Computer Use Not documented

The DTMF system tool presses phone keys, which is not browser or desktop control, and the browser examples stream voice into a web page, the product running in a browser. No browser, desktop or computer control is documented.

SourceCartesia, docs.cartesia.ai (llms.txt, System tools)read 2026-09-27

The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Recent platform changes

2026-09-23·Agent capabilityVerified

Cartesia introduced Multilingual Voices, a library of more than 50 voices that support up to 25 languages while retaining the same vocal identity. Customers can keep a voice ID across its supported languages and select a locale in speech generation requests. The offering is available through the voice library and API, with language and accent support varying by voice.

Bears on: Agent capability

View source
2026-08-27·Agent capabilityPartially Verified

Cartesia released Sonic 3.6, an updated text-to-speech model that adapts intonation, pacing, and emotiveness to transcript context without SSML tags. The release adds support for Odia and Urdu, bringing the total to 44 supported languages, and improves transcript following for Hindi and Hinglish.

Bears on: Agent capability

View source
2026-08-15·Agent capabilityVerified

Cartesia introduced keyterm prompting and configurable turn detection for its Ink-2 model. Keyterm prompting enhances transcription accuracy for specific vocabulary, while configurable turn detection allows developers to adjust how the model identifies the end of a speaker's turn.

Bears on: Agent capability

View source
View all 4 changes for Cartesia →Tracked since Jul 2026 · Verified from public vendor sources

Pricing

Free plan; Pro $5/mo; Startup $49/mo; Scale $299/mo; Enterprise; usage on credits and agent minutes

credits (characters, seconds) and agent minutes

Free tier

Included quota

Free: 20K credits and $1 of agent usage a month. Pro ($5 a month): 100K credits and $5 of agent usage. Startup ($49): 1.25M credits and $49 of agent usage. Scale ($299): 8M credits and $299 of agent usage, priority support. Enterprise: custom credits, agent usage and concurrency. Unlimited workspace seats on every plan; team invitations start at Startup.

What is public

Plan prices, included credits and agent usage, per credit TTS and STT rates and the per minute agent rate are published; Enterprise is by sales.

Billing mechanics

Monthly plan fee with included credits and prepaid agent usage. TTS charges 1 credit per character, STT 1 credit per second, and Managed Agents $0.06 a minute.

Cost watchouts

Free LLM usage in agents ends on 1 October 2026, after which LLM tokens bill on top of the agent minute rate. Inviting team members requires Startup or higher. Free plan credits are small for production.

Variable cost rationale

Plan fees are small floors; real cost is credit usage on characters and seconds plus per minute agent and telephony charges that scale directly with call volume.

Overage / add-ons

Usage above included credits bills at 1 credit per character (TTS) and 1 credit per second (STT); agents bill $0.06 a minute against prepaid agent usage.

Sales call required

Mixed (some tiers require a call)

Free / trial

Free plan for prototyping with limited credits, no commercial use

Lowest paid plan

Pro $5 a month plus usage

Commercial notes

Independent. States SOC 2 Type II certification; Enterprise adds SSO, regional deployments and zero data retention.

Key ambiguities

LLM token prices for agents after the free period ends on 1 October 2026 are listed per model in the API rather than on the pricing page.

Agentic Index verified 2026-09-27

Alternatives to Cartesia

The closest documented capability profiles to Cartesia among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • Sourcegraph11.0 / 14Fuller documented coverage on Workflow Orchestration
  • Paragon10.0 / 14Fuller documented coverage on Workflow Orchestration
  • Thread AI11.0 / 14Fuller documented coverage on Workflow Orchestration and Human Oversight & Guardrails
  • xpander.ai12.0 / 14Fuller documented coverage on Workflow Orchestration and Human Oversight & Guardrails
  • Composio8.5 / 14A lighter documented profile than Cartesia
  • Kestra12.5 / 14Fuller documented coverage on Workflow Orchestration and Human Oversight & Guardrails

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.