Rime
Rime is text to speech infrastructure for voice agents, not a full agent: Coda and Mist speech models over a streaming API, with pronunciation control and on premise deployment.
Rime builds text to speech models for real time voice agents and IVR, delivered as an API rather than as an agent platform. Its own docs describe it as the speech layer: a voice agent built on LiveKit, Pipecat, Vapi, Daily or the customer's own code sends text to Rime and streams back audio. Coda is the flagship model and the default for new applications, with sub 100 millisecond model latency across nine languages; Mist v3 targets the lowest time to first audio, and Mist v2 keeps inline pronunciation control. Rime's voices are trained on its own studio recorded conversational speech.
Pronunciation is where Rime puts its developer tooling: a coverage endpoint reports which words in a script are missing from its pronunciation dictionary, a phonetic alphabet sets custom pronunciations on Mist models, and a spell function reads codes and IDs letter by letter. The API streams over HTTP, SSE or WebSockets from US West or US East endpoints, and a CLI and a hosted MCP server sit beside it.
Enterprise customers can run licensed Coda or Mist engines on their own NVIDIA hosts or in a VPC, with text and audio kept in the deployment. Rime states SOC 2 Type II and HIPAA compliance, keeps only character counts by default and does not train on customer data without opt in. Teams have Owner and Member roles.
Rime is a speech component: orchestration, knowledge, memory, triggers and tool calling live in whatever agent stack calls it. Pricing is usage based per thousand characters, with free starter minutes and custom Enterprise pricing.
Vendor details
Canonical URL
https://rime.ai
Category
Agent infrastructure
Funding status
Independent. Backed by Unusual Ventures, Cadenza Capital and Founders You Should Know.
Company status
independent
Use cases & customers
Deployment options
Integrations
HTTP, SSE and WebSocket synthesis API, CLI and hosted MCP server. Guides for LiveKit, Pipecat, Daily, Vapi, VideoSDK and Twilio. On premise engines over HTTP, gRPC or WebSocket.
In practice
A voice agent company integrates Rime's Arcana model so its phone agents sound like real people on calls, then deploys it on premises to hit the lowest possible latency for callers.
A national brand handling millions of monthly calls uses Rime as the speech layer of its telephony stack, relying on SpeechQA to catch and fix mispronounced product names before they reach customers.
A regulated enterprise runs Rime inside its own virtual private cloud so customer audio never leaves its environment, using code switching to serve callers in more than one language on a single call.
Sources & related URLs
Agentic Index coverage score
3.5 / 14 capabilities · 25%
| Integrations & Tool Calling | Not documented |
|---|---|
|
No tool calling or action in outside systems. Rime is a text to speech API that voice agent stacks call; its documented integrations (LiveKit, Pipecat, Daily, Vapi, VideoSDK, Twilio) plug Rime's voice into those platforms, which is distribution of the model, not an agent acting. SourceRime, docs.rime.ai (llms.txt, integration guides)read 2026-09-27 |
|
| Workflow Orchestration | Not documented |
|
No workflow model. Each request turns text into speech; the voice agent tutorials put the listen, think and speak loop in the customer's own code or in LiveKit or Pipecat, so orchestration lives in the platform that calls Rime. SourceRime, docs.rime.ai (Build a voice agent, Streaming TTS)read 2026-09-27 |
|
| Knowledge Grounding & RAG | Not documented |
|
No retrieval over the customer's content. The docs cover models, voices, pronunciation, streaming and on premise engines; the pronunciation dictionary shapes how words are spoken, not what the agent knows. SourceRime, docs.rime.ai (llms.txt, Pronunciation control)read 2026-09-27 |
|
| Human Oversight & Guardrails | Not documented |
|
No oversight of agent actions. Rime takes no actions to approve or constrain; pronunciation control and the spell function govern how text is spoken, which is output formatting, not a guardrail on what an agent does. SourceRime, docs.rime.ai (Pronunciation control)read 2026-09-27 |
|
| Security, Identity & Governance | Full |
|
Teams carry Owner and Member roles with a published permission table (only owners invite, edit billing, view all API keys, delete others' keys or remove members). On compliance, Rime holds SOC 2 Type II as of May 2025, with the most recent audit completed March 2026 and the report under NDA, has been HIPAA compliant since February 2024 with a BAA on Enterprise, and publishes a subprocessor list. SourceRime, docs.rime.ai (Privacy and compliance, Teams) and rime.ai/securityread 2026-09-27 |
|
| Observability & Auditability | Partial |
|
Reporting is aggregate, with no per run trace. Character usage shows in the dashboard and the CLI, the CLI measures time to first byte per endpoint, and on premise containers expose health and OpenTelemetry metrics for Prometheus; Rime keeps no request content by default, so there is nothing per request to inspect. SourceRime, docs.rime.ai (Monitoring and usage, on-prem Metrics)read 2026-09-27 |
|
| Memory & State Persistence | Not documented |
|
Synthesis is stateless. Rime states it collects only character counts by default and keeps no customer content, and nothing in the docs holds conversation state; each request stands alone. SourceRime, docs.rime.ai (Privacy and compliance)read 2026-09-27 |
|
| Deployment & Data Residency | Full |
|
A generally available on premise option lets enterprise customers run a licensed Coda or Mist v3 engine on their own NVIDIA hosts with Docker Compose or Kubernetes, with synthesis text and audio staying in the deployment and only aggregate usage counts sent to Rime; the pricing page lists cloud, on premise or VPC. The cloud API offers US West and US East endpoints. SourceRime, docs.rime.ai (On-prem quickstart, Regional endpoints) and rime.ai/pricingread 2026-09-27 |
|
| Prebuilt Agents, Templates & Packs | Not documented |
|
No packaged agents. The voice catalog (287 Coda voices, Mist voices) is a set of model output options, not agents, and the voice agent tutorials for Next.js, Vite, Express, Node and FastAPI, like the MCP tool that generates Pipecat or LiveKit code, are sample code for developers to build from. SourceRime, docs.rime.ai (Build a voice agent, Voices)read 2026-09-27 |
|
| Triggers & Channel Coverage | Not documented |
|
Rime runs only when called. It synthesizes speech when the customer's application sends a request over HTTP or WebSockets and owns no phone numbers, channels, schedules or events; Twilio and the voice platforms supply the calls. SourceRime, docs.rime.ai (Streaming TTS, WebSocket API)read 2026-09-27 |
|
| Model Flexibility & Routing | Not documented |
|
Rime is the only provider. The customer picks among Rime's own speech models (Coda by default, Mist v3 for the lowest latency, Mist v2 for inline pronunciation), all Rime's; no other provider or bring your own model is documented. SourceRime, docs.rime.ai (Models)read 2026-09-27 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Rime documents an API for its own platform. The API reference covers HTTP, SSE and WebSocket synthesis and a dictionary coverage endpoint, with quickstarts in cURL, Python, JavaScript and TypeScript, a Rime CLI, and a hosted MCP server with authenticated tools to list voices, check the dictionary, normalize text and synthesize speech; on premise engines take HTTP, gRPC or WebSocket. SourceRime, docs.rime.ai (API reference, MCP Quickstart, Engine protocols)read 2026-09-27 |
|
| Testing, Debugging & Optimization | Not documented |
|
No evaluation of agent behavior. The coverage endpoint and MCP tool report which words in the customer's script are missing from Rime's pronunciation dictionary, a check on input text rather than on what an agent did, and the latency benchmarks are Rime's own model measurements rather than an evaluation surface for the customer. SourceRime, docs.rime.ai (Coverage, Latency)read 2026-09-27 |
|
| Browser & Computer Use | Not documented |
|
No browser, desktop or computer control; Rime is a speech synthesis API. SourceRime, docs.rime.ai (llms.txt)read 2026-09-27 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Rime deprecated its legacy Arcana text-to-speech model for cloud users, migrating all cloud API traffic to its default Coda model. Coda utilizes an LLM backbone paired with a dedicated speech inference engine trained on full-duplex human conversations, delivering sub-100ms model latency across eight languages. The API request shape remains identical, but the update transitions users to Coda's curated voice lineup.
Bears on: Deployment / data residency
View sourceRime has updated its Coda model API to officially support Arabic and Hindi languages. The language matrix and API reference documentation have been expanded to include these new language parameters for the Coda text-to-speech model.
Bears on: Agent capability
View sourceRime launched Coda, a new dual-decoder text-to-speech model built for real-time enterprise conversations, announced alongside a Series A funding round. The release also introduces a native integration with Together AI that co-locates STT, LLM, and TTS to streamline voice agent pipelines.
Bears on: Agent capability
View sourcePricing
Starter: about 800 free minutes, then $0.03 to $0.05 per 1K characters; Enterprise custom
characters
Included quota
Starter: about 800 free minutes and 20 concurrent generations. Enterprise: unlimited concurrent generations and custom voice clones, SLAs and dedicated support.
What is public
Per 1K character rates by model, the free Starter allowance and concurrency limit are published; Enterprise is custom.
Billing mechanics
Pay as you go per 1,000 characters, priced by model.
Cost watchouts
Costs scale with characters synthesized. On premise, VPC, BAA and unlimited concurrency are Enterprise only. Custom voice clones are ordered through Enterprise.
Overage / add-ons
Per 1,000 characters by model after the free minutes: Mist v3 $0.03, Coda $0.05.
Sales call required
Mixed (some tiers require a call)
Free / trial
About 800 free minutes on Starter, no credit card
Lowest paid plan
Starter pay as you go at $0.03 per 1K characters (Mist v3)
Key ambiguities
rime.ai/llms.txt still says 3,000 free minutes and a $0.05 starting rate, while the pricing page shows about 800 free minutes from $0.03; the pricing page is taken as current.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Rime
The closest documented capability profiles to Rime among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Speechmatics3.5 / 14Matches Rime across all 14 documented capabilities
- AgentOps5.0 / 14Adds documented Integrations & Tool Calling and Model Flexibility & Routing, among others
- Mem06.0 / 14Adds documented Integrations & Tool Calling and Memory & State Persistence, among others
- Unikraft6.0 / 14Adds documented Integrations & Tool Calling and Memory & State Persistence, among others
- Veecle3.0 / 14Adds documented Prebuilt Agents, Templates & Packs and Testing, Debugging & Optimization
- Coral5.5 / 14Adds documented Integrations & Tool Calling and Knowledge Grounding & RAG, among others
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded