Modal
Also known as: Modal Labs
Serverless cloud for AI workloads, with per second GPU billing, isolated sandboxes for untrusted agent code, scheduled functions and model endpoints.
Modal is a serverless cloud for AI workloads: model inference and training, batch jobs and sandboxed code execution. Developers define functions in Python, JavaScript or Go with the GPU, image, memory and concurrency they need, and Modal builds the container, schedules it, scales it from zero and bills per second.
For agents, Modal Sandboxes run untrusted code in isolated environments created programmatically, with networking controls, snapshots of filesystem and memory, and VM Sandboxes in beta; Restricted Functions are a second way to run untrusted code. Scheduled Functions, web endpoints and job queues start work without a person, and retries, queues and fan-out handle scale. The Modal Library serves models on Shared Endpoints billed per token or Dedicated Endpoints with the customer's own capacity and region, and any open model can be deployed on GPUs with vLLM or SGLang. Its examples include coding agents, a computer-use agent over VNC and parallel evals.
Workloads can be pinned to a region, with a data residency guide and a price multiplier for pinning. Modal states a SOC 2 Type II audit and HIPAA support, and offers Okta, Entra and SAML SSO, SCIM, role-based access control and, on Enterprise, audit logs, with logs and metrics exportable to OpenTelemetry providers and Datadog.
Pricing is per second usage on top of a plan: Starter is free with 30 dollars a month of compute, Team is 250 dollars a month with 100 dollars of compute, and Enterprise is custom. Sandbox CPU costs about three times the function rate, and network egress is billed from 1 October 2026.
Vendor details
Canonical URL
https://modal.com
Category
Agent infrastructure
Subcategory
Compute and sandboxes
Funding status
Independent.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
Python, JavaScript/TypeScript and Go SDKs and a CLI; Functions, Sandboxes, Endpoints, Scheduled Functions, job queues, Volumes, Dicts and Queues; OIDC to external clouds, cloud bucket mounts, OpenTelemetry and Datadog export, Okta, Entra and SAML SSO, SCIM (beta) and Slack notifications (beta).
In practice
Your agent needs to run untrusted code it generated. Modal Sandboxes spin up isolated environments programmatically with any dependency and scale back to zero when done.
You want serverless GPUs without DevOps. You declare a GPU function in Python, and Modal builds the container, schedules it on an H100 and bills per second.
Your agent should check a data source every morning without anyone starting it. You deploy it as a Modal Scheduled Function that runs on a cron schedule and writes its results to a Volume.
Sources & related URLs
Agentic Index coverage score
8.0 / 14 capabilities · 57%
| Integrations & Tool Calling | Partial |
|---|---|
|
Code on Modal reaches other systems through what the customer writes: OIDC tokens authenticate functions to external clouds, cloud buckets mount as filesystems, secrets hold credentials, Slack notifications report on apps (beta), and examples connect Discord, Google Sheets and Tailscale. No catalog of connectors, OAuth actions or tool gateway of Modal's own is documented; the tools are the customer's code. Sourcemodal.com/llms.txtread 2026-09-22 |
|
| Workflow Orchestration | Partial |
|
Retries are first class on Modal: a Retries policy can sit on any function, alongside job queues, fan out with map and spawn, batch and dynamic batching, and gang scheduled clusters, so a pipeline that mixes model calls with deterministic steps can be built on it. Sequencing and branching live in the customer's Python, and no multistep workflow primitive that records completed steps and resumes after failure is documented. Sourcemodal.com/docs/guide/retriesread 2026-09-22 |
|
| Knowledge Grounding & RAG | Not documented |
|
Retrieval systems the customer builds can run on Modal (its examples embed documents with TEI and run a RAG chatbot over PDFs), but Modal does not maintain a retrieval structure over the customer's content that an agent queries. Volumes, Dicts and mounted buckets are storage, not a Modal index a query returns, and nothing documented lets the product itself index, refresh or permission sources. Sourcemodal.com/llms.txtread 2026-09-22 |
|
| Human Oversight & Guardrails | Not documented |
|
No approval step, consent checkpoint or escalation to a person is documented for work running on Modal. Restricted Functions come closest: they stop untrusted code from touching Modal resources, calling other functions or reaching Modal's APIs. Together with sandbox networking controls, they are isolation the developer configures, not a point where a person reviews an agent's action. Sourcemodal.com/docs/guide/restricted-accessread 2026-09-22 |
|
| Security, Identity & Governance | Full |
|
Modal states a SOC 2 Type II audit, with the report available through its Security Portal at trust.modal.com, HIPAA support, and external penetration testing. Identity runs through SSO with Okta, Microsoft Entra or custom SAML, with SCIM provisioning (beta). Access is managed with role based access control, user groups and service users, and an append only audit log records sensitive actions on Enterprise. Secrets, Restricted Functions that cannot touch Modal resources, and sandbox networking controls round out the controls. Sourcemodal.com/docs/guide/securityread 2026-09-22 |
|
| Observability & Auditability | Partial |
|
The workload itself is well recorded: real time logs and metrics per function and container, GPU metrics, function stats, an append only audit log of sensitive workspace actions on Enterprise, and export of audit logs, function logs and container metrics to any OpenTelemetry provider or Datadog. Audit stays separate from runtime. Prompts, tool calls, retrieved knowledge and outputs can be inspected step by step only as far as the customer's own code logs them; no tracing of each call an agent makes is documented. Sourcemodal.com/docs/guide/otel-integrationread 2026-09-22 |
|
| Memory & State Persistence | Partial |
|
State carries across runs: sandbox snapshots save and restore a sandbox's filesystem and memory, memory snapshots speed cold starts, volumes persist files, and Dicts and Queues hold data shared across function calls. The agent's code reads that state to carry on. Volumes and Dicts are storage the code reads, though, and restored process state belongs to the machine; the buyer can resume or delete it but not review or edit it as memory. No memory layer with a stated scope and lifetime is documented. Sourcemodal.com/docs/guide/sandbox-snapshotsread 2026-09-22 |
|
| Deployment & Data Residency | Full |
|
Workloads can be pinned to a region. Functions and Sandboxes take a region argument for the container region, dedicated Endpoints set both container and routing regions, and a data residency guide sets out per service what can be pinned and when customer data may leave the region. Region selection carries a stated price multiplier (1.15x broad, 1.75x narrow). Modal is a managed cloud with no self hosted option, and shared Endpoints do not take a region. Sourcemodal.com/docs/guide/data-residencyread 2026-09-22 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
Ready made starting points come as a gallery of runnable examples and a Library of models deployable as Shared or Dedicated Endpoints. The agent examples include Cursor cloud agents, a background coding agent with OpenCode, a coding platform, a Claude Agent SDK Slack bot, a LangGraph agent in a GPU sandbox, a computer use agent over VNC and parallel evals. These are examples and deployable models a developer builds from, not a browsable catalog of agents a buyer adopts. Sourcemodal.com/llms.txtread 2026-09-22 |
|
| Triggers & Channel Coverage | Full |
|
Work starts without a person. Scheduled Functions run on a cron expression or a fixed period, web functions and endpoints boot a container from zero when a request arrives, and job queues hand work to functions as it lands. Sourcemodal.com/docs/guide/cronread 2026-09-22 |
|
| Model Flexibility & Routing | Full |
|
Model choice is the customer's throughout: Modal serves models from its Library on Shared Endpoints billed per token or Dedicated Endpoints with the customer's capacity and region, and the customer can deploy any open model on GPUs with vLLM or SGLang, as its examples show for DeepSeek, Nemotron and others. Sourcemodal.com/docs/guide/shared-endpointsread 2026-09-22 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Official SDKs in Python, JavaScript/TypeScript and Go (the latter two in beta) let outside code call Modal, alongside a full CLI that includes a skills command for coding agents. Web functions and endpoints expose the customer's code as HTTP APIs, and a gRPC API sits behind the open source client. Sourcemodal.com/llms.txtread 2026-09-22 |
|
| Testing, Debugging & Optimization | Not documented |
|
Evaluations run on Modal, but scoring them is left to other tools. Its example runs massively parallel evals with Harbor in Modal Sandboxes, and its reinforcement learning examples use verl and TRL, with the datasets, graders and scoring belonging to those frameworks. No fixtures, scoring, quality gates or comparison of agent runs are documented as Modal features. Sourcemodal.com/llms.txtread 2026-09-22 |
|
| Browser & Computer Use | Partial |
|
Computer use agents can run on Modal's machines. Its example runs a Browser Use agent driving Chromium inside a Modal Sandbox, watched over VNC, with the model served from a Modal Endpoint, and VM Sandboxes are in beta. Control of the browser comes from Browser Use, a separate library, and no Modal API for screenshots, clicks or typing is documented: Modal supplies the environment, not the control. Sourcemodal.com/docs/examples/computer_use_vncread 2026-09-22 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Modal Sandboxes can now run on a full virtual machine instead of the gVisor container runtime, set with one option in the Python and JavaScript SDKs. Modal says the VM runtime handles workloads such as running Docker inside a sandbox, while gVisor stays the default for now.
Bears on: Deployment / data residency
View sourcePricing
Free Starter ($30/mo credits); Team $250/mo; Enterprise; per second usage
per second compute (GPU, CPU, memory)
Included quota
Starter: $30 a month of compute, 3 seats, 100 containers, 10 concurrent GPUs. Team at $250 a month: $100 a month of compute, unlimited seats, 5,000 containers, 50 concurrent GPUs. Usage is billed per second on top of the included compute.
What is public
Modal publishes its plans, per second GPU, CPU and memory rates, sandbox rates, Volume storage pricing, region multipliers and the October 2026 egress schedule.
Billing mechanics
Active seconds multiplied by per second rate, plus the plan fee, with regional multipliers of 1.25x to 2.5x and a non preemption multiplier of about 3x where applicable. Scale to zero means no charge during idle time. No egress fees on base rates.
Cost watchouts
Sandbox CPU is billed at about three times the function rate, which compounds for agents that keep sandboxes alive; pinning a region adds 1.15x to 1.75x; Volumes bill $0.09 per GiB-month after the first TiB; and from 1 October 2026 network egress beyond the included amount bills per GiB ($0.04 on Starter after 1 TiB).
Variable cost rationale
There is a small plan floor, but cost is almost entirely per second compute that scales with GPU, CPU, and memory usage, amplified by regional and non preemption multipliers on sandbox workloads.
Additional watchouts
Budget sandbox time at the higher sandbox CPU rate, check region multipliers before pinning a region, and plan for egress charges from October 2026.
Overage / add-ons
Per second usage across GPU, CPU and memory accrues above the plan's included compute, with region multipliers where a region is pinned; Volumes and (from 1 Oct 2026) egress bill beyond their included amounts.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free Starter tier with $30 a month in credits
Lowest paid plan
Team $250 a month plus per second usage
Commercial notes
$87 million Series B in September 2025 at about a $1.1 billion valuation. Customers include Suno, Substack, Ramp, Notion, and Lovable. Roughly a thousand paying customers.
Key ambiguities
Total cost depends on GPU type, CPU and memory time, sandbox use, region pinning, storage beyond 1 TiB and, from October 2026, egress; Enterprise terms are custom.
Missing data
Enterprise pricing is custom; egress allowances beyond Starter were not captured.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Modal
The closest documented capability profiles to Modal among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Beam7.5 / 14Fuller documented coverage on Workflow OrchestrationModal vs Beam →
- E2B8.5 / 14Fuller documented coverage on Integrations & Tool Calling and Browser & Computer UseModal vs E2B →
- Exa6.5 / 14A lighter documented profile than Modal
- Daytona8.0 / 14Fuller documented coverage on Observability & Auditability and Browser & Computer UseModal vs Daytona →
- Unikraft6.0 / 14A lighter documented profile than Modal
- Composio8.5 / 14Adds documented Human Oversight & Guardrails
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded