Back to vendors
V

Vellum

Visit site
Entry priceFree tier (no card); Pro self serve, reported around $500/mo with machine and storage tiers; Enterprise custom; model tokens passed through at costFull pricing detail

Agent builder and evaluation platform with workflow orchestration, observability environments, model flexibility, and a free-tier entry, for product teams and AI developers building and managing production agents.

Vellum is an end-to-end development platform for building, evaluating, deploying, and monitoring LLM applications and AI agents. It sits in the LLMOps and agent-orchestration space, giving product and engineering teams a single place to take an AI feature from a rough idea to a reliable production system, rather than re-plumbing everything each time a new model or requirement appears. A defining goal is letting technical and non-technical teammates collaborate on the same workflow.

The platform centers on a few connected surfaces. A prompt playground lets teams write and compare prompts side by side across many models, with version control so prompts can change without touching application code. A visual workflow builder turns an AI system into a graph whose nodes are model calls, Python or TypeScript code, conditional branches, API calls, and retrieval steps, which makes the order of operations, bottlenecks, and failure modes far easier to see and debug as systems get more agentic. An evaluation framework adds test datasets, custom metrics, and language-model-judge scoring so teams can assert on output quality and validate changes before they go live.

Vellum has leaned hard into agents. Its agent builder exposes an Agent Node that connects to tools, including code, sub-workflows, third-party integrations, and Model Context Protocol servers, with function-calling schemas generated automatically, and a 2026 release lets teams describe an agent in plain language and have Vellum assemble it. Vellum is model-agnostic across major providers, so teams can swap models to balance cost and quality, and it also exposes its own MCP server so coding assistants like Claude Code and Cursor can work with it.

Once built, workflows deploy to an API with one click, with bidirectional sync between the visual editor and code so updates ship without redeploying the surrounding app. Production monitoring then tracks execution logs, latency, and cost so teams can see what their agents actually do in the wild. For regulated buyers, Vellum carries SOC 2 and HIPAA compliance with private-cloud deployment options, and it offers both visual tooling and Python and TypeScript SDKs for code-first teams.

Vendor details

Canonical URL

https://www.vellum.ai

Category

Agent builder

Company status

independent

Use cases & customers

Target customers

product teamsAI developers

Deployment options

SaaS

In practice

A new model drops every few weeks and you want to know if switching helps. Vellum's playground compares your prompts side by side across providers, and its evaluation framework scores the outputs so you can decide on evidence, not vibes.

Your agent's logic has grown into a tangle that's hard to debug. Vellum lets you model it as a visual graph of model calls, code, branches, and tool steps, making the order of operations and failure points easy to see.

Your product managers and engineers keep stepping on each other building AI features. Vellum gives both a shared workflow, visual for non-engineers and code via SDKs for developers, with one-click deployment to an API.

Agentic Index coverage score

10.5 / 14 capabilities · 75%

Integrations & Tool CallingOfficial docs 2026-06-08 Full
Workflow OrchestrationAgent workflow nodes docs 2026-06-08 Full
Knowledge Grounding & RAGOfficial docs 2026-06-08 Partial
Human Oversight & GuardrailsDeployment environments docs 2026-06-08 Full
Security, Identity & GovernanceOfficial docs 2026-06-08 Partial
Observability & AuditabilityObservability deployment docs 2026-06-08 Full
Memory & State PersistenceOfficial docs 2026-06-08 Partial
Deployment & Data ResidencyProduction management docs 2026-06-08 Full
Prebuilt Agents, Templates & PacksOfficial docs 2026-06-08 Partial
Triggers & Channel CoverageOfficial docs 2026-06-08 Partial
Model Flexibility & RoutingMulti-model docs 2026-06-08 Full
APIs, SDKs & MCP ExtensibilityOfficial docs 2026-06-08 Full
Testing, Debugging & OptimizationEval suite docs 2026-06-08 Full
Browser & Computer UseOfficial docs 2026-06-08 Unable to verify

The Agentic Index coverage score grades every vendor Full, Partial or Unable to verify against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Recent platform changes

2026-06-19·Major release / new capabilitiesVerified

Vellum moved to wide release with plugins as a first-class marketplace, a new Advisor capability that pulls in a second more powerful model on hard problems, Memory v3 at section grain for sharper recall, and upgrades to subagents, workflows, Slack, and the Activity page.

Bears on: Agent capability

View source
View all 1 change for Vellum →Tracked since Jun 2026 · Verified from public vendor sources

Pricing

Free tier (no card); Pro self serve, reported around $500/mo with machine and storage tiers; Enterprise custom; model tokens passed through at cost

hybrid (subscription machine and storage tiers plus model tokens at cost)

Free tierTrial available

Included quota

Free: about 50 prompt executions and 25 workflow executions per day, 1 user seat, core toolset (playground, workflow builder, RAG, evaluation). Pro: higher compute and storage via machine tiers, custom subdomain, more seats. Enterprise: custom contracts, BAA and DPA, SSO, audit logs, data retention policies, VPC deployment, dedicated support.

What is public

The free tier and its daily execution caps are public, and Pro is self serve via Stripe with machine and storage tiers; the exact Pro price is not clearly listed (reviewers report around $500/mo) and Enterprise is custom.

Billing mechanics

Three tiers: a free plan (about 50 prompt and 25 workflow executions per day, 1 seat, full core toolset), Pro (self serve via Stripe, choose a compute machine and storage tier, prorated changes, reported around $500/mo), and Enterprise (custom). Model provider token costs are passed through at cost with no markup, so a dollar of credits is a dollar to the model provider.

Cost watchouts

The jump from free to Pro is steep, daily execution caps on the free plan are easy to hit in real testing, heavy RAG or high volume workflow runs raise model token spend, and compliance features (BAA, SSO, VPC, HIPAA) require Enterprise.

Variable cost rationale

A subscription base plus model tokens passed through at cost; token spend scales with how much your workflows and agents run, but Vellum adds no markup on model usage, so exposure is moderate rather than steep.

Additional watchouts

Confirm the current Pro price and machine and storage tiers directly, since the public page is sales led and the free to Pro gap is large with nothing in between.

Overage / add-ons

Model tokens are passed through at cost as you run prompts and workflows; compute and storage are set by the machine and storage tier you pick and resizable anytime, prorated.

Sales call required

Mixed (some tiers require a call)

Free / trial

Free tier, no card: about 50 prompt and 25 workflow executions per day, 1 seat

Lowest paid plan

Pro, self serve via Stripe, reported around $500/mo

Commercial notes

End to end platform for building, evaluating, deploying, and monitoring LLM apps and agents (prompt playground, visual workflow builder, Python SDK, evals, RAG, observability). Founded 2023 (Y Combinator), $29.5M total funding including a $20M Series A in 2025. Supports OpenAI, Anthropic, Google, Cohere, and self hosted models. SOC 2 Type II, HIPAA workflows, and VPC deployment on Enterprise.

Key ambiguities

Vellum's public pricing page is opaque on exact numbers; the Pro price is widely reported around $500/mo but is not clearly listed, and Enterprise is quote only. Pricing changes frequently, so confirm on the live page.

Agentic Index verified 2026-07-07

Alternatives to Vellum

The closest documented capability profiles to Vellum among agent builders tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • OutSystems12.0 / 14Fuller documented coverage on Knowledge Grounding & RAG and Security, Identity & Governance
  • Altilia11.5 / 14Fuller documented coverage on Knowledge Grounding & RAG and Security, Identity & Governance
  • AutoGPT10.5 / 14Fuller documented coverage on Prebuilt Agents, Templates & Packs and Triggers & Channel Coverage
  • Dify9.5 / 14Fuller documented coverage on Knowledge Grounding & RAG
  • Flowise11.5 / 14Fuller documented coverage on Knowledge Grounding & RAG and Memory & State Persistence
  • Gumloop11.5 / 14Fuller documented coverage on Security, Identity & Governance and Prebuilt Agents, Templates & Packs

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Head to head

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.