Back to vendors
V

Vellum

Also known as: Vellum Workflows, Vellum Agent Builder

Visit site
Entry priceVellum assistant: free start; Mighty $30/mo, Super $100/mo, Ultra $200/mo; the workflows and agent-builder platform is no longer priced on vellum.aiFull pricing detail

Agent builder and evaluation platform with workflow orchestration, observability environments, model flexibility, and a free-tier entry, for product teams and AI developers building and managing production agents.

Vellum is an end-to-end development platform for building, evaluating, deploying, and monitoring LLM applications and AI agents. It sits in the LLMOps and agent-orchestration space, giving product and engineering teams a single place to take an AI feature from a rough idea to a reliable production system, rather than re-plumbing everything each time a new model or requirement appears. A defining goal is letting technical and non-technical teammates collaborate on the same workflow.

The platform centers on a few connected surfaces. A prompt playground lets teams write and compare prompts side by side across many models, with version control so prompts can change without touching application code. A visual workflow builder turns an AI system into a graph whose nodes are model calls, Python or TypeScript code, conditional branches, API calls, and retrieval steps, which makes the order of operations, bottlenecks, and failure modes far easier to see and debug as systems get more agentic. An evaluation framework adds test datasets, custom metrics, and language-model-judge scoring so teams can assert on output quality and validate changes before they go live.

Vellum has leaned hard into agents. Its agent builder exposes an Agent Node that connects to tools, including code, sub-workflows, third-party integrations, and Model Context Protocol servers, with function-calling schemas generated automatically, and a 2026 release lets teams describe an agent in plain language and have Vellum assemble it. Vellum is model-agnostic across major providers, so teams can swap models to balance cost and quality, and it also exposes its own MCP server so coding assistants like Claude Code and Cursor can work with it.

Once built, workflows deploy to an API with one click, with bidirectional sync between the visual editor and code so updates ship without redeploying the surrounding app. Production monitoring then tracks execution logs, latency, and cost so teams can see what their agents actually do in the wild. For regulated buyers, Vellum carries SOC 2 and HIPAA compliance with private-cloud deployment options, and it offers both visual tooling and Python and TypeScript SDKs for code-first teams.

Vendor details

Canonical URL

https://www.vellum.ai

Category

Agent builder

Subcategory

LLMOps and agent-building platform with graph workflows, test suites and metric-based evaluation

Company status

independent

Use cases & customers

Primary use cases

Building agent workflows as an explicit graph of model calls, retrieval, code, tool calls and conditional branchesScoring prompt and model changes against test suites and custom metrics before release, and continuing to score them in productionGrounding agents on a persistent Document Index with metadata-filtered retrieval, queryable from workflows or an external Search APIDeploying workflows to an API with environments, release tags and release reviews governing what reaches productionRunning the platform self-hosted, with static IPs and configurable data retention for customers whose requirements exceed managed SaaS

Target customers

product teamsAI developers

Deployment options

SaaS

Integrations

Vellum reaches external systems through generic primitives rather than a connector catalog: an API Node making HTTP requests to any endpoint, a Code Execution Node running Python or TypeScript, and an Agent Node handling tool calling with automatic function-calling schema generation and loop logic. Named third-party integrations in the Product documentation are Datadog for exporting execution telemetry and a worked Zapier and Airtable workflow example. A Document Index holds uploaded material with metadata filtering, queryable from a Search Node or externally through a Search API. Model coverage spans major providers with custom models supported and fallback models documented as a common architecture. Outbound surfaces are a deployed workflow API, SDKs and developer tools under a top-level Developers section, webhook export of execution data, execution URLs addressing individual runs, and HMAC authentication for verifying callback origin. Self-hosting is documented as a top-level section.

In practice

A new model drops every few weeks and you want to know if switching helps. Vellum's playground compares your prompts side by side across providers, and its evaluation framework scores the outputs so you can decide on evidence, not vibes.

Your agent's logic has grown into a tangle that's hard to debug. Vellum lets you model it as a visual graph of model calls, code, branches, and tool steps, making the order of operations and failure points easy to see.

Your product managers and engineers keep stepping on each other building AI features. Vellum gives both a shared workflow, visual for non-engineers and code via SDKs for developers, with one-click deployment to an API.

Agentic Index coverage score

11.0 / 14 capabilities · 79%

Integrations & Tool Calling Full

The Agent Node documents tool calling with automatic function-calling schema generation and loop logic, with tools drawn from custom code, subworkflows and API calls, a dedicated page on function calling with chat models, and a worked example of parallelised function calls. The API Node makes authenticated HTTP requests to any endpoint and the Code Execution Node runs custom Python or TypeScript, and a Zapier and Airtable example and a Datadog export are documented. Documented custom tool support that lets an agent act in outside systems is Full.

Sourcedocs.vellum.ai Agent Node, API Node, Code Execution Node and function-calling pages, 31 August read, re-gradedread 2026-09-29

Workflow Orchestration Full

Fifteen documented node types cover an Agent Node with automatic schema handling and loop logic, prompt and prompt-deployment nodes, templating, a search node, an API node, a code execution node running Python or TypeScript, subworkflow, map, guardrail, conditional, merge, final output, error and note nodes. Node adornments apply retry and try semantics to existing nodes.

Advanced documentation covers long-running workflows and batching executions. Multi-agent composition appears through subworkflows and worked examples including multi-agent content creation and LLMs debating each other, and six common architectures are documented covering RAG, escalation to a human, prompt retry, PDF summarization and fallback models.

Sourcedocs.vellum.ai workflows nodes overview and per-node pages, node-adornments, advanced long-running-workflows and batching-executions, common-architectures and examples pagesread 2026-08-31

Knowledge Grounding & RAG Full

A Documents section documents uploading documents into a Document Index, integrating with a Search API, and metadata filtering for scoped retrieval. A dedicated Search Node searches against a Document Index and is described by the vendor as suited to RAG.

A RAG system appears as one of the documented common architectures, with worked examples covering a basic RAG chatbot, a RAG chatbot with Cohere rerank, and a building-a-RAG-chatbot tutorial. Evaluating RAG pipelines has its own documentation page under Evaluation and Test Suites. Documents persist as an index between runs rather than being assembled per request.

Sourcedocs.vellum.ai documents section, workflows nodes search-node, common-architectures rag-system, workflow examples and evaluation/evaluating-rag-pipelines pagesread 2026-08-31

Human Oversight & Guardrails Partial

The platform documentation offers a Guardrail Node that runs an inline evaluation against a pre-defined Metric inside a workflow, Error and Conditional nodes that stop or branch execution, Retry and Try adornments, release reviews before a change is promoted to a live deployment, role-based access control, and an Escalation to a Human architecture that automatically routes sensitive or complex messages to human operators.

Those are runtime constraints, a design-time release step and workflow-initiated escalation. The documentation index shows no approval node or pause for a person to approve an agent's action before it executes.

Sourcedocs.vellum.ai documentation index, Escalation to a Human, Guardrail Node and release reviews pagesread 2026-09-29

Security, Identity & Governance Full

The vendor's data privacy and storage page states that Vellum maintains SOC 2 Type 2 compliance and is HIPAA compliant, with security practices regularly audited to meet industry standards and healthcare data protection requirements. All data stored in Vellum, including documents in Document Indexes, is encrypted with AES-256 GCM in transit and at rest.

The vendor states it does not send interactions or feedback to LLM providers for training, and that Completion Actuals submitted through the feedback API are stored for the customer's own quality monitoring rather than used to train or fine-tune models. Documented controls include role-based access control, HMAC authentication for verifying request origin, static IPs for customer-side allowlisting, organization access management and configurable data retention policies.

Sourcedocs.vellum.ai security section covering data-privacy-and-storage, rbac-permissions, hmac-authentication and static-ips, and organizations manage-access and pagesread 2026-08-31

Observability & Auditability Full

Observability in production is documented as its own capability, alongside monitoring of production trends, tracking of workflow execution costs, and execution URLs addressing individual runs. A Datadog integration and a webhook integration export execution data to the customer's own monitoring and alerting systems. Online evaluations score production traffic on defined metrics. Deployment lifecycle management, environments and release tags tie each execution to a specific released version. Data retention policies are documented separately at organization level.

Sourcedocs.vellum.ai deployments observability and deployment-lifecycle-management, monitoring section covering production-trends, execution-cost-tracking, datadog, webhooks and execution-urls, evaluation online-evaluations, and organizations pagesread 2026-08-31

Memory & State Persistence Partial

A Document Index persists uploaded material between runs and is queryable through a Search Node or Search API. Long-running workflows are documented for executions that extend beyond a single request, and deployment lifecycle management, environments and release tags persist configuration and version history across releases. No memory module, conversation store, session identity or cross-execution context capability appears anywhere in the Product documentation navigation, and workflows are documented as invoked per execution through the deployed API.

Sourcedocs.vellum.ai workflows advanced long-running-workflows, documents section, deployments section and Product navigation indexread 2026-08-31

Deployment & Data Residency Full

Self-hosting is a top-level documentation section alongside Home, Product and Developers, with its own getting-started introduction. Static IPs are documented under both Security and advanced workflow behavior, providing fixed egress addresses for customer-side allowlisting. Data retention policies are configurable at organization level, and a data privacy and storage page documents how data is held. Managed deployment covers deployment lifecycle management, environments and release tags for promoting versions.

Sourcedocs.vellum.ai self-hosting introduction, security static-ips and data-privacy-and-storage, workflows advanced static-ips, organizations and deployments pagesread 2026-08-31

Prebuilt Agents, Templates & Packs Partial

Twelve named workflow examples are documented covering prompt chaining, basic RAG chatbot, RAG with Cohere rerank, customer support bot, PDF to CSV, summarizing images of websites, parallelized function calls, conference attendee lookup, multi-agent content creation, LLMs debating each other, a Zapier and Airtable integration and automating PR reviews.

Six common architectures cover RAG systems, escalation to a human, prompt retry logic, PDF content summarization and fallback models. Out-of-the-box metrics are supplied by the vendor and reusable across test suites, and subworkflows let customers make their own work reusable. Whether examples are importable into a workspace in one action, rather than followed as documentation, is not stated.

Sourcedocs.vellum.ai workflow examples overview, common-architectures, metrics out-of-the-box-metrics and reusing-metrics, and workflows nodes subworkflow-node pagesread 2026-08-31

Triggers & Channel Coverage Partial

Workflows are invoked programmatically through a deployed API, with a dedicated integrating page covering calling them from application code, plus batching executions for volume and long-running workflows for extended jobs. The webhook integration documented under Monitoring pushes execution data outward to customer systems rather than accepting inbound events to start a workflow. No scheduled, cron, event-driven or chat-channel trigger appears anywhere in the Product documentation navigation.

Sourcedocs.vellum.ai workflows api-integration, advanced batching-executions and long-running-workflows, monitoring webhooks and execution-urls pages, and Product navigation indexread 2026-08-31

Model Flexibility & Routing Full

A Custom Models page documents bringing models beyond the built-in roster, and prompt engineering documentation covers comparing prompts across models. Fallback models are documented as one of Vellum's common workflow architectures. Prompt caching and multimodality have their own pages under Prompts.

Model choice is coupled to measurement through test suites and metrics, so switching providers can be evaluated quantitatively rather than by inspection. Prompt deployment nodes execute deployed prompts within workflows, and release tags and environments govern which prompt version, and therefore which model configuration, is live.

Sourcedocs.vellum.ai prompts custom-models, prompt-engineering, prompt-caching and multimodality, workflows common-architectures fallback-models, and deployments pagesread 2026-08-31

APIs, SDKs & MCP Extensibility Full

A Developers section sits at the top level of the documentation covering SDKs, APIs and developer tools. Workflows deploy to an API with a dedicated page on integrating them into application code, and a Search API reaches the Document Index independently of workflows. Webhook integration exports execution data, and execution URLs address individual runs. A Code Execution Node runs customer Python or TypeScript within a workflow, and HMAC authentication verifies request origin for callbacks.

Sourcedocs.vellum.ai developers overview, workflows api-integration, documents api-integration, monitoring webhooks and execution-urls, workflows nodes code-execution-node and security hmac-authentication pagesread 2026-08-31

Testing, Debugging & Optimization Full

Quantitative evaluation is documented through Test Suites and Metrics, with the vendor stating that prompt changes, parameter adjustments and model switches make regression likely and that quantitative evaluation is the answer. Metrics are documented as out-of-the-box, custom, and reusable across test suites.

Online evaluations extend scoring to production, and evaluating RAG pipelines is documented separately. A Guardrail Node runs an inline evaluation using a pre-defined Metric within a workflow. Experimentation and a prompt playground support side-by-side comparison across models, and Retry and Try node adornments handle failure paths during development.

Sourcedocs.vellum.ai evaluation section covering quantitative-evaluation, online-evaluations and evaluating-rag-pipelines, metrics section, and workflows nodes guardrail-node and node-adornments pagesread 2026-08-31

Browser & Computer Use Not documented

The Product documentation enumerates fifteen node types and none provides browser, desktop or computer-use capability: the external-action nodes are an API Node making HTTP requests to an endpoint, a Code Execution Node running custom Python or TypeScript, and a Search Node querying a Document Index, with tool calling handled by the Agent Node through automatic schema generation.

The nearest example, summarizing images of websites, applies retrieval and vision to fetched material rather than navigating or operating a site. No browser automation, session control, form filling or screen interaction appears anywhere in the Product documentation navigation.

Sourcedocs.vellum.ai workflows nodes overview and per-node pages, workflow examples summarize-images-of-websites, and Product navigation indexread 2026-08-31

The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Recent platform changes

2026-06-19·Major release / new capabilitiesVerified

Vellum moved to wide release with plugins as a first-class marketplace, a new Advisor capability that pulls in a second more powerful model on hard problems, Memory v3 at section grain for sharper recall, and upgrades to subagents, workflows, Slack, and the Activity page.

Bears on: Agent capability

View source
View all 1 change for Vellum →Tracked since Jun 2026 · Verified from public vendor sources

Pricing

Vellum assistant: free start; Mighty $30/mo, Super $100/mo, Ultra $200/mo; the workflows and agent-builder platform is no longer priced on vellum.ai

hybrid (subscription machine and storage tiers plus model tokens at cost)

Free tierTrial available

Included quota

Free: about 50 prompt executions and 25 workflow executions per day, 1 user seat, core toolset (playground, workflow builder, RAG, evaluation). Pro: higher compute and storage via machine tiers, custom subdomain, more seats. Enterprise: custom contracts, BAA and DPA, SSO, audit logs, data retention policies, VPC deployment, dedicated support.

What is public

The pricing page lists the Vellum personal assistant's plans (free start, Mighty $30, Super $100, Ultra $200, custom) and free open-source self-hosting. The workflows platform's Pro and Enterprise pricing is no longer published.

Billing mechanics

Assistant plans pair a compute and storage tier with included usage; credits are $1 each for model responses, web searches and image generation beyond the bundle, and a $10 monthly base fee applies to Super and Ultra.

Cost watchouts

The jump from free to Pro is steep, daily execution caps on the free plan are easy to hit in real testing, heavy RAG or high volume workflow runs raise model token spend, and compliance features (BAA, SSO, VPC, HIPAA) require Enterprise.

Variable cost rationale

A subscription base plus model tokens passed through at cost; token spend scales with how much your workflows and agents run, but Vellum adds no markup on model usage, so exposure is moderate rather than steep.

Additional watchouts

The pricing page now sells the Vellum personal assistant, so pricing for the workflows and agent-builder platform has to be confirmed with the vendor directly.

Overage / add-ons

Model tokens are passed through at cost as you run prompts and workflows; compute and storage are set by the machine and storage tier you pick and resizable anytime, prorated.

Sales call required

Mixed (some tiers require a call)

Free / trial

Free tier, no card: about 50 prompt and 25 workflow executions per day, 1 seat

Lowest paid plan

Mighty at $30 a month (Vellum assistant)

Commercial notes

End-to-end platform for building, evaluating, deploying, and monitoring LLM apps and agents (prompt playground, visual workflow builder, Python SDK, evals, RAG, observability). Founded 2023 (Y Combinator), $29.5M total funding including a $20M Series A in 2025. Supports OpenAI, Anthropic, Google, Cohere, and self-hosted models. SOC 2 Type II, HIPAA workflows, and VPC deployment on Enterprise.

Key ambiguities

The pricing page now prices the Vellum personal assistant; pricing for the workflows and agent-builder platform documented at docs.vellum.ai is not published.

Missing data

Pricing for the workflows and agent-builder platform documented at docs.vellum.ai.

Agentic Index verified 2026-09-29

Alternatives to Vellum

The closest documented capability profiles to Vellum among agent builders tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • Retool11.5 / 14Fuller documented coverage on Triggers & Channel Coverage
  • Joget11.0 / 14Fuller documented coverage on Human Oversight & Guardrails
  • Lyzr12.0 / 14Fuller documented coverage on Memory & State Persistence and Prebuilt Agents, Templates & Packs
  • Moveo.AI12.0 / 14Fuller documented coverage on Prebuilt Agents, Templates & Packs and Triggers & Channel Coverage
  • SnapLogic12.0 / 14Fuller documented coverage on Prebuilt Agents, Templates & Packs and Triggers & Channel Coverage
  • Activepieces11.5 / 14Fuller documented coverage on Human Oversight & Guardrails and Triggers & Channel Coverage

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Head to head

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.