W&B Weave
Also known as: Weave, Weights & Biases Weave, W&B Weave
Agent and LLM observability, evaluation, and guardrails from Weights and Biases, with session aware tracing, LLM as a judge scoring, and safety scorers.
W&B Weave is the agent and LLM application observability, evaluation and guardrails product from Weights & Biases, the machine learning experiment tracking company. It is built for multi-turn, multi-agent systems, organizing traces into conversations, turns, LLM calls, tool calls and sub-agents, and it instruments applications through a decorator in the Python and TypeScript SDKs, built-in integrations for agent SDKs, harnesses, LLM providers and frameworks, or any OpenTelemetry pipeline.
In production, Weave's Agents view shows what an agent did, what it cost and where it went wrong; built-in and custom signals score agent turns and tag quality and safety issues; monitors passively score live traffic; and automations trigger actions from monitor metrics and trace activity.
Guardrails run scorers inline and can block or modify a response before it reaches a user, with built-in safety and quality scorers for toxicity, bias, PII and hallucination.
Evaluations run applications against versioned datasets with built-in, local and custom LLM-judge scorers, compare runs to catch regressions and rank models on leaderboards, and annotation queues bring domain experts in to label traces.
Weave runs on W&B's Multi-tenant Cloud, on Dedicated Cloud on AWS, Google Cloud or Azure in the customer's chosen region, or self-managed. Both cloud options are SOC 2 Type II compliant and Dedicated Cloud is HIPAA compliant, with SSO over OIDC, team and project roles, scoped service accounts and SCIM. The Free plan includes 5 GB of storage and 1 GB a month of Weave data ingestion; Pro starts at $60 a month for teams under 50 employees, with usage billed at $0.03 per GB of storage and $0.10 per MB of ingestion beyond the included amounts; Enterprise is custom.
Vendor details
Canonical URL
https://wandb.ai/site/weave/
Category
Agent infrastructure
Subcategory
Observability and evaluation
Funding status
Product of Weights and Biases, founded 2018, an MLOps standard with more than a million users. Weights and Biases was acquired by CoreWeave in a roughly $1.7 billion deal that closed in May 2025, folding it into CoreWeave's GPU cloud with an interoperability pledge.
Company status
acquired
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
Traces any function through a simple decorator, with native integrations for OpenAI, Anthropic, Google, LangChain, LlamaIndex, and more than fifty frameworks. Connects to coding agents like Claude Code for continuous iteration. Exposes an SDK and API, with SOC 2, HIPAA, and ISO compliance for enterprise deployments.
In practice
Your agent works in testing but reliability collapses in production. You add Weave with a decorator, get session and turn level traces, and use built in signals to surface the failure modes offline evals missed.
You are changing prompts and models and cannot tell what regressed. Weave's evaluation framework with datasets and LLM as a judge scorers measures every change so wins in one area do not quietly break another.
Compliance needs safety enforced on outputs. Weave Guardrails apply prebuilt scorers for toxicity, bias, PII, and hallucination, and produce the audit trail your governance team needs.
Sources & related URLs
Agentic Index coverage score
9.0 / 14 capabilities · 64%
| Integrations & Tool Calling | Partial |
|---|---|
|
Integrations with LLM providers, agent SDKs and harnesses such as the OpenAI Agents SDK, Claude Agent SDK and Google ADK, and frameworks such as LangChain, CrewAI and LlamaIndex let Weave trace calls, and it traces activity between MCP clients and servers, while alerts go out by Slack and email. These bring trace data in and send notifications out; none connects an agent to a real system to take actions. SourceWeights & Biases, docs.wandb.ai Weave index (integrations overview, agent integrations, MCP) and wandb.ai/site/pricingread 2026-09-21 |
|
| Workflow Orchestration | Not documented |
|
Weave traces the customer's own multi-step and multi-agent workflows, including sub-agent delegations, and automations trigger actions from monitor metrics and trace activity. No workflows that sequence, branch or retry steps, or combine deterministic nodes with agent steps, are documented as something Weave runs. SourceWeights & Biases, docs.wandb.ai Weave index (trace sub-agents, automations)read 2026-09-21 |
|
| Knowledge Grounding & RAG | Not documented |
|
RAG applications can be evaluated with LLM judges, and Weave traces the context a customer's retrieval returns. No document ingestion, index, retrieval layer or knowledge API that grounds an agent in company data is documented. SourceWeights & Biases, docs.wandb.ai Weave index (evaluate RAG applications)read 2026-09-21 |
|
| Human Oversight & Guardrails | Full |
|
Guardrails run inline scorers on a user's input or a model's output in real time, before outputs reach users, and can block or modify a response when a score crosses a threshold, with built-in and custom scorers and AWS Bedrock Guardrails as examples; every guardrail result is stored as a monitor. Annotation queues route traces to domain experts for structured feedback. SourceWeights & Biases, docs.wandb.ai/weave/guides/evaluation/guardrails and the Weave index (annotation queues)read 2026-09-21 |
|
| Security, Identity & Governance | Full |
|
Users sign in with SSO through Google, GitHub or enterprise providers such as Okta and Azure Active Directory over OIDC, are organized into teams and projects with a restricted visibility scope, get role-based access at team or project level, automate with scoped service accounts, and are managed through the SCIM API; Multi-tenant Cloud and Dedicated Cloud are both SOC 2 Type II compliant with reports on the W&B Security Portal, Dedicated Cloud is HIPAA compliant, and the pricing page lists audit logs and customizable data retention. SourceWeights & Biases, docs.wandb.ai/weave/guides/platform and wandb.ai/site/pricingread 2026-09-21 |
|
| Observability & Auditability | Full |
|
The Agents view shows conversations, turns, LLM calls, tool calls and sub-agent delegations with their cost, the trace view follows nested execution paths, calls can be filtered and exported through the Python SDK, REST API or UI, PII can be redacted from traces, and the pricing page lists audit logs and customizable data retention. SourceWeights & Biases, docs.wandb.ai Weave index (agents view, trace tree, query and export calls, redact PII) and wandb.ai/site/weave and pricingread 2026-09-21 |
|
| Memory & State Persistence | Not documented |
|
Conversations, threads and turns of the customer's agent are recorded, and datasets, prompts and models are versioned as tracked objects. No session, workflow or long-term memory that an agent reads and writes is documented. SourceWeights & Biases, docs.wandb.ai Weave index (trace threads, track and version objects)read 2026-09-21 |
|
| Deployment & Data Residency | Full |
|
Weave runs on W&B Multi-tenant Cloud in a North America region of W&B's Google Cloud account, on W&B Dedicated Cloud on AWS, Google Cloud or Azure with data in a dedicated ClickHouse cluster in the cloud and region of the customer's choice, with IP allowlisting and private connectivity, or on Self-Managed instances the customer hosts; W&B fully manages the two cloud options. SourceWeights & Biases, docs.wandb.ai/weave/guides/platformread 2026-09-21 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
W&B Skills install into a coding agent to teach it to train models, build agents and analyze experiments on the W&B platform, a packaged, ready-to-install starting point for another agent. No ready-made agents, packaged employees or deployable workflows for a buyer's own work are documented for Weave; the Claude Code, Codex and OpenClaw plugins trace sessions and are integrations, not packaged agents. SourceWeights & Biases, docs.wandb.ai/llms.txt (W&B Skills) and the Weave index, and wandb.ai/site/weaveread 2026-09-21 |
|
| Triggers & Channel Coverage | Full |
|
Monitors passively score production traffic, built-in signals score agent turns and tag quality and safety issues as they arrive, and automations trigger actions from monitor metrics and trace activity, with Slack and email alerts on the Pro plan. Scoring and actions start on incoming traffic without a person asking. SourceWeights & Biases, docs.wandb.ai Weave index (monitors, signals, automations) and wandb.ai/site/pricingread 2026-09-21 |
|
| Model Flexibility & Routing | Full |
|
Weave's documentation index describes connecting OpenAI-compatible inference endpoints to a project as custom runtimes, managed from the UI or the Python and TypeScript SDKs, comparing models in the Playground and Evaluation Playground, and running open source models through W&B Serverless Inference. Customers choose models and bring their own endpoints. SourceWeights & Biases, docs.wandb.ai Weave index (custom runtimes, Playground, Serverless Inference)read 2026-09-21 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
Weave publishes Python and TypeScript SDKs, a service REST API with an OpenAPI description, and an OpenTelemetry endpoint that accepts traces from any pipeline without the SDK; the W&B MCP server lets an IDE or agent query W&B data and documentation, W&B Skills teach coding agents to use the platform, and the Pro plan adds CI/CD automations. SourceWeights & Biases, docs.wandb.ai/llms.txt and the Weave index (OpenTelemetry, query and export calls), and wandb.ai/site/pricingread 2026-09-21 |
|
| Testing, Debugging & Optimization | Full |
|
Evaluations run applications against versioned datasets with built-in, local and custom scorers, including LLM judges, for single-turn and multi-turn agents; evaluations are compared side by side to spot regressions and ranked on leaderboards, monitors passively score production traffic, and the Pro plan adds CI/CD automations. SourceWeights & Biases, docs.wandb.ai Weave index (evaluations, scorers, compare evaluations, monitors) and wandb.ai/site/weave and pricingread 2026-09-21 |
|
| Browser & Computer Use | Not documented |
|
The Weave documentation covers tracing, evaluation, monitoring, guardrails, annotation and deployment, and no browser, desktop or computer control by an agent is documented; the CoreWeave Sandboxes listed beside Weave are isolated environments for running agents, and they execute code rather than control a browser, desktop or computer. SourceWeights & Biases, docs.wandb.ai Weave index and wandb.ai/site/weave and pricing; wandb.ai/site/pricingread 2026-09-21 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
W&B Server v0.85.0 adds links in both directions between Weave calls and the agent spans they produce or invoke. The same release adds a configurable process wide query parallelism budget, allowing deployments to reduce peak memory use when Weave workloads load large text heavy tables by trading off request concurrency.
Bears on: Observability / auditability
View sourceW&B Weave introduced a new set of features encompassing expanded integrations, user feedback collection, and detailed cost calculations. This update enables developers to capture user feedback on agent interactions and track the financial costs associated with LLM calls.
Bears on: Observability / auditability
View sourcePricing
Free tier · Pro from $60/mo + usage · Enterprise custom
Monthly plan fee plus usage meters for storage and Weave data ingestion
Included quota
Free tier for individuals and small teams. Pro adds team features and higher limits with usage meters for storage ($0.03 a gigabyte) and Weave trace ingestion ($0.10 a megabyte), plus per token serverless inference. Free academic Pro includes up to 25 gigabytes a month of Weave ingestion and up to 100 seats.
What is public
The usage meters (storage $0.03 a gigabyte, Weave ingestion $0.10 a megabyte) and the free and academic tiers are clearly published. The exact Pro plan fee is stated differently across recent sources.
Billing mechanics
A monthly plan fee acts as the floor, with three usage meters layered on: storage at $0.03 a gigabyte, Weave trace ingestion at $0.10 a megabyte, and per token serverless inference. Self hosting removes cloud storage limits.
Cost watchouts
Weave data ingestion beyond the included amount is billed at $0.10 per MB and storage at $0.03 per GB, so trace-heavy applications push cost past the plan fee; Pro is for early-stage teams under 50 employees, and larger customers must move to Enterprise.
Variable cost rationale
The Pro fee is a predictable floor, but trace ingestion and storage meters scale with application volume, so heavy GenAI workloads push most of the cost into metered overage.
Overage / add-ons
Weave ingestion, inference, and storage are billed monthly in arrears based on usage over the last 30 days, above included plan quotas.
Sales call required
Mixed (some tiers require a call)
Free / trial
Free plan (5 GB storage, 1 GB a month of Weave data ingestion); 30-day free trial of Pro; free for academic research
Lowest paid plan
Pro from $60 a month plus usage meters
Commercial notes
Weave is Weights & Biases' product for tracing, evaluating and monitoring agents and LLM applications, priced on the shared W&B plans with Weave-specific storage and ingestion meters. Academic research is free.
Key ambiguities
Enterprise pricing is not published, and the pricing page notes volume discounts on ingestion with annual commitments.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to W&B Weave
The closest documented capability profiles to W&B Weave among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- F5 AI Guardrails9.0 / 14Matches W&B Weave across all 14 documented capabilities
- Galileo9.0 / 14Matches W&B Weave across all 14 documented capabilitiesW&B Weave vs Galileo →
- Opik9.0 / 14Matches W&B Weave across all 14 documented capabilities
- Confident AI8.5 / 14A lighter documented profile than W&B Weave
- Fiddler AI8.5 / 14A lighter documented profile than W&B Weave
- HoneyHive8.5 / 14A lighter documented profile than W&B Weave
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded