Back to vendors
Q

QA Wolf

Also known as: QA Wolf, QAWolf, QA Wolf Inc

Visit site
Entry pricePlatform 1¢ per AI credit + 15¢ per runner minute · managed service by quoteFull pricing detail

End-to-end testing platform whose agent maps an app and writes Playwright and Appium tests from prompts, run in parallel on deploy, schedule or pull request, sold self-serve by usage or as a managed service with guaranteed coverage.

QA Wolf is an end-to-end testing platform and managed service for web and mobile applications. On the platform, its agent explores an application with computer-use models, maps its workflows into a coverage outline, and turns plain-language prompts into deterministic Playwright and Appium tests that the customer owns and can export at any time. It can work unattended from a prompt or be guided by a person, reads linked documentation and team skill files, and diagnoses and repairs broken tests after failures.

Tests run in parallel on QA Wolf's infrastructure across browsers, Android emulators and real iPhones and iPads, triggered on deploy through GitHub, GitLab or a webhook, on a schedule, or on pull requests where Smart Smoke Suites pick the affected flows. Run Rules sequence flows and their dependencies, and flows can pass data between each other.

Judge-model assertions check AI-generated responses, and tests can validate MCP servers, email and SMS, phone audio and visual diffs. Results sync to Jira, Linear and test management tools, and a REST API, CI SDK, CLI and an MCP server for coding agents open the platform to other tools. QA Wolf states it is SOC 2 Type II compliant and supports SSO over SAML and OpenID Connect.

QA Wolf sells two ways: a self-serve Platform at 1 cent per AI credit and 15 cents per runner minute, which can be tried for free, and Coverage as a Service, where QA Wolf's own team builds, runs, investigates and maintains the suite with guaranteed coverage, zero flakes and human-verified bug reports, priced per test under management.

Vendor details

Canonical URL

https://www.qawolf.com

Category

Agent infrastructure

Subcategory

Agentic Coverage as a Service E2E testing with human in the loop

Funding status

QA Wolf is independent and headquartered in Seattle. It was founded in 2019 by CEO Jon Perl, a former Zipdrug engineer. It raised a thirty six million dollar Series B in July 2024, led by Scale Venture Partners with participation from Threshold Ventures, Ventureforgood, and existing investors Inspired Capital and Notation Capital. Total funding is reported at roughly fifty seven million dollars across rounds. It serves more than one hundred thirty customers, including Salesloft, Drata and AutoTrader.ca.

Company status

independent

Use cases & customers

Primary use cases

Guaranteed 80 percent plus maintained E2E test coverageAI generated Playwright and Appium tests for web and mobileCI integrated pre merge smoke and regression testingEvaluating non deterministic AI outputs with LLM as a judge

Target customers

Engineering and QA teams needing guaranteed coverageTeams that want to avoid in house QA headcountEnterprises testing web, mobile, and Salesforce appsTeams testing AI features with LLM as a judge

Deployment options

Cloud

Integrations

Integrates directly with the CI pipeline so tests run pre merge on pull request branches and developers ship at the speed of AI, and covers web, iOS, Android, Electron, and Salesforce applications. Agents ingest videos of user flows along with DOM snapshots and browser logs to generate open source Playwright code for web and Appium code for mobile, which the customer owns.

In practice

A team's releases keep getting stuck in QA. QA Wolf's agents map the app and generate Playwright and Appium tests while its engineers maintain the suite, delivering eighty percent coverage in weeks and pre merge smoke tests in CI.

A flaky test would normally page an engineer at 2am. QA Wolf's Zero Flake Guarantee means a human QA engineer verifies every failure before the dev team is ever alerted.

A team ships an AI feature with non deterministic output. QA Wolf uses LLM as a judge assertions to evaluate it, extending E2E testing to the agentic parts of the app.

Agentic Index coverage score

10.0 / 14 capabilities · 71%

Integrations & Tool Calling Full

Bug reports flow into Jira and Linear, where QA Wolf automatically creates, syncs and closes issue tickets, and results sync to test management tools (TestRail, Xray, Zephyr, Qase and Testmo). Deployments are heard from GitHub and GitLab, notices go to Slack and Microsoft Teams, and webhooks reach any other bug tracker. Its tests can call APIs, seed databases, toggle feature flags and use email inboxes. Writes into the customer's issue trackers mean access runs both ways, and webhooks and helper libraries extend it to other tools.

SourceQA Wolf, docs.qawolf.com Jira, Linear, test management integrations and webhooksread 2026-09-21

Workflow Orchestration Full

Run Rules sequence the order of flows and their dependencies in a suite, and flows pass authentication tokens, user IDs and other values between each other in a coordinated run. Tags build custom suites, runs execute in parallel with automatic reruns of failures, and triggers choose which flows run on each deploy or schedule.

SourceQA Wolf, docs.qawolf.com Run Rules, pass data between flows, tags and triggers, and qawolf.com/pricingread 2026-09-21

Knowledge Grounding & RAG Partial

Before it builds or explores, the agent reads documentation linked in a prompt, test plans, product requirements and help docs, and it applies team conventions written into custom skill files in every session. Company knowledge enters as linked pages and skill files the agent reads per session. No index, retrieval layer or refresh of the customer's sources is documented.

SourceQA Wolf, docs.qawolf.com test automation, coverage mapping and custom skillsread 2026-09-21

Human Oversight & Guardrails Partial

Tests can be created unattended from a prompt, or the agent can work with a person who guides it, and a person can take over the browser during mapping while the agent watches. The generated tests are standard code the team reviews and versions. Choosing guided or unattended working sets an autonomy mode, but no approval step that holds the agent's action for a person is documented.

SourceQA Wolf, docs.qawolf.com test automation and coverage mappingread 2026-09-21

Security, Identity & Governance Full

QA Wolf states it is SOC 2 Type II compliant, verified through independent audits, with HIPAA ready infrastructure. Customers can set up single sign-on over SAML 2.0 and OpenID Connect, and credentials are stored securely for use in tests. That pairs a held attestation with sign in through the customer's identity provider.

SourceQA Wolf, docs.qawolf.com why QA Wolf (security and data protection) and single sign-onread 2026-09-21

Observability & Auditability Partial

Each run of the customer's tests is recorded with results, the verdict reached when evaluating each deployment for a trigger is kept, and a person can watch the agent open a browser and explore during mapping. A step by step record of the agent's own prompts, tool calls and decisions is not described.

SourceQA Wolf, docs.qawolf.com triggers and diagnosing a trigger, coverage mapping, and qawolf.com/pricingread 2026-09-21

Memory & State Persistence Not documented

The coverage map, the test code and run history persist, and skill files the team writes apply in every agent session. No session, workflow or long term memory that the agent itself writes and reuses is documented.

SourceQA Wolf, docs.qawolf.com custom skills and coverage mappingread 2026-09-21

Deployment & Data Residency Partial

Tests run on QA Wolf's own managed infrastructure, with a shared, allowlisted or private iOS device pool, local runs through its CLI, and Playwright tests customers can export at any time. No deployment of the platform in the customer's environment or choice of data region is documented, so where data is kept is not stated.

SourceQA Wolf, docs.qawolf.com iOS device pools, CLI local execution, and qawolf.com/pricingread 2026-09-21

Prebuilt Agents, Templates & Packs Not documented

QA Wolf offers solution recipes and helper libraries for testing scenarios such as emails, barcodes, video injection and web performance. No catalog of ready made agents, templates or packaged workflows a buyer adopts for its own work is documented.

SourceQA Wolf, docs.qawolf.com solutionsread 2026-09-21

Triggers & Channel Coverage Full

Triggers run chosen flows whenever an application deploys, reported through GitHub or GitLab deployments or a webhook, or on a schedule. Smart Smoke Suites automatically select the flows affected by each pull request, and the automation agent diagnoses broken tests and refactors test code after failures. Work starts on a deploy, a schedule or a pull request without a person asking.

SourceQA Wolf, docs.qawolf.com triggers, Smart Smoke Suites, and qawolf.com/llms.txtread 2026-09-21

Model Flexibility & Routing Full

The judge model is the customer's choice in the non-deterministic AI assertions, with ChatGPT documented as a drop in replacement with the same signature and return shape. They are designed for apps that use different AI providers, so the customer makes the model choice in its own tests.

SourceQA Wolf, docs.qawolf.com non-deterministic AI assertionsread 2026-09-21

APIs, SDKs & MCP Extensibility Full

QA Wolf publishes a REST API giving direct HTTP access to core platform actions, a TypeScript CI SDK for triggering and reading runs, a CLI for running flows locally, flow and test libraries with API references, and deployment webhooks. An MCP server lets coding agents such as Claude Code and Codex build and run the suite. Together these give a stable API and SDK, a served MCP interface and a fit into CI/CD.

SourceQA Wolf, docs.qawolf.com REST API overview, CI SDK reference, CLI and QA Wolf MCPread 2026-09-21

Testing, Debugging & Optimization Full

An application's end to end workflows are mapped into a coverage outline, and the agent generates deterministic Playwright and Appium tests from plain language prompts and runs them in parallel on deploy, on a schedule or on pull requests. Judge model assertions return a structured verdict with hard and soft thresholds on a customer's AI generated responses, and tests validate MCP connections and tool execution. Testing happens before release, as a gate in CI and continuously after.

SourceQA Wolf, docs.qawolf.com why QA Wolf, test automation, non-deterministic AI assertions and triggersread 2026-09-21

Browser & Computer Use Full

Computer use models let the agent open a browser and explore each feature of the customer's application, documenting the workflows it finds. A person can take control of the browser while the agent watches, and tests run in real browsers and on Android emulators and real iPhones and iPads. That is documented control of a real interface.

SourceQA Wolf, docs.qawolf.com coverage mapping and why QA Wolf, and qawolf.com/pricingread 2026-09-21

The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Recent platform changes

2026-09-29·MCP / tool calling / APIVerified

QA Wolf made its MCP server generally available for Codex, Claude Code and other MCP compatible coding agents. A coding agent can now ask QA Wolf's agents to find existing test coverage, write or update tests, run them, investigate failures and file bug reports with video, traces and logs. The QA Wolf CLI offers the same capabilities when local files or local runs are needed.

Bears on: MCP / tool calling / API

View source
View all 1 change for QA Wolf →Tracked since Sep 2026 · Verified from public vendor sources

Pricing

Platform 1¢ per AI credit + 15¢ per runner minute · managed service by quote

Platform: $0.01 per AI credit and $0.15 per runner minute, no seat fees; Coverage as a Service: priced by the number of tests under management

Trial available

What is public

The Platform's rates are public, at 1 cent per AI credit and 15 cents per runner minute after a free trial. Coverage as a Service is public in shape only, priced per test under management with everything included, so costs are locked to the managed test count, and its rates are by quote.

Billing mechanics

The Platform is self serve and billed by usage, per AI credit and per runner minute. Coverage as a Service is sales led and outcome based, billed per test for the number of tests QA Wolf manages and bundling investigation, maintenance, setup, and cleanup, so spend is locked to the managed test count rather than hourly labor.

Cost watchouts

Runner minutes accrue with every parallel test run, so large suites run on each deploy add up; the managed service scales with the number of tests under management.

Variable cost rationale

Platform cost is pure consumption, rising with AI credits used and runner minutes, so suite size and run frequency drive the bill; the managed service scales with tests under management.

Additional watchouts

On Coverage as a Service, billing scales with the number of managed tests, so growing coverage raises cost. Confirm the per test rate and any minimums directly, and account for the multi week onboarding ramp.

Sales call required

Mixed (some tiers require a call)

Free / trial

The Platform can be tried for free

Lowest paid plan

Platform pay as you go: 1¢ per AI credit, 15¢ per runner minute

Commercial notes

Two ways to buy: a self-serve, usage-based Platform where the customer's team automates and maintains its own tests, and Coverage as a Service where QA Wolf's team builds, runs, investigates and maintains the suite for a quote per test under management.

Key ambiguities

How many AI credits typical exploration, creation and maintenance tasks consume is not stated; Coverage as a Service rates are by quote. Platform browsers are Chrome, Firefox and WebKit, while iOS, Android and Electron are listed under the managed service.

Missing data

Per test rates and volume tiers for Coverage as a Service are not published.

Agentic Index verified 2026-09-21

Alternatives to QA Wolf

The closest documented capability profiles to QA Wolf among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • Anchor Browser11.0 / 14Adds documented Prebuilt Agents, Templates & Packs
  • Paragon10.0 / 14Fuller documented coverage on Knowledge Grounding & RAG and Observability & Auditability
  • Pipedream11.0 / 14Adds documented Memory & State Persistence
  • Apify10.5 / 14Adds documented Memory & State Persistence and Prebuilt Agents, Templates & Packs
  • BotGauge7.5 / 14Fuller documented coverage on Observability & AuditabilityQA Wolf vs BotGauge →
  • Inworld AI8.5 / 14Adds documented Memory & State Persistence

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.