Firecrawl
Web data API for AI agents that searches, scrapes, crawls and acts on web pages through hosted browsers, with an open source core.
Firecrawl is a web data API for AI agents. It searches the web, scrapes pages into clean markdown or structured JSON, crawls whole sites, maps a site's URLs, parses PDFs and Office documents, and lets an agent act inside a page through a hosted browser. The output is shaped for a model's context, stripping navigation and boilerplate.
The endpoints cover the path from finding to using web data. Search returns full page content rather than links. Scrape turns any URL into markdown or JSON, with change tracking between scrapes. Crawl and Map traverse a site. Interact continues on a scraped page by prompt or Playwright code, and Browser Sandbox opens standalone sessions where an agent fills forms, clicks and signs in, with persistent profiles that keep a login between runs. Monitor checks pages, sites or searches on a schedule and notifies a webhook, email or Slack when something changes. The /agent endpoint, a research preview, gathers data from a prompt without known URLs.
Firecrawl ships SDKs in nine languages, a CLI, an MCP server and skills for coding agents. The core is open source under the AGPL 3.0 license and can be self hosted with Docker Compose; Agent, Browser, Interact, the dashboard and enterprise controls stay on the cloud service. Enterprise adds SOC 2 Type II documentation, SSO and SCIM, IP and key restrictions, Threat Protection policies on every fetched URL, SIEM audit logging and Zero Data Retention.
Pricing is a credit based subscription. A free tier gives 1,000 credits a month, then Hobby is sixteen dollars a month billed yearly (nineteen monthly), Standard eighty three, Growth three hundred thirty three and Scale five hundred ninety nine, with Enterprise custom. A basic scrape costs one credit per page, while structured formats, search, browser minutes and monitor checks cost more, and error pages still cost a credit.
Vendor details
Canonical URL
https://www.firecrawl.dev
Category
Agent infrastructure
Subcategory
Web scraping and crawling
Funding status
Independent. The core is open source under the AGPL 3.0 license (github.com/firecrawl/firecrawl), offered as a managed cloud service and self hostable with Docker. SOC 2 Type II per its enterprise documentation.
Company status
independent
Use cases & customers
Primary use cases
Target customers
Deployment options
Integrations
REST API with OpenAPI specifications, SDKs for Python, Node, Go, Java, Ruby, Rust, .NET, PHP and Elixir, a CLI, an MCP server (keyless, OAuth or API key) and skills for Claude Code and Cursor. Listed integrations include LangChain, LlamaIndex, Mastra, the Vercel AI SDK, n8n, Make, Zapier and Stripe Projects; SIEM audit events stream to Microsoft Sentinel.
In practice
Your cold outbound needs a personalized first line per prospect. You scrape each company's homepage and about page with Firecrawl, feed the clean markdown to an LLM, and generate specific opening lines at scale.
Your RAG knowledge base is fed by brittle scrapers that break on JavaScript heavy sites. You replace them with Firecrawl, which renders the page and returns clean markdown ready for ingestion.
You need to know when a competitor changes its pricing page. You set a Firecrawl Monitor on the page with a daily schedule and get a webhook or Slack alert only when the page changes.
Sources & related URLs
Related / legacy domains
Research sources
Agentic Index coverage score
8.5 / 14 capabilities · 61%
| Integrations & Tool Calling | Partial |
|---|---|
|
An agent gets the web as tools: search, scrape, crawl, map, parse, monitor and Firecrawl's research agent, through its API, SDKs, CLI and an MCP server, and it is listed with LangChain, LlamaIndex, n8n, Make and Zapier. It connects to websites, not to the customer's other systems: there are no native connectors that act in a CRM, ticketing or messaging tool, and no per user credentials to scope, rotate or revoke for them. Acting inside a web page happens through the hosted browser rather than through connectors. Sourcefirecrawl.dev/llms.txtread 2026-09-22 |
|
| Workflow Orchestration | Not documented |
|
No workflow runtime for the customer is documented. The /agent endpoint searches, navigates and gathers data on its own from a prompt, optional URLs and a schema, but that pipeline is Firecrawl's own ready-made agent; the customer cannot define, branch or sequence its steps. Crawl, batch scrape and monitors are jobs Firecrawl runs to a fixed pattern. Nothing lets the customer sequence, branch, retry or route work while mixing deterministic nodes with agent steps; the customer's own logic lives in its framework or in tools such as n8n. Sourcedocs.firecrawl.dev/features/agentread 2026-09-22 |
|
| Knowledge Grounding & RAG | Not documented |
|
The public web is what Firecrawl searches, scrapes and crawls, and its Research and Developer indexes cover public papers and developer artifacts; none of it is a structure Firecrawl maintains over the customer's own documents. Parse turns a customer's PDFs and Office files into markdown and JSON for the customer to use, which is extraction, not an index that persists and stays queryable. Nothing documents Firecrawl indexing, refreshing or permissioning customer content. Clean output feeding the customer's own RAG layer is that layer's grounding, not Firecrawl's. Sourcefirecrawl.dev/llms.txtread 2026-09-22 |
|
| Human Oversight & Guardrails | Partial |
|
Threat Protection lets an organization set a policy, once, that every URL a request would fetch is checked against before Firecrawl fetches it: a scrape target, a search result, a link found during a crawl, or an agent's starting URL. The policy takes a custom blocklist and allowlist of domains, blocked top-level domains, a risk score threshold and a fail-closed setting, can be driven by the organization's own Zscaler categories, and can be locked so no request weakens it. That is a structured guardrail and policy constraint, enforced before the action. No approval step, checkpoint or escalation to a person is documented. It is an enterprise feature enabled per organization. Sourcedocs.firecrawl.dev/features/threat-protectionread 2026-09-22 |
|
| Security, Identity & Governance | Full |
|
SOC 2 Type II, independently audited, is listed in Firecrawl's enterprise documentation, and its security page carries the SOC 2 Type 2 mark. The customer-facing controls are documented beside it: SAML and OIDC single sign-on through Okta, Entra ID or Google Workspace, SCIM provisioning and deprovisioning, team roles, API keys restricted to IP ranges or locked to specific endpoints and formats, Zero Data Retention for scrape and search, and PII redaction, with DPAs and custom contracts. The enterprise controls are provisioned per organization. Sourcedocs.firecrawl.dev/enterpriseread 2026-09-22 |
|
| Observability & Auditability | Full |
|
What the jobs and agent did is recorded one step at a time: SIEM Audit Logging streams one structured event for every URL Firecrawl fetches on the customer's behalf, whether from a scrape, crawl, batch, search, extract, parse or agent run, with the target, status, outcome (including fetches blocked by policy), the API key and workflow that caused it, a request identifier that groups a whole crawl or agent run, and the customer's correlation IDs, delivered to Microsoft Sentinel. The dashboard's Activity Logs show request history with status, duration and credits. That gives an audit trail separate from the runtime, exported to a SIEM, covering each tool call. Page content is not logged, and retention by plan is not documented. SIEM logging is an enterprise feature enabled per organization. Sourcedocs.firecrawl.dev/features/siemread 2026-09-22 |
|
| Memory & State Persistence | Not documented |
|
No memory the agent reads to decide is documented. The closest thing is the saved login: a scrape can save its browser session to a named profile that later scrapes and Interact sessions load, and Browser Sandbox sessions can persist. Cookies, local storage and an authenticated profile are loaded by the browser when a session starts, so the product applies them as settings rather than the agent reading them as context. No context persisted across a run, conversation or workflow and no longer term memory layer is documented, and nothing lets memories be reviewed, edited, deleted or scoped. Change tracking's comparison with the previous scrape is likewise applied as a rule. Sourcedocs.firecrawl.dev/features/interactread 2026-09-22 |
|
| Deployment & Data Residency | Full |
|
Customers can use the managed cloud service or run Firecrawl on their own infrastructure: the core is open source under the AGPL 3.0 license and ships a pinned Docker Compose self-hosting guide, and the docs set out what each mode includes. Self-hosting exposes the core scraping APIs with the customer owning authentication, TLS, persistence, monitoring and upgrades and connecting its own OpenAI-compatible provider or Ollama for LLM formats, while Agent, Browser, Interact, the dashboard and enterprise controls are Cloud only. No region choice for the cloud service is documented. Sourcedocs.firecrawl.dev/contributing/open-source-or-cloudread 2026-09-22 |
|
| Prebuilt Agents, Templates & Packs | Partial |
|
A single prebuilt agent ships: /agent searches, navigates and gathers structured data from a prompt and optional schema, also offered in the dashboard with no code; Firecrawl labels it a Research Preview in early access, and it is callable by any account. Firecrawl also publishes skills for Claude Code, Cursor and other coding agents covering Search, Scrape, Crawl and Interact. These are vendor made working assets for one job, web data gathering. One agent in preview and a skill set are not a browsable catalog of workflows, templates or role specific agents a buyer adopts. Sourcedocs.firecrawl.dev/features/agentread 2026-09-22 |
|
| Triggers & Channel Coverage | Full |
|
Monitor runs recurring checks on a schedule the customer sets, against named URLs, a scheduled crawl of a whole site, or a web search, with an optional plain-language goal; each check labels every page same, new, changed, removed or error, and notifications go to a webhook per page, a webhook per completed check, an email summary or a Slack channel. Work starts on a schedule with no person initiating it. Crawl, batch and agent jobs also post lifecycle events to the customer's webhook. Sourcefirecrawl.dev/llms.txtread 2026-09-22 |
|
| Model Flexibility & Routing | Full |
|
Web data goes to whatever model the customer's agent runs, through LangChain, LlamaIndex, Mastra, the Vercel AI SDK and any MCP client, and a self-hosted deployment connects its own OpenAI-compatible provider or Ollama for LLM-backed formats. On the cloud service its own extraction and /agent models are Firecrawl's, with a choice among its agent models. Model choice sits with the customer for its own agent, and for extraction when self hosted. Sourcedocs.firecrawl.dev/contributing/open-source-or-cloudread 2026-09-22 |
|
| APIs, SDKs & MCP Extensibility | Full |
|
A documented REST API makes Firecrawl callable from outside, with published OpenAPI specifications for v2 and for webhooks, official SDKs for Python, Node, Go, Java, Ruby, Rust, .NET, PHP and Elixir, a CLI, and an MCP server that works keyless, with OAuth sign-in or with an API key, plus skills for coding agents. Sourcedocs.firecrawl.dev/llms.txtread 2026-09-22 |
|
| Testing, Debugging & Optimization | Not documented |
|
Firecrawl's Playground lets a developer try a request and inspect the raw response, and the Activity Logs' Debug issue button runs Firecrawl's support agent on a failed job to suggest a fix; both test and debug Firecrawl calls, not the customer's agent. No fixtures, test datasets, scoring or quality gates for the customer's agent workflows are documented. Sourcefirecrawl.dev/llms.txtread 2026-09-22 |
|
| Browser & Computer Use | Full |
|
Firecrawl hosts real browsers an agent controls: Browser Sandbox gives a secure, isolated session where agents fill out forms, click buttons and authenticate, with Playwright and agent-browser preinstalled, a CDP URL and a live view, and Interact continues on any scraped page by prompt or by Playwright code for multi-step flows such as signing in and clicking through pagination. That is documented control of a real interface, managed and sandboxed remotely. Browser Sandbox and Interact are Cloud only. Sourcedocs.firecrawl.dev/features/browserread 2026-09-22 |
|
The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded
Recent platform changes
Firecrawl launched Alexandria, a common interface for agents to discover and retrieve information from official data providers, custom connectors and its Research, Developer and Government indexes. Agents can query providers and use specialized tools alongside live web search through Firecrawl.
Bears on: MCP / tool calling / API
View sourceFirecrawl launched the Developer Index, a specialized search index covering over 70 million developer artifacts including READMEs, documentation, issues, pull requests, and OpenAPI specs. The index features semantic retrieval and metadata filters and is accessible via the /search/developer API endpoint, CLI, MCP, SDKs, and a companion skill.
Bears on: Agent capability
View sourceFirecrawl upgraded its /search API endpoint with a custom relevance model that evaluates and scores every paragraph, list, and table on retrieved web pages against the user's query. Rather than returning full page content, the API now extracts and returns only the most relevant structural excerpts.
Bears on: Browser/computer use
View sourcePricing
From $16/mo billed yearly · free tier + open source
Credits per API request (1 credit per page for scrape, crawl, map, monitor), within monthly subscription tiers
Included quota
Free 1,000 credits a month. Hobby 5,000 credits, Standard 100,000, Growth 500,000, Scale 1,000,000. One credit is one page on a basic scrape, crawl or map; credits are shared across all endpoints.
What is public
Firecrawl publishes six tiers from Free to Enterprise, monthly and annual prices, and per endpoint credit costs.
Billing mechanics
A credit based monthly or annual subscription; the annual figure is the effective monthly rate, not a separate plan. Each tier grants a monthly credit pool shared across all endpoints, with most calls one credit per page and higher rates for structured formats, search, browser minutes and monitor checks. Annual Scale and Enterprise plans allow some rollover.
Cost watchouts
A page that returns an error status such as 403 or 404 still costs a credit. Structured formats (JSON, Question, Highlight) add 4 credits per page, Search costs 2 credits per 10 results, Interact 2 credits per browser minute, and Monitor 1 credit per page per check.
Variable cost rationale
Tiers are fixed monthly subscriptions with predictable credit pools, but structured formats, search, browser minutes and monitor checks consume credits faster than the one credit per page base, and unused credits are generally lost.
Additional watchouts
Unused credits generally do not roll over. Structured output adds four credits per page, so an extraction-heavy crawl costs several times the one credit base, and error pages still cost a credit.
Overage / add-ons
Credit based subscription tiers; exceed a tier by upgrading or buying more credits. Annual Scale and Enterprise plans allow some credit rollover; other credits do not carry over.
Sales call required
No, self serve available
Free / trial
Free tier: 1,000 credits per month, no card. Open source core self hostable.
Lowest paid plan
Hobby $16/mo billed yearly or $19/mo billed monthly (5,000 credits)
Commercial notes
Sold bottom up to developers and AI teams through a free tier, open source distribution under AGPL 3.0, an MCP server, skills and framework integrations. Self serve tiers run up to Scale; Enterprise adds custom credits, SSO and SCIM, IP and key restrictions, Threat Protection, SIEM audit logging, Zero Data Retention, pooled credits, spend limits and dedicated support.
Key ambiguities
Real monthly cost depends on which endpoints and formats a workload uses, since structured formats, search, browser minutes and monitor checks draw credits at different rates from the one credit per page base.
Cancellation / refund
Self serve monthly or yearly subscriptions via Stripe; standard subscription cancellation. Enterprise terms are contractual.
Support SLA / resale
Self serve support by tier; Enterprise adds dedicated support and SLAs.
Missing data
Enterprise credit volumes and discounts are custom, and Agent run pricing is not on the pricing summary read.
Related vendors
- AgentOps — Agent observability and debugging platform: open source SDKs trace…
- Agno — Python agent framework and AgentOS runtime (formerly Phidata) for…
- AIsa — Resource and payment gateway for AI agents: one key to 110+ models…
- AlphaBitCore — AI control plane for regulated financial firms: one gateway enforces…
- Anchor Browser — Cloud hosted browser infrastructure that lets AI agents operate real…
- Apify — Cloud platform and marketplace of more than 73,000 ready-to-run…
Alternatives to Firecrawl
The closest documented capability profiles to Firecrawl among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.
- Bright Data8.0 / 14Fuller documented coverage on Human Oversight & Guardrails
- Daytona8.0 / 14Adds documented Memory & State Persistence
- Confident AI8.5 / 14Adds documented Testing, Debugging & Optimization
- E2B8.5 / 14Adds documented Memory & State Persistence
- HoneyHive8.5 / 14Adds documented Testing, Debugging & Optimization
- Langfuse8.5 / 14Adds documented Testing, Debugging & Optimization
Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded