Repository navigation
firecrawl vs ai browser agents
Here is the answer before the article. If your task is "URLs go in, LLM-ready content comes out" - feeding a RAG pipeline, summarizing pages, building a corpus - Firecrawl or something shaped like it wins over a browser agent on cost, speed and simplicity, and it is not close. If your task needs things to happen - logged-in sessions, forms filled, multi-step flows where the next click depends on what the page said - a crawl API is the wrong tool category, and an agent driving a real browser is the right one. Everything below is the reasoning, the numbers, and the overlap where the choice gets genuinely interesting.
Disclosure first: this wiki belongs to invisible_playwright_mcp, an open-source AI browser agent, so we are a vendor of exactly one side of this comparison. Every Firecrawl fact below was read from Firecrawl's own repository, site or documentation on 2026-09-03, and the paragraph conceding their side is not smaller than the one favoring ours.
Firecrawl calls itself "the context API to search, scrape, and interact with
the web at scale", and the shape of the product is exactly that sentence: an
API. You send a URL and get back clean markdown or structured JSON built for
model consumption; a /crawl endpoint follows links across a site; a
/search endpoint returns results with full-page markdown included. It is
the most-adopted project of its kind by a huge margin - about 176,000 GitHub
stars when read on 2026-09-03.
The licensing is worth thirty seconds because it is layered. The core repository is AGPL-3.0; the SDKs and some UI components are MIT. And the hosted cloud is more than the open code: Firecrawl's own self-hosting guide states that Fire-engine, the fetching layer it credits with the cloud's "advanced anti-bot behavior" and some capabilities like screenshots, "is not included" in the self-hosted stack. Their words, their architecture, and a fair thing to know before you assume the AGPL repo equals the product at firecrawl.dev.
One more of their claims, quoted because it is the honest center of their pitch: clean markdown "with 93% fewer input tokens for your model" compared to raw page content. That number is Firecrawl's own measurement, but the direction is obviously right - stripped markdown is far smaller than raw HTML, and token cost is the tax every LLM pipeline pays.
An agent is a different machine entirely: a language model in a loop with a real browser, reading the page, deciding an action, performing it, and repeating until the goal is met. The explainer covers the loop; the property that matters here is that nothing is scripted in advance. The model decides at run time, which is what lets an agent do errands - sign in, navigate, fill, submit, retry - and also what makes it slower and more expensive per page than any API call will ever be.
Does your task read the web, or act on it?
Reading at volume is Firecrawl's home game. Acting is the agent's. The boundary is sharper than it first looks, because Firecrawl does have an interaction feature and it is worth describing precisely: its scrape can run a caller-provided sequence of actions before extraction - write, press, click, wait, screenshot. That is scripting, not deciding. You write the sequence in advance against a page structure you already know, which puts it in the same family as a Playwright script: excellent when the steps are known and stable, brittle when the page changes, and unable to handle a flow whose next step depends on judgment. The distinction between scripted and decided is the entire subject of AI browser agents vs traditional scraping; this page is the named-tool version of it, and the general cost math lives there, not here.
For content-to-LLM at any real volume, the crawl API side wins on every axis that pays the bills:
- Cost per page. An API call versus an agent loop that sends the page through a model at least once, usually several times. The agent's cost per page is orders of magnitude higher, and their "93% fewer input tokens" framing shows they know exactly which axis they are selling.
- Speed and parallelism. Requests fan out; an agent's loop turns are serial and model-latency bound.
- Operational simplicity. No browser fleet, no loop to supervise, no non-deterministic wandering to cap. URL in, markdown out.
- The ecosystem agrees. This is the category that feeds RAG pipelines and model context, and Firecrawl's adoption is the measure of the fit.
If that list is your task, use the API. An agent aimed at bulk extraction is a misallocation: slower, costlier, and non-deterministic where determinism was on offer, and this is an agent's own wiki saying so.
The agent's territory is everything an API request cannot be:
- State. A logged-in session that persists across steps, a cart, a wizard, a dashboard behind auth that expects a real, continuous visitor.
- Actions with consequences. Submitting, booking, applying, replying, changing settings: tasks where the deliverable is a changed website, not a document (the forms page is the honest account of how well that goes).
- Judgment mid-flow. "Find the cheapest of these that ships this week and order it" cannot be pre-scripted, because the branch points depend on content nobody has seen yet.
- Unknown structure. One instruction across twenty differently-built sites, where writing twenty action sequences costs more than the answer.
invisible_playwright_mcp sits on this side, and the differentiator we bring is the browser itself: a real Firefox patched at the C++ level rather than a headless Chromium in a datacenter, running locally, with identity derived from a seed so a failing run replays. The boundary stays where our wiki always puts it: that addresses what a page reads from the browser, and does nothing for your IP, pacing, or a site's limits - why agents get blocked is the full map, and it applies to hosted crawl APIs and agents alike.
The two categories compose better than they compete:
- Agent for the errand, API for the volume. Use an agent to reach the state that matters (find the section, work out the path), then hand the URL space to a crawl API for the thousand-page sweep.
- API for the corpus, agent for the exceptions. Route everything through the cheap path and send only the pages that need interaction or judgment to the agent, keeping the blended cost near the API's.
- Monitoring split. Watching pages for change is API-shaped work; reacting to a change with a multi-step task is agent-shaped (the monitoring page walks that line).
Is Firecrawl an AI agent? No, and it does not claim to be. It is a hosted API that turns web pages into LLM-ready markdown and JSON, with optional pre-scripted page actions. No model decides the next step inside Firecrawl's loop; your pipeline consumes its output.
Firecrawl vs Playwright? Different layers. Playwright is a library you program to drive a browser you run; Firecrawl is a hosted service that does the fetching and cleaning for you. A Playwright script and Firecrawl's action sequences are cousins; an agent is the third thing, a model deciding at run time.
Firecrawl vs Browserbase? Also different layers, easy to conflate because both are infrastructure: Firecrawl sells extracted content, Browserbase sells the browsers themselves for your code to drive. The infrastructure explainer covers that category.
Can Firecrawl log in and fill forms? It can execute a scripted sequence you provide (write, click, press, wait) before extracting. What it cannot do is decide the sequence: flows that branch on page content need an agent.
Is Firecrawl open source? The core repo is AGPL-3.0 and the SDKs are MIT, but the cloud's fetching layer (Fire-engine) is a separate service their self-host docs say is not included. Self-hosters get a real product, not the whole product.
Which is cheaper? For reading pages at volume, Firecrawl-shaped APIs, by orders of magnitude per page. For a task, the comparison stops making sense: an API cannot do a task, so the agent's token bill is the price of the category, not a premium over the API.
See also: AI browser agents vs traditional scraping for the category-generic version of this trade-off, cloud browser infrastructure for AI agents for the layer Firecrawl gets confused with, extracting data to CSV with an agent for the agent-side extraction workflow, and open-source AI browser agents for the agent field itself.
All retrieved 2026-09-03.
- firecrawl/firecrawl: stars, description, AGPL-3.0 core and MIT SDK licensing.
- firecrawl.dev: the markdown, crawl and search endpoints, and the "93% fewer input tokens" claim quoted above.
- Firecrawl self-hosting guide: the Fire-engine exclusion quoted above.
- Firecrawl scrape documentation: the pre-extraction action types.
- feder-cr/invisible_playwright_mcp, for the agent-side claims about our own tool.
From the invisible_playwright_mcp wiki. invisible_playwright_mcp is an agent, and this page's first paragraph hands bulk extraction to the other category anyway. If we only won half the comparisons, we would rather be trusted on that half.
- OpenAI Operator alternatives
- Open-source Operator-style agents
- Is OpenAI Operator still available?
- OpenAI Operator vs Claude computer use
- browser-use alternatives
- Choosing an AI browser agent
- Open-source AI browser agents
- Open-source computer-use agents
- What is an AI web agent?
- AI browser agents vs traditional scraping
- Cloud browser infrastructure for AI agents, explained
- Browserbase alternatives
- Firecrawl vs an AI browser agent
- Skyvern alternatives
- Stagehand vs browser-use
- Project Mariner is gone: what replaced it
- Manus alternatives
- Gemini computer use vs Claude computer use
- A stealth browser MCP, reviewed honestly by its own wiki
- AI browser vs AI browser agent: which one do you want?
- AI browser agent vs RPA: which one fits the job
- AI browser agent vs n8n, Zapier and Make
- Vercel agent-browser alternatives, compared honestly
- What is an agentic browser? Definition and the two kinds
- Open-source agentic browsers: the three layers, compared
- Choosing an MCP server for browser automation: four axes
- Stealth MCP servers compared: Camoufox, nodriver, Patchright
- Playwright MCP alternatives, and the three you don't need
- Autonomous browser agents: the four rungs of autonomy
- What is actually free in the AI browser agent stack
- browser-use on GitHub: what the repo actually gives you
- Playwright MCP vs Chrome DevTools MCP: different jobs
- How to choose among MCP servers: a map by category
- Which MCP servers are worth adding to Claude Code
- MCP on GitHub: finding servers and judging them fast
- MCP vs an API: the decision, and what the wrapper costs
- MCP alternatives: when the protocol is the wrong shape
- Why does my AI agent get blocked?
- The timing signal AI agents give off
- Agent retry loops trip rate limits, not fingerprints
- Claude computer use detected as a bot
- browser-use getting blocked: what you can and cannot change
- Playwright MCP session blocked: four causes, four fixes
- Playwright MCP and captchas: what actually gets you past
- Cloudflare and a browser MCP server: what is being read
- Can an AI agent solve a captcha? The honest answer
- Getting an AI agent to fill out forms
- The best model for an MCP browser agent, and what it really costs
- Browser problem or model problem?
- How to use a Playwright MCP server with Claude Code
- Extracting data to a CSV with an AI agent
- Monitoring a page for changes with an AI agent
- How to let Claude Desktop control a browser
- How to add a browser to Cursor as an MCP server
- Using an AI agent to hunt for apartments
- Getting website data into Google Sheets with an AI agent
- Using an AI agent to download invoices from portals
- AI agents for web research
- Using an AI agent to test your own website
- How to add a browser to Cline as an MCP server
- Posting to social media with an AI agent
- Posting to Facebook with an AI agent
- Posting to Instagram with an AI agent
- Posting to X with an AI agent
- Automating LinkedIn posts: read this first
- Appointment bots: what they are and what an agent can legitimately do
- Track prices across sites with an AI agent
- Build a lead list with an AI browser agent
- Run an AI browser agent on a schedule
- AI browser agent with a local LLM: what changes
- Should you log your AI agent into your accounts?
- How to write a task an AI browser agent can follow
- Move data between two web apps with an AI agent
- The MCP server
- How the tools are shaped, and why
- Playwright MCP vs the Playwright CLI: which fits when
- Playwright MCP: browser is already in use, and the fix
- Playwright MCP best practices: four decisions that matter
- Playwright MCP with a proxy, and the three leaks it leaves
- A browser MCP server in GitHub Copilot: setup and limits
- Using a browser MCP server for web scraping: the pattern
- Which LLM for browser automation: the four properties
- How to build a browser agent, and what to take instead
- Getting an AI agent to log into a website: three routes
- MCP tools, resources and prompts: who controls each
- How many MCP tools is too many? The context arithmetic
- How to build an MCP server: the decisions, not the scaffold
- Local or remote MCP server: what changes, and what does not
- Writing an MCP client in Python: the thirty-line version
- Self-hosted AI agent: what one actually costs to run
- How long an AI browser agent takes per step, measured
- Text, HTML, snapshot or screenshot: what the agent should read
- Giving an AI browser agent a stopping condition
- Keeping an AI browser agent out of destructive actions
- Why did the AI agent click the wrong thing
- When the page changes under the AI agent
- Running one AI agent task across a list of sites
- Seeing a page as it appears in another country
- Getting data out of a dashboard with no export button
- Two browsers in one session: main and support
- Finding the dead links on a site with an AI agent
- Filling a CRM record from a company's website
- One form submission per spreadsheet row, with an AI agent
- Dated screenshots of a page as evidence
- Checking order and delivery status with an AI agent
- Reading a PDF that opens inside the browser
- Summarising a long page or thread with an AI agent
- Collecting every image on a page with its caption
- Collecting event and course listings with an AI agent
- Cancelling a subscription with an AI agent
- What an AI agent can and cannot do inside an iframe
- Shadow DOM and an AI agent: you can click it, you cannot read it
- How to add a browser to Codex as an MCP server
- What a page snapshot costs, per control
- How to let Gemini CLI use a browser
- Native selects and the ones that only look like selects
- Clicking by selector or by coordinates
- How long the agent waits before it gives up
- What a second browser costs
- Uploading a file with an AI agent, and why this one cannot
- Watching the agent work, and when it is worth it
- When not to use an AI browser agent
- Agent or script: deciding once instead of every time
- Using the keyboard instead of the mouse
- Secrets in an agent task: where they end up
- What an agent run should log
- Deduplicating what an AI agent collects
- Normalising values across sites
- Validating an AI agent's output
- Reading a table with an AI agent
- Driving a site's own search and filters
- The task works headed and fails headless
- How to let the AdaL CLI use a browser