-
Notifications
You must be signed in to change notification settings - Fork 221
how to test bot detection
Test bot detection by comparing your automated browser against a stock browser on the same machine, field by field, run at least ten times through the same proxy production uses - and treat any suppressed or empty signal as a failure, not a pass. A single suite's verdict cannot tell a working feature from a broken one.
Most people test this by opening a suite, reading the verdict, and stopping. That method has a specific failure mode: it cannot tell "working correctly" from "not working at all", and it will happily report success on a browser that is broken in the exact way you were trying to avoid.
This page is what each public suite proves, the false pass that method produces, the comparison that replaces the verdict, where and how many times to run it, and what none of these tools cover.
Five public tools, each answering a different question. Using them interchangeably is most of why people get confused by contradictory results.
| Suite | Question it actually asks | What a pass proves |
|---|---|---|
| sannysoft | Are you an unmodified, old headless browser? | Table stakes only - nothing about a modern one |
| CreepJS | Are you lying about your environment? | Nothing here contradicts anything else here |
| BotD | Which engine are you really running? | A verdict, not a fingerprint |
| FingerprintJS | Can you be recognised again? | Linkability, not "looks human" |
| BrowserLeaks | What does one specific surface report? | A single value, not a score |
bot.sannysoft.com is a smoke test. Its main table describes headless Chrome as it behaved around 2018, and every serious tool fixes those on day one. Passing it proves you are not running an unmodified headless browser. The interesting part is not the table: three canvas tests are also run inside an iframe and compared, which is a consistency check rather than a fingerprint check.
CreepJS asks whether you are lying, not what you report. It takes a clean copy of the built-ins from a fresh iframe, inspects stack traces, walks descriptors and prototypes, and records a blocked probe as a lie by name. A high score means nothing here contradicts anything else here.
BotD returns a verdict rather than a fingerprint, and most of its twenty detectors are really asking which engine you are, by testing behaviours that differ between engines.
FingerprintJS gives you a visitor ID, which is a hash of roughly forty-one components. It answers "can I be recognised again", which is a different question from "do I look automated". The commercial version adds signals the open-source library does not have.
BrowserLeaks is per-surface rather than a verdict: WebRTC, canvas, WebGL, fonts. It is the one to reach for when you want to read a specific value rather than a score.
None of these is a superset of the others, and a green result on one says nothing about the rest.
A verdict-based test cannot distinguish a working feature from an absent one.
We learned this the expensive way. Our WebRTC gate asserted the sensible things: the host candidate does not expose the LAN address, no IPv6 candidate appears. Both passed, run after run. Behind a proxy on real pages, WebRTC was returning nothing at all, and the gate passed because a dead feature leaks nothing.
Every negative assertion is trivially satisfied by a feature that does not run.
So the rule, which applies to every surface and not just that one:
Assert the presence of the right signal, not the absence of a wrong one. A page that comes back empty, blocked or still loading is a failure, not a pass.
Concretely: the WebRTC section must complete and show a host candidate and a server reflexive one. The canvas must produce a hash, twice, matching. The font list must be non-empty and belong to the platform you claim. A suppressed signal is itself a signal, and CreepJS records blocking by name.
The same trap shows up one level down, inside a test itself rather than in what it tests. A cleanup-identification check once passed for exactly this reason - the input it claimed to reject never actually reached the code path being checked, so removing the guard entirely still left it green.
The single highest-value change to how you test bot detection is direct comparison, not a verdict.
Open the same page in your automated browser and in a stock browser on the same machine, and diff the two reports field by field. Not the scores, the fields.
What that catches which a verdict does not:
- Values that are individually plausible and disagree with each other.
- Values you fixed that now differ from a real browser in the other direction, which is what happens when a spoof overcorrects.
- Everything about the machine, which no verdict separates from everything about the automation.
The stock browser is the reference. Anything that differs between the two, other than the address, is a candidate. Anything that matches is not your problem, whatever the score says.
When something differs, ask which of the two it belongs to:
- Automation tells are things like
navigator.webdriver, leftover globals, an untrusted event. Which are mostly solved and mostly not your problem. - Machine tells are the GPU, the fonts, the audio device, the screen, the codecs. None of which any stealth layer touches.
They need completely different fixes, and most container failures are the second kind while most people debug the first.
Three ways a test can be true and useless.
Testing on your laptop. Your laptop has a GPU, fonts, an audio device and a real screen. The server has none of those. A test that passes at home tells you about home.
Testing on localhost. Localhost bypasses the proxy, so you have measured the path you do not use. Run the page through the same proxy the job uses, and confirm the exit address inside the browser rather than assuming it.
Testing the wrong browser version. If your tool pins one version and you tested another, you tested something you do not ship.
The general form: feed the test the same inputs production gets. A test that does not is not evidence.
This domain is not deterministic, and single runs mislead in both directions.
- Repeat at least ten times. A verdict that appears once in ten is a verdict, and a single green run is not a pass.
- Read the same value twice in one session. Canvas, audio and WebGL hashes must match. If they do not, something is randomising per call, which is the cheapest tampering check there is.
- Relaunch the same identity and compare. Values that should be stable across sessions must be.
- Space the runs out. Hammering one scoring endpoint from one address creates the velocity signal you are trying to measure. We flagged our own product for this once, and the flag belonged to the test harness.
A text log tells you what your code extracted. A screenshot tells you what the page actually rendered, including the parts your extractor did not know to look at. Several findings in these notes came from opening a PNG after the log said everything was fine.
Worth stating so a clean sweep is not mistaken for a clean session.
- The TLS handshake, which is decided before any page loads. And which no in-page test can see.
- Behaviour, including pointer motion, typing rhythm and, for agents, the pause shaped like model latency.
- IP reputation, which is not a browser property at all.
- What a commercial system does with all of it together, which is a model rather than a checklist.
A perfect score on every public suite means your browser is internally consistent and does not announce automation. It does not mean the session passes.
Testing well is mostly a method rather than a tool. Assert what should be present rather than what should be absent, compare against a stock browser instead of reading a verdict, run it where you deploy and through the proxy you deploy with, repeat it, and open the screenshot.
Do that and the public suites become useful instruments. Read them as scores and they will eventually tell you that a broken browser is fine.
Which bot detection test should I use? Several, because they answer different questions. sannysoft as a smoke test, CreepJS for tampering, BotD for a verdict, FingerprintJS for linkability, BrowserLeaks to read a specific surface.
I pass every test and still get blocked. Why? Because the suites do not see the TLS handshake, your behaviour, or your address, and because a consistent browser on a datacenter IP is still on a datacenter IP.
Is passing sannysoft enough? It proves you are not running an unmodified headless browser from around 2018.
Why do I get different results on different runs? Because the domain is non-deterministic. That is why ten runs, not one.
Should I test on my laptop? Only to find bugs. To find out what production looks like, test on the machine that runs production.
What does a blank or stuck section mean? A failure. An empty result is not a clean result, and several checks record suppression explicitly.
See also: the checklist for being detected on one site, which is the order to work in once a test tells you something, and the per-suite pages linked throughout.
- The five public suites named above, each with its own page in this set, read from their own source rather than from their rendered output.
- This project's release gates, including the WebRTC gate whose negative-only assertions produced the false pass described above, and the velocity flag that turned out to be the harness.
From the notes of invisible_playwright, a Firefox patched at the C++ level. Every rule on this page exists because breaking it cost us something first.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2024
- selenium-stealth hasn't been updated since December 2021
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder