-
Notifications
You must be signed in to change notification settings - Fork 221
how to scrape infinite scroll playwright
Most infinite-scroll code looks the same: scroll down, sleep a couple of seconds, scroll again, repeat a fixed number of times. It works on the page you tested it on and breaks on the next one, because a fixed sleep is a guess about a network and a render pipeline you do not control.
This page is the loop that does not guess: how to wait for content to actually arrive, how to know when the page is really done rather than just slow, how to dedupe what you collect, and the part generic scroll tutorials skip entirely - a perfectly mechanical scroll loop is itself something a page can notice, and the block that shows up minutes into a session is usually that, not your fingerprint.
A reliable infinite-scroll scraper makes four decisions, and the fragile version of each is the one most tutorials ship. Wait on a condition, stop on a streak, key dedup on an identifier, and vary the motion:
| Decision | Fragile version | Reliable version |
|---|---|---|
| Know a batch loaded |
sleep(2) between scrolls |
wait_for_function on scrollHeight + wait_for_load_state("networkidle")
|
| Know the feed ended | stop on the first flat round | stop on a streak of consecutive flat rounds |
| Collect without repeats | index or scroll position | a stable per-item identifier |
| Stay unremarkable | constant wheel delta at a fixed interval | seeded, varied delta and dwell |
Each row is a section below.
A sleep(2) between scrolls encodes an assumption: that two seconds is always enough for
the next batch of items to load, render, and settle. It usually is not always true in
either direction. Slow connections and heavy pages make two seconds too little, so you
scroll past content that has not arrived yet and read a shorter page than exists. Fast
pages make it too much, and a thousand-item feed now takes ten times longer to collect
than it needs to.
The fix is the same one behind most Playwright reliability advice: wait for a condition, not for a duration. For infinite scroll the condition is content growth, which you can read directly from the DOM.
The reliable pattern is one loop: scroll, then wait for document.body.scrollHeight to
actually increase before scrolling again, with a per-round timeout standing in for the
sleep. When the height stops growing for several rounds in a row, the feed is done.
Start from a normal launch. The browser object below is a real Playwright Browser,
so everything from here on is standard Playwright, not a project-specific API:
from invisible_playwright import InvisiblePlaywright
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
page.goto("https://example.com/listings")
page.wait_for_load_state("networkidle")Now the loop. Each round scrolls, then waits for document.body.scrollHeight to have
actually increased, with a timeout instead of a sleep:
def scroll_until_stable(page, max_stable_rounds=3, round_timeout_ms=10_000):
stable_rounds = 0
last_height = page.evaluate("document.body.scrollHeight")
while stable_rounds < max_stable_rounds:
page.mouse.wheel(0, 2400)
try:
page.wait_for_function(
"prevHeight => document.body.scrollHeight > prevHeight",
arg=last_height,
timeout=round_timeout_ms,
)
except Exception:
# height did not grow within the timeout - this may be the end
stable_rounds += 1
continue
stable_rounds = 0
last_height = page.evaluate("document.body.scrollHeight")
page.wait_for_load_state("networkidle")Two conditions are doing the work here, not one. wait_for_function catches the DOM
actually growing, and wait_for_load_state("networkidle") catches the requests that fill
it in settling down before the next scroll. A page that appends a placeholder and fills it
in later can pass the first check and still be mid-load, which is why the second one runs
too.
A scroll loop knows the feed ended when max_stable_rounds consecutive scrolls fail to
grow the page: once that streak is long enough, "still loading" is no longer the likely
explanation.
One round of no growth is not proof the feed ended. A slow response, a lazy-loaded image block still resolving, or a scroll that landed just before a batch was pushed can all produce a single flat round on a page with plenty left. Two or three consecutive flat rounds is a much stronger signal, which is why the check is a streak and not a single miss.
If the page also gives you something more direct - a "no more results" element, an item count in the page, an API response that returns an empty array - prefer that over height alone. Height is the general-purpose fallback, not the best signal when a better one exists.
Scroll-and-collect naturally re-reads items you already have, because the DOM keeps everything loaded so far rather than replacing it. Track what you have already recorded by a stable identifier, not by position:
seen_ids = set()
items = []
def harvest(page):
cards = page.locator("[data-item-id]")
count = cards.count()
for i in range(count):
card = cards.nth(i)
item_id = card.get_attribute("data-item-id")
if item_id in seen_ids:
continue
seen_ids.add(item_id)
items.append({
"id": item_id,
"title": card.locator(".title").inner_text(),
})Call harvest(page) after each successful growth round rather than only once at the end.
Two reasons: it lets you inspect progress while the loop still runs, and it means a crash
or a block partway through still leaves you with everything collected up to that point
instead of nothing.
If the page has no stable identifier at all, the URL of a detail link or a normalized title-plus-position hash is a reasonable fallback, but prefer whatever attribute the site already uses to key the item - it is far less likely to collide.
Waiting for growth instead of a fixed sleep solves the reliability problem, but it does nothing about a second problem that only shows up once a session runs for a while: the loop is a perfect, identical motion repeated hundreds of times, and that pattern is itself something a page can watch for.
A constant wheel delta, a scroll that never pauses, never overshoots, never slows down to read anything, and fires at a rate no human scrolling with a mouse or trackpad produces - that is a behavioral signature, and it is a different layer entirely from the fingerprint. The checklist for one-site detection puts behavior at step five for a reason: blocks that arrive after an interaction, minutes into a session rather than at the first request, are usually this rather than anything about the browser's reported GPU or fonts. A perfectly consistent fingerprint does not help if the scrolling pattern on top of it is the part that stood out.
This is easy to miss because it produces no error. The loop above will run cleanly for a while and then simply start getting shorter pages, emptier responses, or a challenge instead of the next batch - which is the same shape of failure this project hit with a WebRTC gate that asserted absence instead of presence: a suppressed or truncated result looked fine to a check that was only watching for an exception.
The fix is the same one this project applies to pointer movement: vary the motion instead of repeating it, and derive the variation from a seed so a run that misbehaves can be replayed rather than re-guessed.
import random
sf = InvisiblePlaywright(seed=42)
with sf as browser:
page = browser.new_page()
page.goto("https://example.com/listings")
rng = random.Random(sf.seed)
def human_scroll_step(page):
delta = rng.randint(500, 1800)
page.mouse.wheel(0, delta)
page.wait_for_timeout(rng.randint(200, 900))That pause is not the thing telling you content arrived - wait_for_function still is,
exactly as before. It only changes how long the loop dwells between wheel events, so the
gaps look like a device scrolling rather than a timer firing at a fixed interval. Wire
human_scroll_step in as the scroll action inside scroll_until_stable and the growth
check around it does not change at all.
The same idea already ships for clicks and hovers: humanized mouse movement is on by default in this project, driven from the same per-session seed, so the arc to a button and the pacing of a scroll come from one consistent identity rather than two unrelated randomizers. Bezier curves get the pointer path right; what actually gets read on a scroll loop is closer to timing than shape, since a wheel event does not carry the same rich field set a pointer move does - but a constant interval is exactly as loud as a teleporting cursor, and for the same reason.
Infinite scroll sometimes ships with client-side routing bolted on: a "load more" click that also swaps the URL, or a feed that quietly replaces the whole document once you pass a certain point instead of appending to it. Either one is a navigation, and a navigation mid-evaluation produces the classic error:
Error: Execution context was destroyed, most likely because of a navigation
Most of the time this is your own code holding a handle across a scroll that triggered a
route change, and the fix is the ordinary one: re-query elements after the growth check
instead of carrying a Locator snapshot across it. Occasionally it is the site deciding,
mid-scroll, that this session should go somewhere else, and the two look identical in the
stack trace. The full breakdown of that error covers how
to tell them apart - the short version is to read page.url the moment the error fires:
the URL you expected means your code raced, a different one means the page moved you.
Given the previous section, there is a specific reason to check this on infinite scroll in particular: if the context keeps dying at the same scroll position, after the same number of rounds, every single run, that is not a race. Races are intermittent. A page that reacts at a consistent point in a repeated action is behaving like a threshold, not a bug.
An infinite-scroll loop that works on every page you throw it at waits for the DOM to actually grow and for the network to actually settle, stops on a streak of flat rounds rather than one, and keys its dedup on an identifier rather than position. None of that is specific to this project; it is just Playwright used the way its own waiting primitives are meant to be used.
The part worth adding on top is the one generic tutorials leave out: the motion of the loop is itself a signal, separate from anything about the browser's fingerprint, and a uniform scroll at a uniform interval is exactly as noticeable as a teleporting cursor. Vary it, derive the variation from a seed so a bad run is reproducible, and read a block that arrives mid-session as a hint about behavior rather than a fingerprint bug to chase.
How do I know when infinite scroll has finished loading? Wait for
document.body.scrollHeight to grow after each scroll, with a timeout instead of a sleep,
and stop after a streak of rounds that did not grow, not after a single one.
Why does my scroll loop miss items or duplicate them? Because it is scrolling on a fixed schedule instead of waiting for content, so it either reads before a batch arrives or re-reads a batch it already has. Key what you collect by a stable identifier and only harvest after a confirmed growth round.
Should I just use wait_for_timeout between scrolls? Only for pacing, never as the
signal that content arrived. Use it to vary how long the loop dwells; use
wait_for_function or wait_for_load_state("networkidle") to know whether the page
actually grew.
Why do I get blocked partway through a long scroll session and not at the start? A block that shows up after a while, rather than on the first request, usually means behavior rather than fingerprint. A scroll loop that never varies its pace or distance is one of the more obvious things there is to watch for.
What does "Execution context was destroyed" mean during a scroll? Usually that your code held a reference across a navigation the scroll itself triggered. Re-query after each round. If it happens at the exact same point every run, it may be the page reacting rather than a race.
Does humanizing the scroll matter if the site does not watch behavior? No. Check whether the block arrives at the first load or only after scrolling for a while. The first points at the fingerprint or the address; the second points here.
-
Playwright's own waiting primitives
(
wait_for_function,wait_for_load_state,expect_navigation) and its auto-waiting / actionability model, used here for content growth rather than a fixed sleep. - This project's release gates, including the WebRTC gate whose absence-only assertion produced a false pass, cited here as the same shape of mistake a scroll loop can make by treating a truncated or suppressed page as a clean end of feed.
See also: the checklist for being detected on one site for where behavior sits in the overall order, how to scrape without getting blocked for the layers above the scroll loop itself, how to scrape a "load more" button for the click-driven cousin of this same pattern, and how to test whether your browser is detected for why a clean-looking run is not the same thing as a passing one.
From the notes of invisible_playwright, a Firefox patched at the C++ level driven by stock Playwright. The scroll loop in this project's own test suites started as a fixed-sleep version and got rewritten after it silently under-collected on a slower connection.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2024
- selenium-stealth hasn't been updated since December 2021
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder