-
Notifications
You must be signed in to change notification settings - Fork 230
how to scrape infinite carousels playwright
To scrape an infinite carousel with Playwright, treat it as a ring with duplicated ends instead of a list: pause the autoplay before reading anything, mark the cloned wrap slides by the marker the widget puts on them rather than by matching their text, read a real slide index off each slide as you advance, wait for that index's own content before extracting it, and stop when an index you have already recorded comes round again rather than when a step adds nothing new.
A looping carousel earns its unbroken wrap by lying about how many slides it has. The widget copies the first slides and appends them after the last, copies the last slides and puts them in front of the first, then snaps the track back with transitions off while a copy is on screen. The eye sees one ring. The DOM holds extra nodes carrying duplicate content.
That is the opposite failure from the vertical case. A virtual scrolling table keeps too few nodes and recycles them, so records vanish before you read them. A carousel keeps too many, so a plain collector records the first slides twice and the loop never runs out of slides to count.
Count what is in the track, then compare it against what the widget says it holds. Those two numbers disagree on every looping carousel, and the gap is how many copies sit on each side.
from invisible_playwright import InvisiblePlaywright
SURVEY = """
(root, sel) => {
const slides = [...root.querySelectorAll(sel)];
const cloneish = (el) => {
const cls = (el.className || '').toString().toLowerCase();
return cls.includes('clone') || cls.includes('duplicate')
|| el.hasAttribute('data-cloned');
};
const readIndex = (el) => {
for (const name of ['data-slide-index', 'data-index', 'data-position']) {
const raw = el.getAttribute(name);
if (raw !== null && raw !== '') return parseInt(raw, 10);
}
const pos = el.getAttribute('aria-posinset');
return pos === null ? null : parseInt(pos, 10) - 1;
};
return {
nodes: slides.length,
marked_clone: slides.filter(cloneish).length,
aria_hidden: slides.filter(s => s.getAttribute('aria-hidden') === 'true').length,
indexes: slides.map(readIndex),
set_size: parseInt(slides[0]?.getAttribute('aria-setsize') || '-1', 10),
dots: root.querySelectorAll('[role="tab"], [class*="pagination"] > *').length,
};
}
"""
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
page.goto("https://example.com/collection", wait_until="domcontentloaded")
root = page.locator("div.carousel")
root.wait_for(state="visible")
print(root.evaluate(SURVEY, ".slide"))
# {'nodes': 16, 'marked_clone': 4, 'aria_hidden': 4, 'set_size': 12, 'dots': 12,
# 'indexes': [10, 11, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 0, 1]}Sixteen nodes for twelve slides, and look where the index list starts. The first node in
the DOM is slide 10, because two copies of the tail sit in front of it. Every habit that
reaches for slides.first or nth(0) picks a copy of the second to last slide and calls
it the beginning.
Extend the attribute list before trusting it. Most widgets stamp their own prefixed name,
in the shape data-<widget>-slide-index, and the generic ones above are only what survives
across implementations. When every entry comes back null, the next section is the work.
Three markers appear in practice and only one is unambiguous. A class containing clone
or duplicate says what it means, and so does an index outside the real range: negative in
front, at or above the set size behind. Some widgets stamp the original index on the copy,
which is better still, because the copy names what it duplicates.
aria-hidden="true" is the one people reach for first and the one that misleads. Plenty
of carousels set it on genuine off-screen slides too, so it over-marks. Treat it as a
hint, never as the classifier.
Validate with arithmetic before the loop starts: the unmarked slides must equal
aria-setsize, the pagination dot count, or the total the feed declared. Under-marking
leaves duplicates in the output, which is loud. Over-marking drops real slides, and the
result looks clean.
Text is the tempting shortcut and it fails in both directions at once. Promotional slides repeat a caption, so a text key merges records that were never duplicates, while a copy whose lazy image has not promoted yet reads differently from its original and slips through as new.
An auto-rotating carousel moves between your read and your click. You read index 4, the timer fires, your click lands, and you are on 6. Slide 5 is never recorded and nothing failed.
Racing that timer does not help, because the gap is milliseconds wide and the widget owns
the clock. Ask the component to stop instead. Carousels that respect
prefers-reduced-motion suspend their own rotation when it is set, and Playwright sets
that media feature at page creation, before the widget initialises. Setting it afterwards
through emulate_media helps only when the widget subscribed to the change event, and
many read it once at startup.
import re
def pause_autoplay(page, root, slide_sel, quiet_ms=6000):
# The ARIA pattern requires a pause control on an auto-rotating carousel.
control = root.get_by_role("button", name=re.compile(r"paus|stop", re.I))
if control.count() and control.is_visible():
control.click()
# Verify from inside the page that it stopped, instead of assuming it did.
before = root.evaluate(ACTIVE_INDEX, slide_sel)
page.wait_for_timeout(quiet_ms)
if root.evaluate(ACTIVE_INDEX, slide_sel) != before:
raise RuntimeError("carousel is still rotating after the pause attempt")
page = browser.new_page(reduced_motion="reduce") # set before the widget loadsKeep the verification: a pause you did not confirm is a guess, and its cost is a gap in the middle of a tidy dataset. Freezing the clock looks clever and breaks the run, because it stops the CSS transition along with the timer, so the track never settles and every read lands mid-flight.
Two stopping rules get written here and both are wrong. "Stop when a step adds no new slides" never fires, because a ring produces slides forever. "Stop when the index returns to 0" fires only sometimes, because a carousel that moves a group at a time steps by two or three and can jump straight over index 0 without landing on it.
The rule that survives both is a set. Record every real index you visit and stop the first time the current index is already in it. Whatever the step size and whatever the starting slide, the ring closes on the first revisit.
ACTIVE_INDEX = """
(root, sel) => {
const slides = [...root.querySelectorAll(sel)];
const real = slides.filter(s => !s.matches('[class*="clone"], [class*="duplicate"]'));
const box = root.getBoundingClientRect();
const centre = box.left + box.width / 2;
let best = null, bestGap = Infinity;
for (const s of slides) {
const r = s.getBoundingClientRect();
if (r.width === 0) continue;
const gap = Math.abs((r.left + r.width / 2) - centre);
if (gap < bestGap) { bestGap = gap; best = s; }
}
if (!best) return null;
for (const name of ['data-slide-index', 'data-index', 'data-position']) {
const raw = best.getAttribute(name);
if (raw !== null && raw !== '') return parseInt(raw, 10);
}
const pos = best.getAttribute('aria-posinset');
if (pos !== null) return parseInt(pos, 10) - 1;
return real.indexOf(best); // -1 means: on a clone with no index to read
}
"""Picking the active slide geometrically, as the one whose horizontal centre sits nearest
the track's centre, outlasts every active-class convention. -1 is an answer and not an
error: the slide in front of you is a copy carrying no index, so advance without recording
rather than writing an entry you cannot key. Do not assume the starting index is zero
either, since carousels restore a position from session storage or open on whichever slide
a URL anchor asked for.
Slides hydrate on approach. A slide read in the frame it arrives is half built: the image is still a placeholder, the caption is empty, the price has not been written. Native lazy loading makes it worse, because a slide translated off to the side never intersects the viewport, so the browser leaves its images alone until the track brings it in.
The wait has to name the index, not a position. nth(2) after a transform is a different
slide than it was, and the copy of that slide sits elsewhere in the same track looking
identical.
SLIDE_READY = """
(a) => {
const root = document.querySelector(a.root);
if (!root) return false;
const slide = [...root.querySelectorAll(a.sel)].find(
s => !s.matches('[class*="clone"], [class*="duplicate"]')
&& (s.getAttribute('data-slide-index') === String(a.index)
|| s.getAttribute('data-index') === String(a.index))
);
if (!slide) return false;
const images = [...slide.querySelectorAll('img')];
// complete alone is true for a finished placeholder; the width is the test.
return images.every(i => i.complete && i.naturalWidth > 1)
&& slide.innerText.trim().length > 0;
}
"""
READ_SLIDE = """
(root, a) => {
const slide = [...root.querySelectorAll(a.sel)].find(
s => !s.matches('[class*="clone"], [class*="duplicate"]')
&& (s.getAttribute('data-slide-index') === String(a.index)
|| s.getAttribute('data-index') === String(a.index))
);
if (!slide) return null;
const img = slide.querySelector('img');
return {
index: a.index,
text: (slide.innerText || slide.textContent || '').trim(),
href: slide.querySelector('a')?.href || null,
image: img ? (img.currentSrc || img.src) : null,
};
}
"""img.complete on its own is the trap: it turns true the moment the 1x1 placeholder
finishes loading, which is immediately. naturalWidth > 1 separates a real asset from a
spacer, and currentSrc gives the candidate the browser picked out of the srcset rather
than the fallback in src. The same problem without the ring around it is
scraping lazy-loaded images.
One caveat on the text test. Carousels that hide inactive slides with visibility: hidden
return an empty innerText for content that is fully present, because innerText reports
rendered text. Use textContent on those, and prefer innerText everywhere else, since
textContent also returns the strings inside hidden helper markup.
Arrows are often invisible until the pointer enters the carousel, and how they are hidden
decides whether that matters. An arrow at opacity: 0 still has a bounding box and no
visibility: hidden, so Playwright treats it as visible and clicks it without complaint.
An arrow at visibility: hidden or display: none fails that check, and the click waits
out its timeout on an element sitting right there in the DOM.
def advance(page, root, next_sel):
control = root.locator(next_sel)
if control.count():
if not control.is_visible():
root.hover() # reveal, then click
control.wait_for(state="visible", timeout=3000)
control.click()
return
root.press("ArrowRight") # keyboard-driven widget
box = root.bounding_box()
y = box["y"] + box["height"] / 2
page.mouse.move(box["x"] + box["width"] * 0.8, y)
page.mouse.down()
page.mouse.move(box["x"] + box["width"] * 0.2, y, steps=12)
page.mouse.up() # steps matter: 0 reads as a clickSome widgets render no arrow at all until hover, so count() is zero on the first look
and one a moment later. Hover the container first and wait for the control to attach
rather than concluding there is none.
The drag is the last resort and it has a hard edge. Those moves produce pointer and mouse
events, and a carousel bound only to touchstart and touchmove ignores every one of
them. The steps argument is not cosmetic: a single jump from press to release carries no
intermediate movement, and the widget scores that as a click on whatever sits underneath.
The pieces compose into one loop. The extra helper is a settle check, because a read taken while the track is still transitioning finds the outgoing slide nearest the centre and returns its index.
import random
from playwright.sync_api import TimeoutError as PlaywrightTimeout
from invisible_playwright import InvisiblePlaywright
def settled(root, slide_sel, previous, tries=25):
"""Return once the index has changed and then held still for two reads."""
last, stable = None, 0
for _ in range(tries):
now = root.evaluate(ACTIVE_INDEX, slide_sel)
if now == last and now != previous:
stable += 1
if stable >= 2:
return now
else:
stable = 0
last = now
root.page.wait_for_timeout(120)
return last
def scrape_carousel(url, root_sel, slide_sel, next_sel, seed=42, max_steps=200):
rng = random.Random(seed)
with InvisiblePlaywright(seed=seed) as browser:
page = browser.new_page(reduced_motion="reduce")
page.goto(url, wait_until="domcontentloaded")
root = page.locator(root_sel)
root.wait_for(state="visible")
survey = root.evaluate(SURVEY, slide_sel)
pause_autoplay(page, root, slide_sel)
expected = survey["set_size"] if survey["set_size"] > 0 else survey["dots"]
args = {"root": root_sel, "sel": slide_sel}
seen, collected = set(), []
for _ in range(max_steps):
index = root.evaluate(ACTIVE_INDEX, slide_sel)
if index is not None and index >= 0:
if index in seen:
break # the ring closed
seen.add(index)
try:
page.wait_for_function(
SLIDE_READY, arg={**args, "index": index}, timeout=10000
)
except PlaywrightTimeout:
pass # record it, report the gap below
row = root.evaluate(READ_SLIDE, {"sel": slide_sel, "index": index})
if row:
collected.append(row)
advance(page, root, next_sel)
page.wait_for_timeout(rng.randint(280, 900))
if settled(root, slide_sel, index) == index:
break # finite carousel, at its end
return collected, expected, len(seen)Comparing len(seen) against expected at the end is the honest report. A ring that
closes after nine of twelve slides means the step skipped three, and the fix is a smaller
step or a click on each pagination dot in turn. The pause before each advance comes from
the same seed as the browser identity, so the run replays exactly while no two gaps match,
which is the cadence argument that applies to
any repeated action in a long session.
Two different widgets get called an infinite carousel and only one is a ring. The ring holds a fixed set of slides and wraps through clones, and everything above applies. The other is an endless rail that fetches more slides as it reaches its tail, which is horizontal infinite scroll with arrows on it. No first index ever comes back, so the stop has to come from the feed's own paging, the way a vertical infinite scroll takes it.
When either kind is fed by an endpoint, read the endpoint. A recommendation rail almost always fetches a JSON list and renders slides from it, and that response carries the real count, stable ids, full-size image URLs and the fields the caption truncates, with no clones and no transitions to wait out. The hooks are in capturing XHR and API responses.
Three other cases end the DOM route. A carousel inside a closed shadow root is not
reachable by a locator at all, though open roots are pierced for you and are covered in
scraping shadow DOM content. A track painted
into a <canvas> has no per-slide markup at any position. And when the slides are
thumbnails whose full assets live behind a lightbox, the ring gets you the list and
scraping image galleries gets the pictures.
A looping carousel is a ring with duplicated ends, and every mistake here comes from reading it as a list. The clones are ordinary nodes holding real content, so counting overcounts and the first node in the track is not the first slide. Find the marker the widget already puts on those copies, check it against the declared set size, and do not fall back to matching text, which merges different slides and misses identical ones in the same pass. Stop the autoplay through the component's own supported path and confirm it stopped. Then track the real index, wait for that index's content, and stop on the first repeat. The advancing is easy. Knowing which of the sixteen slides in front of you are the twelve that exist is the whole job.
Why does my carousel scraper record the first slides twice? Because a looping carousel appends copies of the opening slides after the last one so the wrap looks continuous. Those copies are real nodes with the same content. Detect them by a clone class or an index outside the real range, and skip them.
How do I know when the carousel has come full circle? Keep a set of the real slide indexes you have visited and stop the first time the current index is already in it. Returning to index 0 is not the test, because a carousel that moves a group at a time can step straight over it.
Can I dedupe the slides by their text? No, and it fails twice. Different promotional slides often share a caption, so a text key merges records that were never duplicates, and a clone whose lazy image has not loaded reads differently from its original, so the same key misses the duplicate it was meant to catch.
How do I stop the carousel advancing on its own? Create the page with reduced motion so the widget suspends its own rotation, click its pause control if it has one, then verify by reading the index twice with a wait between. Do not freeze the clock: that stops the transition too, so the track never settles.
The next arrow is in the DOM but the click times out. Why? It is hidden with
visibility: hidden or display: none until hover, and both fail Playwright's visibility
check. Hover the container first, then wait for the control. An arrow hidden with
opacity: 0 is a different case and clicks fine.
Why is the slide I just extracted half empty? Carousels hydrate slides on approach, so
a read in the arrival frame gets the placeholder image and an empty caption. Wait on that
specific index and require naturalWidth > 1, since complete turns true as soon as the
1x1 placeholder loads.
- Playwright's
emulate_mediaand thereduced_motionoption on context and page creation, which is what asks the component to suspend its own rotation, retrieved 2026-08-28. - Playwright's actionability rules,
which define a visible element as one with a non-empty bounding box and without
visibility: hidden, so an element atopacity: 0is still actionable, retrieved 2026-08-28. - Playwright's
wait_for_function,hoverandmouse.movewith itsstepsargument, retrieved 2026-08-28. - The WAI-ARIA carousel pattern this article reads rather than infers:
aria-setsizeandaria-posinseton the slides,aria-hiddenon the ones not presented, and the rotation control an auto-rotating carousel is expected to expose.
See also: scraping virtual scrolling tables for the vertical sibling that recycles nodes instead of cloning them, capturing XHR and API responses for the feed that carries the slide list without the ring, scraping lazy-loaded images for the placeholder problem on its own, and scraping infinite scroll for the endless rail that fetches rather than wraps.
Written while maintaining invisible_playwright, a Firefox patched at the C++ level driven by stock Playwright. The clone read is the one that cost a run here: the first three slides landed in the output twice, a text dedupe went in to clean that up, and it quietly merged two different slides sharing a caption, so the file came back shorter and looked correct.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
- How to scrape course catalogs with Playwright
- How to scrape store locator pages with Playwright
- How to scrape stock levels with Playwright
- How to scrape accordion and tab content with Playwright
- How to scrape size charts with Playwright
- How to scrape delivery slots with Playwright
- How to scrape appointment availability with Playwright
- How to scrape auction listings with Playwright
- How to scrape public transport timetables with Playwright
- How to scrape GraphQL endpoints with Playwright
- How to scrape virtual scrolling tables with Playwright
- How to scrape shipping rates with Playwright
- How to scrape cursor-based pagination with Playwright
- How to scrape multi-select facet filters with Playwright
- How to scrape currency exchange rates with Playwright
- How to scrape WebSocket streams with Playwright
- How to scrape book metadata with Playwright
- How to scrape professional directories with Playwright
- How to scrape range slider filters with Playwright
- How to scrape currency and locale switchers with Playwright
- How to scrape software changelogs and release notes with Playwright
- How to scrape breadcrumb hierarchies with Playwright
- How to scrape microdata and RDFa markup with Playwright
- How to scrape server-sent events with Playwright
- How to scrape open data portals with Playwright
- How to scrape infinite carousels with Playwright
- How to scrape printer-friendly pages with Playwright
- How to handle A/B test variants when scraping with Playwright
- How to scrape recipe data with Playwright
- How to scrape vehicle recall notices with Playwright
- How to scrape public tender notices with Playwright
- How to scrape nutrition labels with Playwright
- How to scrape podcast episode listings with Playwright
- How to scrape weather station data with Playwright
- How to scrape newsletter archives with Playwright
- How to scrape wine and spirits catalogs with Playwright
- How to scrape insurance quotes with Playwright
- How to scrape fitness class schedules with Playwright
- How to scrape flight seat maps with Playwright
- How to scrape concert and tour dates with Playwright
- How to scrape museum and gallery exhibition dates with Playwright
- How to scrape warranty terms with Playwright
- How to scrape sortable data tables with Playwright
- How to scrape salary and pay scale data with Playwright
- How to scrape live sports scores with Playwright
- How to scrape video game prices with Playwright
- How to scrape domain WHOIS records with Playwright
- How to scrape podcast transcripts with Playwright
- How to scrape patent listings with Playwright
- How to scrape clinical trial listings with Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2023
- selenium-stealth hasn't been updated since November 2020
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder