-
Notifications
You must be signed in to change notification settings - Fork 230
how to scrape tour dates playwright
To scrape concert and tour dates with Playwright, key every row on the event page's URL instead of its date, stamp each status read with the time you read it, treat a missing ticket link before the announced on-sale moment as an unopened sale rather than a broken page, and read the cancelled badge as its own state instead of inferring cancellation from a show dropping out of a list. Get those four things right and the rest of the extraction, the artist name, the venue, the price, is the easy part.
A tour date looks like four fields on the page: artist, venue, date, status. Treat it as four static fields and the scraper works once and quietly breaks on the second run. Status turns over faster than anything else in the row and nothing on the page tells you when it was last true unless you write that down yourself. The date field itself is not stable either, because a postponed show keeps its page and swaps the date underneath it. This page walks through the parts of a tour-date scraper that only show up on a re-run: the timestamp, the re-key, the on-sale countdown, the cross-partner match, the presale window, and the cancelled badge that never leaves the list.
| What breaks a tour-date scraper | The fix |
|---|---|
| A sellout read last week looks identical to one read an hour ago | Stamp every status read with a checked_at timestamp |
| A postponed show swaps its date but keeps its page | Key rows on the event URL, not on the date |
| No ticket link exists before the on-sale moment | Read the on-sale timestamp; do not treat an empty link as an error |
| Two ticketing partners list the same show with no shared ID | Match on artist, venue and date together, not on an identifier |
| A presale code changes what the page shows | Record which sale window you read, general or presale |
| A cancelled show can stay listed with a badge | Read the badge state; do not infer cancellation from an absence |
A status field without a timestamp is a claim you cannot check. "Sold out" read once tells you the state at that instant and nothing about whether it is still true an hour, a day, or a week later. Store the moment of the read alongside the value, not just the value, or every downstream comparison, "is this still current," "did this change since yesterday," has nothing to compare against.
from datetime import datetime, timezone
from invisible_playwright import InvisiblePlaywright
def read_row(page, url):
page.goto(url, wait_until="domcontentloaded")
return {
"url": url,
"artist": page.inner_text("[data-testid='artist-name']").strip(),
"venue": page.inner_text("[data-testid='venue-name']").strip(),
"date_raw": page.get_attribute("[data-testid='show-date']", "data-date"),
"status": page.get_attribute("[data-testid='sale-status']", "data-state"),
"checked_at": datetime.now(timezone.utc).isoformat(),
}
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
row = read_row(page, "https://example.com/tour/summer-2026/city-a")
print(row)Reading the machine-readable data-state attribute rather than the visible label
matters here for the same reason it matters on a single-event listing: the label
is copy, and copy gets restyled and localized. The state attribute is the site's
own canonical value, and it is what you diff against on the next visit. Once
checked_at is on every row, a downstream job can answer "how stale is this" with
a subtraction instead of a guess.
Postponed and rescheduled shows are the case that quietly deletes rows from a
series if you key on the original date. Many ticketing systems keep the same
event page for a postponed show and simply overwrite the date field on it. If your
storage keys a row as (artist, venue, original_date), the next re-check reads a
different date on the same page, the old key no longer matches anything current,
and the row you already had just falls out of your dataset while the show itself
is still very much alive under a new date.
The event page's URL is almost always the stable part. Key storage on the URL,
and treat a changed date_raw on a re-visit as a signal to write, not a reason to
open a new row.
def reconcile(previous_rows_by_url, url, fresh_row):
"""previous_rows_by_url: dict keyed on event URL, from your own storage."""
prior = previous_rows_by_url.get(url)
if prior is None:
return fresh_row | {"event": "new"}
if prior["date_raw"] != fresh_row["date_raw"]:
# same page, different date: a postponement, not a new show
return fresh_row | {
"event": "rescheduled",
"previous_date": prior["date_raw"],
}
if prior["status"] != fresh_row["status"]:
return fresh_row | {"event": "status_changed", "previous_status": prior["status"]}
return fresh_row | {"event": "unchanged"}A series keyed on the date instead of the URL never sees the rescheduled branch
at all. It sees one row vanish and, if the postponed show later reappears in a
listing sorted by date, a second row that looks brand new. Both of those are the
same show, and the URL is what proves it.
On-sale time is frequently a future timestamp with its own countdown, entirely separate from the show date. A show can be announced weeks before tickets go on sale, and the page in that window has no ticket link at all: no price, no buy button, sometimes nothing but the countdown widget itself. Read that page with a scraper that expects a purchase link and it looks exactly like an error, a broken selector, a dead endpoint, when the page is doing precisely what it was built to do.
The fix is to read the on-sale timestamp explicitly and branch on it, rather than treating a missing link as a failure to retry.
from datetime import datetime, timezone
def read_availability(page, url):
page.goto(url, wait_until="domcontentloaded")
onsale_iso = page.get_attribute("[data-testid='onsale-countdown']", "data-onsale-utc")
buy_link = page.query_selector("[data-testid='buy-tickets']")
if buy_link is None and onsale_iso:
onsale_at = datetime.fromisoformat(onsale_iso)
if onsale_at > datetime.now(timezone.utc):
return {"state": "not_yet_on_sale", "onsale_at": onsale_iso}
# onsale_at is in the past but no link rendered: this one is worth a retry
return {"state": "unclear", "onsale_at": onsale_iso}
if buy_link is None:
return {"state": "no_link_no_countdown"}
return {"state": "on_sale"}The distinction that matters is between "nothing is here yet, on purpose" and "something should be here and is not." Only the second case is worth retrying or alerting on; the first is just the calendar not having arrived.
Multiple partners commonly list the same show at different prices with different fee structures, and there is no identifier shared between them: each partner's event ID is internal to that partner. Matching across sources has to run on the fields a human would use, artist, venue, and date, normalized enough that spelling and formatting differences do not split one show into two rows.
import re
def match_key(artist, venue, date_iso):
def norm(s):
s = s.casefold().strip()
return re.sub(r"[^a-z0-9]+", " ", s).strip()
# date only, no time: partners round or omit showtime differently
date_only = date_iso[:10]
return (norm(artist), norm(venue), date_only)
partner_a = {"artist": "The Long Winters", "venue": "The Grand Hall", "date_raw": "2026-09-14T19:00:00"}
partner_b = {"artist": "the long winters", "venue": "Grand Hall, Downtown", "date_raw": "2026-09-14T00:00:00"}
print(match_key(**{k: v for k, v in partner_a.items() if k in ("artist", "venue")}, date_iso=partner_a["date_raw"]))Venue names are the weakest field in that key, because one partner adds a neighborhood or a legal suffix and another does not; treat an exact match on artist and date plus a fuzzy match on venue as strong evidence, and hold anything softer than that for a person to confirm rather than silently merging two different shows into one price comparison.
A presale code opens an earlier window than the general on-sale, and a page can show radically different availability depending on whether that code was entered: a presale session might show tickets available while the same event, read without a code, shows nothing yet on sale. Neither read is wrong. They are two different, legitimate states of the same page, and a dataset that mixes them without saying which one it captured is not comparable to itself from one row to the next.
The concrete rule is to be explicit in the row about which state was read. If the
scraper never submits a presale code, every row it produces is the general
on-sale view, and it is worth recording that as a field, row["sale_window"] = "general", rather than leaving it implicit, because a downstream consumer has no
other way to know the session never saw the presale window at all.
If a project later adds presale tracking, it needs its own explicit sale_window
value and its own re-check cadence, because a presale window opens and closes on
a schedule that has nothing to do with the general sale's own timeline.
Cancelled shows sometimes stay on the listing page with a cancelled badge rather than being removed. A scraper that only checks "is this show still in the list" misses every one of those, because the show is still there, just marked. Treating absence from a list page as the sole cancellation signal also produces false cancellations whenever a listing page is paginated, re-sorted, or temporarily short a row for reasons that have nothing to do with the show itself.
Read the state on the page directly instead of inferring it from a set difference between two crawls.
def is_cancelled(page):
badge = page.query_selector("[data-testid='sale-status'][data-state='cancelled']")
return badge is not None
def reconcile_with_cancellation(page, url, previous_rows_by_url):
row = read_row(page, url)
row["cancelled"] = is_cancelled(page)
return reconcile(previous_rows_by_url, url, row)A show that disappears from a listing page and does not resolve when its own URL is visited directly is a much stronger cancellation signal than a missing row on its own. Visiting the event page and reading its badge before marking anything cancelled avoids the false positives a set difference alone produces.
Every fix above assumes the scraper runs more than once against the same list of
event URLs and remembers what it saw last time. That is a different shape from a
one-off crawl: instead of writing rows, the job writes events, new,
rescheduled, status_changed, unchanged, cancelled, each stamped with
checked_at, and the current state is just the latest event per URL rather than
the only thing kept.
def run_recheck(urls, previous_rows_by_url):
events = []
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
for url in urls:
row = read_row(page, url)
row["cancelled"] = is_cancelled(page)
events.append(reconcile(previous_rows_by_url, url, row))
return eventsThe output of one run becomes the previous_rows_by_url input to the next, and
the event log itself is the useful artifact: it is what lets a later question,
"when did this show go from on sale to sold out," get answered by reading history
instead of being asked of a page that has already moved on.
Scheduling scrapes with cron
covers running this on a recurring interval instead of by hand.
A tour-date row is not four static fields, it is a small state machine, and the part that trips up a first pass is always the part that changes after the row was first written. Stamp every status read with the time you read it so a stale sellout and a current one are distinguishable. Key rows on the event URL so a postponed show does not vanish when its date field changes underneath it. Read the on-sale timestamp explicitly so an unopened sale does not look like a broken page. Match shows across partners on artist, venue and date together, because no shared identifier exists. Record which sale window a row came from, general or presale, and read the cancelled badge directly instead of trusting an absence. None of this is exotic scraping; it is ordinary extraction plus the bookkeeping that survives a second run.
Why does my sellout status look wrong sometimes? It is probably not wrong, it is old. Store the time you read the status alongside the value, so a stale read and a current one are distinguishable instead of looking identical.
A show I was tracking disappeared from my dataset. Where did it go? Check whether it was postponed. Many event pages keep the same URL and swap the date field, so a row keyed on the original date stops matching and looks gone, while the show is still live under a new date on the same page.
Why does a newly announced show have no ticket link at all? On-sale time is often a separate future timestamp with its own countdown. Before that moment there is no link to find, which is not the same as a broken page. Read the on-sale timestamp and branch on it instead of retrying blindly.
How do I compare prices for the same show across two ticketing partners? Match on artist, venue and date together, normalized for case and punctuation. There is no identifier shared between partners, so the match has to run on the fields a person would use.
Why does the same show look sold out on one visit and available on another? A presale code opens an earlier window with different availability than the general on-sale. Record which window your scraper actually read; do not assume every row represents the same state.
A show is missing from the listing page. Is it cancelled? Not necessarily. Visit the event page directly and read its status badge. Some cancelled shows stay listed with a cancelled state rather than being removed, and some absences are just pagination.
- Playwright's own
get_attributeandquery_selector, used exactly as documented upstream to read machine-readable state rather than rendered labels; the browser returned by this library is a real PlaywrightBrowser. - This project's own configuration notes on reading state attributes over visible text for status fields, the same approach used for single-event availability tracking.
- Retrieved 2026-08-28.
See also: scraping event and ticket listings for the availability XHR and calendar-widget mechanics behind a single show's page, tracking product stock for the same diff-and-timestamp shape applied to a boolean instead of a status enum, capturing XHR and API responses for reading a countdown or availability payload directly instead of the rendered page, cleaning scraped prices and dates for turning the raw date and price strings into typed values, and scheduling scrapes with cron for running the re-check on a recurring cadence instead of by hand.
Written while maintaining invisible_playwright, a Firefox patched at the C++ level driven by stock Playwright. The first version of this scraper keyed rows on artist, venue and date; a postponed show swapped its date field on the same page and the row silently dropped out of a series that was keyed on the date I no longer had.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
- How to scrape course catalogs with Playwright
- How to scrape store locator pages with Playwright
- How to scrape stock levels with Playwright
- How to scrape accordion and tab content with Playwright
- How to scrape size charts with Playwright
- How to scrape delivery slots with Playwright
- How to scrape appointment availability with Playwright
- How to scrape auction listings with Playwright
- How to scrape public transport timetables with Playwright
- How to scrape GraphQL endpoints with Playwright
- How to scrape virtual scrolling tables with Playwright
- How to scrape shipping rates with Playwright
- How to scrape cursor-based pagination with Playwright
- How to scrape multi-select facet filters with Playwright
- How to scrape currency exchange rates with Playwright
- How to scrape WebSocket streams with Playwright
- How to scrape book metadata with Playwright
- How to scrape professional directories with Playwright
- How to scrape range slider filters with Playwright
- How to scrape currency and locale switchers with Playwright
- How to scrape software changelogs and release notes with Playwright
- How to scrape breadcrumb hierarchies with Playwright
- How to scrape microdata and RDFa markup with Playwright
- How to scrape server-sent events with Playwright
- How to scrape open data portals with Playwright
- How to scrape infinite carousels with Playwright
- How to scrape printer-friendly pages with Playwright
- How to handle A/B test variants when scraping with Playwright
- How to scrape recipe data with Playwright
- How to scrape vehicle recall notices with Playwright
- How to scrape public tender notices with Playwright
- How to scrape nutrition labels with Playwright
- How to scrape podcast episode listings with Playwright
- How to scrape weather station data with Playwright
- How to scrape newsletter archives with Playwright
- How to scrape wine and spirits catalogs with Playwright
- How to scrape insurance quotes with Playwright
- How to scrape fitness class schedules with Playwright
- How to scrape flight seat maps with Playwright
- How to scrape concert and tour dates with Playwright
- How to scrape museum and gallery exhibition dates with Playwright
- How to scrape warranty terms with Playwright
- How to scrape sortable data tables with Playwright
- How to scrape salary and pay scale data with Playwright
- How to scrape live sports scores with Playwright
- How to scrape video game prices with Playwright
- How to scrape domain WHOIS records with Playwright
- How to scrape podcast transcripts with Playwright
- How to scrape patent listings with Playwright
- How to scrape clinical trial listings with Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2023
- selenium-stealth hasn't been updated since November 2020
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder