-
Notifications
You must be signed in to change notification settings - Fork 230
how to scrape wine catalog playwright
To scrape a wine and spirits catalog with Playwright, key every row by name, vintage and bottle size together rather than by name alone, treat the size selector the way you would treat a variant selector on any e-commerce page and wait for its own price and stock response, pass the age gate once per browser context instead of solving it on every page, keep "NV" as a real vintage value rather than coercing it to null, normalize each critic score to a common scale before averaging anything, and match the proxy exit to the region the catalog is gating its assortment on, since that is a different question from whether one item happens to be in stock.
A wine catalog reads like a product catalog with a label glued on, and that reading throws away the one field the whole trade prices against. Two bottles under the identical name, from the identical producer, in the identical region, are two different products the moment their vintages differ: different stock, a different price, sometimes a rating that exists for one year and not the other. A scraper that groups rows by name alone has already merged data that the retailer, the critic and the regulator all keep apart on purpose.
A bottling's name is stable across years; almost nothing else is. The 2018 and the 2020 of the same wine can differ in price by a wide margin, sell out independently of each other, and carry a rating that one vintage earned and the other never received. Collapse them into a single row keyed on the name and you either overwrite one vintage's price with the other's or silently keep whichever the page happened to render last.
The fix is mechanical: carry vintage and size in the key from the first line of code that reads a row, not as an afterthought added once a bug report shows two prices merged into one.
from dataclasses import dataclass
@dataclass(frozen=True)
class BottleKey:
name: str
producer: str
vintage: str # "2019", or "NV" - never coerced to a number here
size_ml: int # 375, 750, 1500 ...
def row_key(cell):
return BottleKey(
name=cell["name"].strip(),
producer=cell.get("producer", "").strip(),
vintage=(cell.get("vintage") or "NV").strip(),
size_ml=int(cell.get("size_ml", 750)),
)Treat BottleKey as the unit you store, sort and deduplicate on. A dictionary keyed
only by name will happily let a later vintage clobber an earlier one during a naive
rows[name] = data write, and nothing in that line will tell you it happened.
Most catalog pages sell more than one bottle size under a single listing: a half bottle, the standard 750ml, sometimes a magnum. That size choice behaves exactly like a clothing size or a shoe width on a general retail page, a selector that swaps out price and availability underneath the same product photo. The e-commerce variant pattern applies here without modification: the price shown when the page first loads belongs to whichever size the storefront defaulted to, and every other size only prices itself once you select it and the request fires.
from invisible_playwright import InvisiblePlaywright
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
page.goto("https://example.com/product/456")
sizes = page.locator("[data-size-option]")
rows = []
for i in range(sizes.count()):
option = sizes.nth(i)
label = option.get_attribute("data-size-option") # "375ml", "750ml", "1.5L"
with page.expect_response(
lambda r: "price" in r.url or "availability" in r.url
) as caught:
option.click()
payload = caught.value.json()
rows.append({"size_ml": label, "price": payload.get("price"),
"in_stock": payload.get("inStock")})Read the page once and stop, and you have captured one size's price mislabeled as "the" price for the whole listing. The magnum is not an accessory field on the 750ml row; it is its own row with its own price and its own stock count.
Some regions do not gate on age alone. A spirit above a certain proof, or a vintage tied to a futures allocation, can trip a stricter legal path than an ordinary bottle of the same name at a lower strength or a later release. The catalog reacts by hiding the price entirely, showing a "contact to purchase" label, or refusing to render the product page until a stronger age or residency check passes, and it can do this for one SKU on a listing while a neighboring size or vintage of the same product sails through untouched.
A missing price on one row and a normal price on the sibling row of the same listing is not a broken selector. Check the ABV and vintage of the specific row before assuming the scraper failed, because the site made a per-product legal decision, not a general one.
Passing the gate is a session fact, not a page fact. The confirmation writes a cookie or a local storage flag to the browser context, and every later page opened in that same context inherits it automatically. Re-running the click-through logic on every catalog page you visit is wasted work at best, and at worst it looks like a bot that never remembers anything between requests, which is its own kind of tell.
GATE_COOKIE_NAMES = {"age_verified", "over21"}
def pass_age_gate_once(context, start_url):
page = context.new_page()
page.goto(start_url)
gate = page.get_by_role("button", name="Yes, I am of legal age")
if gate.count():
gate.click()
page.wait_for_load_state("networkidle")
cookies = {c["name"] for c in context.cookies()}
if not cookies & GATE_COOKIE_NAMES:
raise RuntimeError("gate did not set a session cookie; check the selector")
page.close()Call this once, then open every remaining product page from context.new_page()
inside that same context and let the cookie ride along. The mechanics of reading and
seeding that cookie jar, and the caveat that a hand-set cookie can look unearned next
to one a real click produced, are covered in
reading and setting cookies in a Playwright context.
A single product page routinely stacks scores from several independent critics or publications, and each one brought its own scale. A 100-point score, a 20-point score from an older European tradition, and a star rating with half-step increments can all sit in the same ratings block, describing the same bottle.
| Scale | Typical range | Example value |
|---|---|---|
| 100-point | 50-100 | 94 |
| 20-point | 0-20 | 17.5 |
| Star | 0-5, often half steps | 4.5 |
Average 94 and 4.5 directly and the star score looks like it is describing a
mediocre bottle next to an excellent one, when both actually mean roughly the same
thing on their own scale. The averaging step has to normalize first, and the
normalized value should sit next to the raw one rather than replace it, because a
reader who trusts one critic's scale specifically wants that original number kept.
def normalize_score(value, scale):
if scale == "100pt":
return value
if scale == "20pt":
return (value / 20) * 100
if scale == "5star":
return (value / 5) * 100
raise ValueError(f"unknown rating scale: {scale}")
def blended_score(ratings):
# ratings: list of {"value": float, "scale": str, "critic": str}
normalized = [normalize_score(r["value"], r["scale"]) for r in ratings]
return sum(normalized) / len(normalized) if normalized else NoneStore ratings as the list it is, not as a single averaged number thrown away after
the fact. The blended score is a convenience field for sorting, the per-critic list
is the data.
"NV" on a sparkling wine or a blended spirit means the producer mixed several years
on purpose, and it is exactly as real a value as "2019" is. A parser that tries
int(vintage) and falls back to None on failure turns every NV bottle into a
missing value, and a missing value drops silently out of any sort or any filter that
expects a year.
def vintage_sort_key(vintage):
if vintage == "NV":
return (0, 0) # NV bottles sort together, deliberately first
return (1, int(vintage)) # everything else sorts by year
rows.sort(key=lambda r: vintage_sort_key(r["vintage"]))A filter for "vintage 2015 or later" needs the same explicit decision: does NV pass
that filter or not. Pick the rule once, in one place, rather than letting the answer
depend on whatever int("NV") happens to raise deep inside a sort call.
Locale-correct handling of the price and date fields sitting next to vintage follows
the same principle of not guessing at a format; see
cleaning scraped prices and dates
for the parsing side of that.
A product being in or out of stock, covered already for variant pricing above, is a different question from whether the catalog will show you the product at all. Wine and spirits importers hold regional licenses, and a retailer's catalog can vary by detected buyer location: some spirits are simply absent from the assortment outside their licensed territory, independent of anything the stock field says. Querying from an exit in the wrong region can produce a thinner catalog, a different set of prices, or a bottle that appears to not exist at all, none of which is a parsing bug.
Match the proxy exit to the region the catalog is meant for and let the timezone and locale follow that exit instead of pinning them by hand. The full set of surfaces that has to agree with the exit for a geofenced catalog, and why fixing the IP alone makes the mismatch worse rather than better, is covered in scraping geotargeted content. The day-to-day question of watching one item's price and stock over time, as distinct from the assortment question here, belongs to tracking product prices.
Every piece above lands in one function: the identity key, the size loop, the vintage-specific price and stock, and the normalized score, combined into one row per SKU rather than one row per listing.
def build_rows(context, product_url):
page = context.new_page()
page.goto(product_url, wait_until="domcontentloaded")
base = read_product_summary(page) # name, producer, abv
for size_row in read_size_variants(page): # the loop from above
for vintage_row in read_vintage_variants(page, size_row):
yield {
"name": base["name"],
"producer": base["producer"],
"vintage": vintage_row["vintage"], # "NV" stays "NV"
"size_ml": size_row["size_ml"],
"price": vintage_row["price"],
"in_stock": vintage_row["in_stock"],
"score": blended_score(vintage_row["ratings"]),
"ratings": vintage_row["ratings"], # keep the raw scale list
}
page.close()Every yielded dict is one vintage of one size of one product, which is the grain buyers actually compare against. Anything coarser than that throws away the exact dimension the catalog was built around.
A wine and spirits catalog fails scrapers that treat it like a normal product grid, because the field that actually varies, vintage, does not look like a variant to someone skimming the page once. Key rows by name, vintage and size together, read the size selector as the variant pattern it is, pass the age gate once per context instead of on every request, keep NV as a value instead of a null, normalize scores before you average them, and match your exit to the region the assortment is gated on. Do all of that and the row you extract is the actual product a buyer would see, not a name with the wrong year's price stapled to it.
Why can't I group rows by wine name alone? Because vintage changes the price, the stock level and sometimes the rating, and two vintages under one name key will overwrite each other the moment you write a row for the second one.
Why does the price differ between two bottles that look identical on the page? They are usually different vintages or different bottle sizes rendered under the same product photo. Check both fields before assuming a scraping error.
Why did the age gate block the whole page instead of just hiding the price? ABV and vintage together can trigger a stricter legal path for a specific SKU, and the site can choose to block the page entirely rather than partially disclose it.
Do I need to solve the age gate on every product page? No. It sets a cookie or a local storage flag on the browser context, and every page opened afterward in that same context inherits it. Solve it once per context, not once per request.
Should I average a 92-point score with a 4.5-star score directly? No. Normalize both to the same scale first. Averaged as-is, the star rating reads as much weaker than it actually is.
What do I do with "NV" in the vintage field? Keep it as the string "NV" and give it an explicit place in your sort order and your filters. Coercing it to an integer and catching the failure as a missing value drops those rows out of anything that depends on vintage.
Why does the catalog itself look different from another location, not just the prices? Availability is geofenced independent of stock. A licensed spirit in one region can be entirely absent from the assortment in another, which is a proxy-exit question, not a parsing question.
- Playwright's
expect_responseandget_by_role, used as documented upstream to catch the variant and age-gate requests. - Playwright's
context.cookies(), which is how an age-gate confirmation is verified as a context-level fact rather than assumed from a page redirect. - This project's own configuration behaviour: locale, timezone and number format follow the proxy exit by default, which is what keeps a geofenced catalog's surfaces in agreement without hand-pinning any of them.
See also: scraping e-commerce product pages for the variant XHR pattern bottle size reuses, reading and setting cookies in a Playwright context for the age-gate session mechanics, scraping geotargeted content for the surfaces that must agree with a region-locked exit, and cleaning scraped prices and dates for parsing the fields that sit next to vintage on the same row.
Written while maintaining invisible_playwright, a Firefox patched at the C++ level driven by stock Playwright. The row key on an early version of this pattern was name and producer alone, and it ran for months before anyone noticed a later vintage had been silently overwriting an earlier one's price on every bottling that got re-released.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
- How to scrape course catalogs with Playwright
- How to scrape store locator pages with Playwright
- How to scrape stock levels with Playwright
- How to scrape accordion and tab content with Playwright
- How to scrape size charts with Playwright
- How to scrape delivery slots with Playwright
- How to scrape appointment availability with Playwright
- How to scrape auction listings with Playwright
- How to scrape public transport timetables with Playwright
- How to scrape GraphQL endpoints with Playwright
- How to scrape virtual scrolling tables with Playwright
- How to scrape shipping rates with Playwright
- How to scrape cursor-based pagination with Playwright
- How to scrape multi-select facet filters with Playwright
- How to scrape currency exchange rates with Playwright
- How to scrape WebSocket streams with Playwright
- How to scrape book metadata with Playwright
- How to scrape professional directories with Playwright
- How to scrape range slider filters with Playwright
- How to scrape currency and locale switchers with Playwright
- How to scrape software changelogs and release notes with Playwright
- How to scrape breadcrumb hierarchies with Playwright
- How to scrape microdata and RDFa markup with Playwright
- How to scrape server-sent events with Playwright
- How to scrape open data portals with Playwright
- How to scrape infinite carousels with Playwright
- How to scrape printer-friendly pages with Playwright
- How to handle A/B test variants when scraping with Playwright
- How to scrape recipe data with Playwright
- How to scrape vehicle recall notices with Playwright
- How to scrape public tender notices with Playwright
- How to scrape nutrition labels with Playwright
- How to scrape podcast episode listings with Playwright
- How to scrape weather station data with Playwright
- How to scrape newsletter archives with Playwright
- How to scrape wine and spirits catalogs with Playwright
- How to scrape insurance quotes with Playwright
- How to scrape fitness class schedules with Playwright
- How to scrape flight seat maps with Playwright
- How to scrape concert and tour dates with Playwright
- How to scrape museum and gallery exhibition dates with Playwright
- How to scrape warranty terms with Playwright
- How to scrape sortable data tables with Playwright
- How to scrape salary and pay scale data with Playwright
- How to scrape live sports scores with Playwright
- How to scrape video game prices with Playwright
- How to scrape domain WHOIS records with Playwright
- How to scrape podcast transcripts with Playwright
- How to scrape patent listings with Playwright
- How to scrape clinical trial listings with Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2023
- selenium-stealth hasn't been updated since November 2020
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder