-
Notifications
You must be signed in to change notification settings - Fork 230
how to scrape podcast transcripts playwright
To scrape podcast transcripts with Playwright, find the direct link or API response carrying the VTT or SRT file before touching the rendered panel, parse it into cues that each keep a start time, an end time and text, normalize every timestamp to seconds while keeping the original string, and record whether the transcript is auto-generated or human-edited, because that one fact decides whether the text is safe to quote by the sentence or only searchable as a block.
This is not the episode-listing problem. Finding a show's episodes is about the feed, the guid and the archive. A transcript lives one level down, inside a single episode, and it is a different shape of data entirely: not a list of items but a list of timed cues, and the timing is the part worth keeping. Flatten a transcript into one string too early and you can still search it, but you can no longer answer "where in the episode does she say this", which is the question a transcript is usually scraped to answer.
The rendered transcript panel looks like a wall of text, sometimes broken into speaker turns, sometimes not. Underneath, it is built from a sequence of cues, and each cue is a small record: a start time, often an end time, and the words spoken in that span. A flat string can be searched. A list of cues can be searched AND jumped to, because every match still carries the moment it happened, which is the entire value of scraping a transcript instead of copying it by hand.
Losing the cue boundaries is a one-way trip. Once forty cues have been joined into one paragraph with a space between them, there is no reliable way to work backward to "this sentence started at 14:32". Store the cue list itself, one row per cue, and generate any flat text view from that afterward, never the other way around.
# The unit worth keeping. Every field below survives to the stored row.
{
"start_seconds": 872.4,
"end_seconds": 876.1,
"start_raw": "00:14:32.400",
"speaker": "Host", # None if the transcript does not label speakers
"speaker_source": "field", # "field", "inline", or "absent"
"text": "So that's when we decided to rebuild the whole pipeline.",
}Plenty of podcast players render the transcript panel lazily, drawing only the cues near
the currently playing timestamp and swapping them out as playback moves forward. Reading
that panel means driving playback, or faking a scroll and a wait for every few seconds of
audio, for the length of the episode. An hour-long show becomes an hour-long scrape, and
a paused or muted player often renders nothing at all because the panel is wired to the
timeupdate event, not to a click.
The file behind that panel is almost always available as one request. Podcast platforms
that support transcripts typically expose the cue file as a .vtt or .srt document,
either linked directly in the page or returned by an API call the player itself makes to
populate the panel. Fetching that file is one request against a static document instead of
minutes of simulated scrolling, and capturing the underlying API response
is the general technique for finding it when there is no visible link.
import re
from invisible_playwright import InvisiblePlaywright
TRANSCRIPT_LINK = "link[type='text/vtt'], a[href$='.vtt'], a[href$='.srt']"
TRANSCRIPT_IN_RESPONSE = re.compile(r"\.(?:vtt|srt)(?:\?|$)")
def fetch_vtt_or_srt(page, url):
caught = {}
def watch(response):
if TRANSCRIPT_IN_RESPONSE.search(response.url):
caught["body"] = response.text()
page.on("response", watch)
page.goto(url, wait_until="networkidle")
node = page.query_selector(TRANSCRIPT_LINK)
if node and "body" not in caught:
caught["body"] = page.request.get(node.get_attribute("href")).text()
if "body" not in caught:
raise SystemExit("no VTT/SRT found; the panel is the only source")
return caught["body"]
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
cue_text = fetch_vtt_or_srt(page, "https://example.com/show/episode-42")Attach page.on("response", ...) before the navigation, or a request fired during page
load is missed. Try the link only as a fallback: the response listener already caught the
file if the player fetched it itself, and the link check costs nothing when it did not.
WebVTT and SRT differ in punctuation but agree on structure: a block of numbered or timestamped entries, each with a time range on its own line and the spoken text below it. Parse both into the same row shape so the rest of the pipeline does not care which format a given show used.
import re
TIME_LINE = re.compile(
r"(\d{2}):(\d{2}):(\d{2})[.,](\d{3})\s*-->\s*"
r"(\d{2}):(\d{2}):(\d{2})[.,](\d{3})"
)
def timestamp_to_seconds(h, m, s, ms):
return int(h) * 3600 + int(m) * 60 + int(s) + int(ms) / 1000
def parse_cues(raw_text):
"""Works on VTT (dot milliseconds) and SRT (comma milliseconds) alike."""
cues = []
lines = raw_text.replace("\r\n", "\n").split("\n")
i = 0
while i < len(lines):
match = TIME_LINE.search(lines[i])
if not match:
i += 1
continue
start_raw = lines[i].split("-->")[0].strip()
h1, m1, s1, ms1, h2, m2, s2, ms2 = match.groups()
text_lines = []
i += 1
while i < len(lines) and lines[i].strip() and not TIME_LINE.search(lines[i]):
text_lines.append(lines[i].strip())
i += 1
cues.append({
"start_seconds": timestamp_to_seconds(h1, m1, s1, ms1),
"end_seconds": timestamp_to_seconds(h2, m2, s2, ms2),
"start_raw": start_raw,
"text": " ".join(text_lines).strip(),
})
return cuesTwo things matter here more than the regex itself. The parser tolerates both delimiter
styles in the same function, because a batch of episodes from different tools will mix
them. And start_raw is kept next to start_seconds: seconds are what you sort, filter
and jump to, but the source string is what you show back to a person or hand to a media
player's own seek parameter, and reconstructing it from a float loses precision that the
original never had.
Speaker attribution shows up in two different shapes, and a transcript uses one or the
other, rarely both. Some feeds mark the speaker as a distinct field or a <v Speaker>
tag inside VTT. Others bake it into the text itself, as Host: welcome back at the start
of a line, with no separate field at all.
Detecting which form a file uses is a cheap check: try the structured tag first, and only fall back to the inline pattern if the structured one is absent for the whole file. Guessing per cue instead of per file causes worse damage than either format alone, because a normal sentence that happens to start with a capitalized word and a colon, "Note: this changed in season two", gets misread as a speaker turn.
import re
VTT_VOICE_TAG = re.compile(r"^<v\s+([^>]+)>(.*)$")
INLINE_SPEAKER = re.compile(r"^([A-Z][\w .]{1,24}):\s*(.+)$")
def detect_speaker_form(cues):
"""One decision for the whole file, not one guess per cue."""
if any(VTT_VOICE_TAG.match(c["text"]) for c in cues):
return "tag"
hits = sum(1 for c in cues if INLINE_SPEAKER.match(c["text"]))
# require a clear majority, or a handful of coincidental matches
# gets treated as if every cue in the file were labeled
if hits >= max(3, len(cues) * 0.4):
return "inline"
return "absent"
def apply_speaker_form(cues, form):
for cue in cues:
if form == "tag":
match = VTT_VOICE_TAG.match(cue["text"])
cue["speaker"] = match.group(1) if match else cue.get("speaker")
cue["text"] = match.group(2).strip() if match else cue["text"]
cue["speaker_source"] = "tag" if match else "absent"
elif form == "inline":
match = INLINE_SPEAKER.match(cue["text"])
cue["speaker"] = match.group(1) if match else None
cue["text"] = match.group(2).strip() if match else cue["text"]
cue["speaker_source"] = "inline" if match else "absent"
else:
cue["speaker"] = None
cue["speaker_source"] = "absent"
return cuesStore speaker_source alongside speaker. A cue whose label came from a distinct tag is
a claim the transcription tool made deliberately; a cue whose label came out of a colon
match is an inference, and knowing which one you have matters the first time two speakers'
turns run together without a blank line between them.
The single fact worth capturing before anything else is whether a transcript came from speech-to-text or from a person editing it. Auto-generated transcripts frequently drop speaker labels entirely, punctuate on pauses rather than grammar, and mis-hear proper nouns and technical terms. Human-edited transcripts usually carry reliable punctuation and correctly labeled turns, because someone read the audio against the text.
Most platforms that distinguish the two say so somewhere on the page, as a badge, a caption near the transcript tab, or a field in the API response that served the VTT file. Look for it before you parse: a human-edited transcript can be quoted sentence by sentence with confidence in the wording, while an auto-generated one is safer treated as a fuzzy full-text index into the audio, where any one cue's exact phrasing is a probability rather than a fact.
def transcript_quality(page):
"""Look for an explicit signal before assuming either way."""
badge = page.query_selector("[data-transcript-type], .transcript-badge")
if badge:
label = (badge.get_attribute("data-transcript-type")
or badge.text_content() or "").strip().lower()
if "auto" in label or "generated" in label:
return "auto_generated"
if "edited" in label or "human" in label:
return "human_edited"
return "unknown"Store the result as its own field next to the cue list, not folded into a comment.
"unknown" is a real, honest value here. Guessing auto-generation from the text quality
alone, no punctuation, no speaker labels, is possible but noisy enough that it belongs in
a separate, clearly-named heuristic field rather than mixed with a page's own declared
signal.
Bring the pieces together into one row per episode: the cue list with timing, speakers and text, plus the two facts that describe the whole transcript rather than one cue. Nothing here needs its own network round trip beyond the single file fetch above.
def scrape_transcript(page, url):
cue_text = fetch_vtt_or_srt(page, url) # the fetch shown earlier
cues = parse_cues(cue_text)
form = detect_speaker_form(cues)
cues = apply_speaker_form(cues, form)
return {
"cues": cues,
"speaker_form": form,
"quality": transcript_quality(page),
"cue_count": len(cues),
}cue_count is a cheap sanity check worth keeping on the row: a transcript that parsed to
zero or one cue almost always means the fetch returned an error page or an empty file
rather than a real absence of cues, and it is easier to catch that from a stored number
than from re-reading every row by hand later.
Some players draw the transcript as an image, generated server-side as part of a "shareable quote" card, with no text node and no underlying file behind it. Others reveal transcript text only while audio is actively playing, tied to a playback event with no static document and no API response carrying the full cue list ahead of time. Neither case has a file to fetch, and neither has a DOM node worth reading.
That is a real limit, not a missing selector. Re-transcribing the audio yourself, running it through a speech-to-text model, is a legitimate answer to that gap, but it is a different technique with its own accuracy and cost tradeoffs, and it is out of scope for a page about reading a transcript the page already produced.
A podcast transcript is worth more as a list of timed cues than as a paragraph, and the fastest way to get that list is the VTT or SRT file behind the panel, not the panel itself. Fetch that file first, parse both timestamp styles into seconds while keeping the source string, detect whether speaker labels sit in their own field or inside the text, and record whether the transcript came from speech-to-text or a human editor, because that single fact tells the next person whether a cue is a quote or an approximation. When there is no file and no text node, a rendered image or a playback-gated panel, stop there: that gap is a different job.
Should I scroll the transcript panel to read it? Only when there is nothing else.
Many panels lazy-render near the current playback position, which makes reading the whole
transcript as slow as playing the episode. Check for a direct .vtt/.srt link or the
API response that feeds the panel first.
How do I normalize SRT and VTT timestamps together? Convert both to seconds with one
parser: VTT uses a dot before milliseconds, SRT uses a comma, and both otherwise follow
HH:MM:SS. Keep the original string in a separate field, because a media player's seek
parameter and a citation both want the source format, not a recomputed one.
How do I know if speaker labels are a field or part of the text? Check the whole file
for a structured tag first, such as VTT's <v Speaker>. Only fall back to matching
"Name:" at the start of a line if no structured tag appears anywhere, and require several
matches before trusting the inline pattern, since one coincidental colon is not a
transcript convention.
Why does it matter if a transcript is auto-generated? Auto-generated transcripts often lack speaker labels and can mis-hear words, so a single cue's exact wording is closer to a probability than a fact. Human-edited transcripts are safer to quote sentence by sentence. Record which kind you got when the page says so.
What if there is no file and the panel needs audio playing to show text? That is a real stop point. A transcript rendered only as an image, or revealed only during active playback with no static file or API response, has to be re-transcribed from the audio, which is a different task than scraping text the page already has.
- The WebVTT specification, including the
<v Speaker>voice tag and theHH:MM:SS.mmmtimestamp format used for cue timing. - Playwright's
Page.on("response")andAPIRequestContext.get(), used exactly as documented upstream, since the browser this library returns is a real PlaywrightBrowser. Retrieved 2026-08-28. - SRT's
HH:MM:SS,mmm --> HH:MM:SS,mmmcue format, the comma-delimited counterpart to WebVTT's dot, as implemented by every common SRT-producing tool this timestamp regex was checked against.
See also: scraping podcast episode listings for finding the feed and keying episodes before you ever reach a transcript, capturing XHR and API responses for the general technique behind finding the VTT/SRT request, cleaning scraped prices and dates for the same normalize-and-keep-the-source habit applied elsewhere, and scraping video listings and metadata for the sibling problem on video pages that carry captions instead of an audio transcript.
Written while maintaining invisible_playwright, a Firefox patched at the C++ level driven by stock Playwright. The first version of this scraper joined every cue into one string before storing it, and the day someone asked "what minute does she mention the merger" there was no way to answer without re-parsing the original file from scratch.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
- How to scrape course catalogs with Playwright
- How to scrape store locator pages with Playwright
- How to scrape stock levels with Playwright
- How to scrape accordion and tab content with Playwright
- How to scrape size charts with Playwright
- How to scrape delivery slots with Playwright
- How to scrape appointment availability with Playwright
- How to scrape auction listings with Playwright
- How to scrape public transport timetables with Playwright
- How to scrape GraphQL endpoints with Playwright
- How to scrape virtual scrolling tables with Playwright
- How to scrape shipping rates with Playwright
- How to scrape cursor-based pagination with Playwright
- How to scrape multi-select facet filters with Playwright
- How to scrape currency exchange rates with Playwright
- How to scrape WebSocket streams with Playwright
- How to scrape book metadata with Playwright
- How to scrape professional directories with Playwright
- How to scrape range slider filters with Playwright
- How to scrape currency and locale switchers with Playwright
- How to scrape software changelogs and release notes with Playwright
- How to scrape breadcrumb hierarchies with Playwright
- How to scrape microdata and RDFa markup with Playwright
- How to scrape server-sent events with Playwright
- How to scrape open data portals with Playwright
- How to scrape infinite carousels with Playwright
- How to scrape printer-friendly pages with Playwright
- How to handle A/B test variants when scraping with Playwright
- How to scrape recipe data with Playwright
- How to scrape vehicle recall notices with Playwright
- How to scrape public tender notices with Playwright
- How to scrape nutrition labels with Playwright
- How to scrape podcast episode listings with Playwright
- How to scrape weather station data with Playwright
- How to scrape newsletter archives with Playwright
- How to scrape wine and spirits catalogs with Playwright
- How to scrape insurance quotes with Playwright
- How to scrape fitness class schedules with Playwright
- How to scrape flight seat maps with Playwright
- How to scrape concert and tour dates with Playwright
- How to scrape museum and gallery exhibition dates with Playwright
- How to scrape warranty terms with Playwright
- How to scrape sortable data tables with Playwright
- How to scrape salary and pay scale data with Playwright
- How to scrape live sports scores with Playwright
- How to scrape video game prices with Playwright
- How to scrape domain WHOIS records with Playwright
- How to scrape podcast transcripts with Playwright
- How to scrape patent listings with Playwright
- How to scrape clinical trial listings with Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2023
- selenium-stealth hasn't been updated since November 2020
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder