-
Notifications
You must be signed in to change notification settings - Fork 230
how to scrape server sent events playwright
To scrape server-sent events with Playwright, do not reach for a network event at all:
install a page.add_init_script that replaces window.EventSource and wraps fetch
before any of the page's own scripts run, pair it with page.expose_function so each
message crosses back into Python as it arrives, and stamp every row with a UTC receive
time and the last event id so a reconnect resumes instead of duplicating. Playwright
exposes WebSocket frames on a first-class event and exposes no equivalent for SSE at any
level of the public API, so the hook that reads one is not the hook that reads the other.
SSE has a reputation as the easy one. One direction, plain text, a documented line format,
no handshake, no binary frames. All true of the protocol, and none of it true of the
tooling. For a Playwright scraper the two invert: a socket hands you framereceived with
the payload sitting in the argument, while an event stream hands you a request whose
resource_type is eventsource and then goes quiet for as long as the stream lives.
This page is what fills that gap. Why the obvious hook hangs instead of failing, why the route everybody recommends is narrower than it looks, the instrumentation that works on all three browsers, and a parser faithful enough to the spec that a comment line and an id with no data both do the right thing.
The asymmetry is worth stating precisely, because it is the whole reason this page exists.
page.on("websocket") fires once per socket and hands you a WebSocket object with
framesent and framereceived on it, each carrying the payload as an argument. That is a
message-level API, and
scraping WebSocket streams is mostly a
matter of using it well.
There is no page.on("eventsource"). There is no message event on any object Playwright
hands you for a stream. What does exist is the classification: eventsource is a real
resource_type, sitting in the same list as xhr, fetch, websocket and the rest. So
you can see the stream open, count the streams a page holds, and block one with
page.route. You cannot read a single message of it. The support stops at the request.
page.on("response") does fire for an SSE request, and it fires early. Playwright emits
that event when the status and headers arrive, not when the body does, which for an
ordinary document is a distinction nobody notices. Here it is the whole trap. The listener
runs, the URL is right, the content type says text/event-stream, and the capture looks
like it is working.
Then response.body() waits for a body that never completes, because a live stream not
completing is the point of it. Nothing raises. No timeout you configured applies. The
script sits there.
from invisible_playwright import InvisiblePlaywright
def on_response(response):
if response.request.resource_type != "eventsource":
return
print("stream seen:", response.status, response.headers.get("content-type"))
# Do NOT do this on a live stream. The body never finishes, so the call
# does not raise and does not time out. It just never comes back.
# body = response.body()
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
page.on("response", on_response) # fires on headers, not on the body
page.goto("https://example.com/live")
page.wait_for_timeout(10_000)The call that would fix this is Response.bodyAsStream(), and it does not exist.
Playwright issue 17199 asked for it on 2022-09-08. At the retrieval date it is still open
and labelled P3-collecting-feedback, so there is no supported way to read a response body
incrementally, and no partial read to fall back on.
One caution, stated plainly because I would rather not assert something I have not run.
The hang follows from response.finished() semantics, which resolve when the response
finishes, plus that open issue, and it is widely reported by people who hit it. It is a
strong inference and not a measurement of mine. Spend ten lines reproducing it against
your own stream before you build anything on the claim, and put a timeout around the
reproduction so confirming it costs a few seconds rather than an evening. For requests
that do end, capturing XHR and API responses
is the hook that works.
The Chrome DevTools Protocol carries the answer people find first.
Network.eventSourceMessageReceived fires with exactly five fields: requestId, timestamp,
eventName, eventId and data. That is the parsed frame, already split at the colon and
already stripped, delivered without any parsing work on your side. It is a good event, and
recommending it flatly is still wrong, because it is restricted twice.
The first restriction is the browser. Playwright states that CDP sessions are only supported on Chromium-based browsers, so the route does not exist on Firefox or WebKit at all, which is one more entry on the ledger between BiDi and CDP.
The second restriction is the one that hurts. The event only fires for real EventSource
objects, and Chromium issue 40659493 is titled "DevTools: XHR and fetch to
text/event-stream resources don't show events". A large share of modern SSE traffic never
constructs an EventSource, because EventSource cannot send a custom header and fetch
can, so anything needing an authorization header reads its stream with fetch and a
ReadableStream. Chat interfaces that stream tokens are almost all in that group. On
those pages CDP reports nothing, and reports it silently.
The portable route is instrumentation rather than interception, and it rests on two
documented guarantees. page.add_init_script is evaluated after the document has been
created but before any of the page's scripts have run, which is early enough to swap
window.EventSource before page code takes a reference to it. page.expose_function adds
a function on the window object of every frame, and that binding outlives every page load
in the frame, so each message crosses into Python without polling.
from datetime import datetime, timezone
from invisible_playwright import InvisiblePlaywright
EVENTSOURCE_HOOK = r"""
(() => {
const Native = window.EventSource;
if (!Native) return;
function Wrapped(url, init) {
const es = new Native(url, init);
const add = EventTarget.prototype.addEventListener.bind(es);
const mirrored = new Set();
const watch = (type) => {
if (mirrored.has(type)) return;
mirrored.add(type);
add(type, (ev) => window.__sse({
url: String(url),
event: type,
data: ev.data,
last_event_id: ev.lastEventId || "",
}));
};
// es.onmessage = fn never passes through addEventListener, so subscribe to
// the default type here instead of waiting for the page to ask for it.
watch("message");
// then mirror every custom type, at the moment the page subscribes to it
es.addEventListener = function (type, fn, opts) {
watch(type);
return add(type, fn, opts);
};
return es;
}
Wrapped.prototype = Native.prototype;
Wrapped.CONNECTING = 0; Wrapped.OPEN = 1; Wrapped.CLOSED = 2;
window.EventSource = Wrapped;
})();
"""
def on_message(msg):
msg["received_at"] = datetime.now(timezone.utc).isoformat(timespec="milliseconds")
print(msg)
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
page.expose_function("__sse", on_message) # binding outlives page loads
page.add_init_script(EVENTSOURCE_HOOK) # runs before the page's scripts
page.goto("https://example.com/stream")
page.wait_for_timeout(30_000)Both registrations go before goto, or the page opens its stream against the native
constructor and the hook mirrors nothing. Two lines in that script are the difference
between catching everything and catching half. Subscribing to message up front covers
the pages that assign onmessage directly, which never reaches addEventListener.
Wrapping addEventListener covers the custom event names, which you cannot know in
advance and which carry most of the interesting payloads.
Say the cost out loud: this replaces a window global, so a page that inspects
EventSource finds something whose toString does not report native code.
The fetch path needs a different shape, because there is no object to replace. What there
is instead is a response body you can copy. ReadableStream.tee() splits it into two
branches: read one, hand the other back to the page, and the page behaves as though
nothing happened.
FETCH_HOOK = r"""
(() => {
const nativeFetch = window.fetch;
window.fetch = async function (input, init) {
const res = await nativeFetch(input, init);
const ctype = res.headers.get("content-type") || "";
if (!res.body || !ctype.includes("text/event-stream")) return res;
const url = typeof input === "string" ? input : (input.url || String(input));
const [ours, theirs] = res.body.tee(); // one branch each
const reader = ours.getReader();
const decoder = new TextDecoder();
(async () => {
for (;;) {
const { value, done } = await reader.read();
if (done) { window.__sseRaw({ url: url, chunk: "", done: true }); return; }
// drain promptly: tee buffers whatever the slower branch has not taken
window.__sseRaw({
url: url,
chunk: decoder.decode(value, { stream: true }),
done: false,
});
}
})();
return new Response(theirs, {
status: res.status, statusText: res.statusText, headers: res.headers,
});
};
})();
"""Two limits, both worth knowing before you ship it. A tee holds whatever the slower branch
has not consumed, so a reader that stalls turns into memory growth on a long stream, which
is why the loop above forwards each chunk and does no work of its own. And the
reconstructed Response is not the original: it reports an empty url, loses
redirected, and its type is default. A page that reads any of those will notice. When
that happens, keep the EventSource hook and drop the fetch one rather than fighting it.
The fetch branch gives you bytes, so the parsing is yours, and the format has six rules
that each cost you data when skipped. Lines split at the first colon, not at every colon,
so a data line holding JSON survives. Exactly one leading space is removed from the
value, not all whitespace, so a payload that begins with two spaces keeps one. A line
starting with : is a comment and is ignored, which is how servers send keepalives, and a
parser that treats it as data fills your file with blanks. A data field appends its
value and then a single LF, so consecutive data lines concatenate with newlines between
them. An id sets the last event ID unless the value contains a NULL, in which case the
whole line is ignored. A retry sets the reconnection time only if the value is entirely
ASCII digits, and unknown field names are ignored outright.
import re
LINE_BREAK = re.compile(r"\r\n|\r|\n") # the three the spec allows, and only these
def parse_sse(chunk, state):
"""Yield one dict per dispatched event. state carries across chunks:
{"buf": "", "data": "", "type": "", "id": None, "retry": None}"""
buf, hold = state["buf"] + chunk, ""
if buf.endswith("\r"):
buf, hold = buf[:-1], "\r" # may be the first half of a CRLF pair
lines = LINE_BREAK.split(buf)
state["buf"] = lines.pop() + hold # the tail may be half a line
for line in lines:
if line == "": # blank line dispatches
if state["data"] == "": # empty buffer: no event,
state["type"] = "" # but the id stays set
continue
payload = state["data"]
if payload.endswith("\n"):
payload = payload[:-1] # remove exactly one LF
yield {"event": state["type"] or "message", # default type
"data": payload,
"id": state["id"],
"retry_ms": state["retry"]}
state["data"], state["type"] = "", ""
continue
if line.startswith(":"): # comment: the keepalive
continue
field, _, value = line.partition(":") # first colon only
if value.startswith(" "):
value = value[1:] # one space, not lstrip()
if field == "data":
state["data"] += value + "\n" # value, then a single LF
elif field == "event":
state["type"] = value
elif field == "id":
if "\x00" not in value: # a NULL voids the line
state["id"] = value
elif field == "retry":
if value.isascii() and value.isdigit(): # ASCII digits or nothing
state["retry"] = int(value)
# every other field name is ignoredTwo shapes in there are not decoration. str.splitlines() also breaks on form feed and on
U+2028, both of which can appear inside a JSON payload, so the split is a regex over the
three terminators the spec names. And an empty data buffer returns without dispatching,
while the last event ID is set before that return, so an id line arriving with no data
still moves your resumption point forward even though no message reaches the page.
Reconnection is automatic and it is where a naive capture doubles its rows. When the
browser reconnects it sends a Last-Event-ID header carrying the last id it saw, and a
well-behaved server replays from there. Those replayed messages are real messages with
real ids, so the deduplication key is the id, not the arrival order and not the payload
hash.
The delay before that reconnect has no specified value. The spec calls it
implementation-defined and describes it only as probably in the region of a few seconds,
so a test that asserts a number is asserting your browser's build rather than the
protocol. A retry field overrides it, and only when the value is all ASCII digits.
The ways a stream stops are worth telling apart, because from Python they look identical.
readyState is CONNECTING at 0, OPEN at 1 and CLOSED at 2. If the status is not 200, or
the content type is not exactly text/event-stream, the connection fails, and once the
user agent has failed the connection it does not attempt to reconnect. HTTP 204 is the
documented way for a server to tell a client to stop reconnecting, so a capture that goes
quiet may have been dismissed rather than broken. Record the status and the last id you
held, and the three cases separate themselves, which is the same bookkeeping that makes
an interrupted scrape resumable.
The pieces assemble into one object: a parser state per stream, a seen set keyed on stream and id, a counter per event type, a deadline, and a recorded reason for stopping.
import collections, json, time
from datetime import datetime, timezone
from invisible_playwright import InvisiblePlaywright
class StreamRecorder:
def __init__(self, path, max_seconds=120):
self.out = open(path, "a", encoding="utf-8")
self.deadline = time.monotonic() + max_seconds
self.state, self.seen = {}, set()
self.kinds = collections.Counter()
self.reason = "running"
def write(self, url, event, data, event_id):
if event_id and (url, event_id) in self.seen: # replayed after a reconnect
self.kinds["replayed"] += 1
return
if event_id:
self.seen.add((url, event_id))
self.kinds[event] += 1
row = {"received_at": datetime.now(timezone.utc).isoformat(timespec="milliseconds"),
"stream": url, "event": event, "id": event_id or None, "data": data}
print(json.dumps(row), file=self.out, flush=True)
def on_message(self, msg): # from the EventSource wrapper, already parsed
self.write(msg["url"], msg["event"], msg["data"], msg["last_event_id"])
def on_raw(self, msg): # from the fetch wrapper, still raw text
if msg["done"]:
self.reason = "stream closed"
return
st = self.state.setdefault(msg["url"],
{"buf": "", "data": "", "type": "", "id": None, "retry": None})
for ev in parse_sse(msg["chunk"], st):
self.write(msg["url"], ev["event"], ev["data"], ev["id"] or "")
def should_stop(self):
if self.reason != "running":
return True
if time.monotonic() > self.deadline:
self.reason = "deadline"
return self.reason != "running"
with InvisiblePlaywright(seed=42) as browser:
page = browser.new_page()
rec = StreamRecorder("events.jsonl", max_seconds=120)
page.expose_function("__sse", rec.on_message)
page.expose_function("__sseRaw", rec.on_raw)
page.add_init_script(EVENTSOURCE_HOOK)
page.add_init_script(FETCH_HOOK)
page.goto("https://example.com/stream", wait_until="domcontentloaded")
while not rec.should_stop():
page.wait_for_timeout(500)
print(rec.reason, dict(rec.kinds))wait_until="networkidle" is the wrong wait here and will time out like a broken page,
because an open stream keeps the network busy forever, which is one of the traps in
waiting for the page to load. Rows go out as
JSON Lines so a killed run keeps everything written up to that moment, the same reason
writing to JSON Lines suits any open-ended
capture. Read the counter first: a run that ends with a thousand keepalives and four
messages is telling you the subscription never took.
Server-sent events are simpler on the wire and harder in Playwright, and the gap is
entirely in the tooling. There is no frame event, so the response hook that looks right
hangs on a body that never completes, and the CDP event that carries the parsed message is
limited to Chromium and to pages that build a real EventSource. What survives all three
browsers is instrumentation: replace the constructor and wrap fetch in an init script that
runs before page code, bring each message back through an exposed function, and parse the
format the way the spec parses it, colon by colon. Then key your deduplication on the
event id, because a reconnect will replay and the browser will not tell you it happened.
Does page.on("response") work for server-sent events? It fires, which is the trap.
Playwright emits the response event when status and headers arrive, so the stream shows up
immediately and the capture looks correct. response.body() then waits for a body that
never completes.
Why is there no framereceived event for SSE? Because Playwright models SSE as an HTTP
response rather than as a message-carrying connection. WebSocket has its own class with
framesent and framereceived; for SSE, eventsource is only a value of
request.resource_type.
Can I read the messages over CDP? On Chromium, for a page that constructs a real
EventSource, yes: Network.eventSourceMessageReceived gives requestId, timestamp,
eventName, eventId and data. Not on Firefox or WebKit, and not when the page reads its
stream with fetch.
How do I catch a stream opened with fetch? Wrap window.fetch in an init script,
check for text/event-stream in the content type, and tee() the body so one branch is
yours and one goes back to the page. Drain your branch, since a tee buffers the difference.
Why does my parser lose the later lines of a multi-line message? Because each data
field appends its value and then a single LF, so consecutive data lines join with newlines
between them and only the final LF is removed at dispatch. Concatenating without that LF
collapses them into one line.
The stream stopped and nothing raised. What happened? Three candidates that look
alike from Python. The server may have answered 204, the documented way to say stop
reconnecting. The response may not have been 200 or not text/event-stream, which fails
the connection with no retry. Or the connection dropped and the browser is inside its
reconnection delay.
- The WHATWG HTML standard,
server-sent events,
for every parsing rule on this page: the comment line, the first-colon split, the single
stripped space, the data buffer and its trailing LF, the NULL rule on
id, the ASCII-digits rule onretry,Last-Event-ID, thereadyStatevalues, the 204 case and the rule that a failed connection is not retried. Retrieved 2026-08-28. - Playwright's Request class, for
resource_typeand theeventsourcevalue in its list. Retrieved 2026-08-28. - Playwright's Response class,
for
body()andfinished(), which is where the hang argument comes from. Retrieved 2026-08-28. - Playwright's BrowserContext class,
for
add_init_scriptrunning before the page's own scripts, forexpose_functionadding a binding on every frame that outlives page loads, and for the statement that CDP sessions are only supported on Chromium-based browsers. Retrieved 2026-08-28. - Playwright's WebSocket class,
for the
framesentandframereceivedevents this page contrasts against. Retrieved 2026-08-28. - Playwright issue 17199,
requesting
Response.bodyAsStream(), opened 2022-09-08, open and labelled P3-collecting-feedback. Retrieved 2026-08-28. - Chromium issue 40659493, "DevTools: XHR and fetch to text/event-stream resources don't show events". Retrieved 2026-08-28.
- The Chrome DevTools Protocol
Network domain,
for the five parameters of
Network.eventSourceMessageReceived. Retrieved 2026-08-28.
See also: scraping WebSocket streams for the protocol Playwright does expose at message level, capturing XHR and API responses for the requests that actually finish, writing to JSON Lines for the append-safe row file a long capture needs, and resuming an interrupted scrape for picking up from the last id you held.
Written while maintaining invisible_playwright, a Firefox patched at the C++ level driven by stock Playwright. The CDP event was the first answer tried here and it is the one every thread recommends, and it reported nothing on a page that read its stream with fetch: not a broken hook, just a hook waiting for an EventSource that was never constructed.
Documentation
Guides
-
Browser Identity
- navigator.webdriver is not the tell you think it is
- hardwareConcurrency, deviceMemory and storage quota
- Screen size and viewport tells in headless browsers
- Playwright headless vs headed: what detectors see
- Playwright User Agent: Why You Should Not Set It
- Client Hints and Sec-Fetch: headers that must agree
- Codec fingerprinting: canPlayType and MediaCapabilities
- Permissions API: the two answers that must agree
- CSS fingerprinting: what media queries reveal
- What privacy.resistFingerprinting actually does
- speechSynthesis.getVoices() returns an empty array
- Browser extensions are a fingerprint surface
- BFCache and pageshow.persisted under browser automation
- Service workers, storage partitioning and automation
- Web Workers: where page-level fingerprint patches fail
- fake-useragent is archived: what changes and what doesn't
- navigator.buildID and the stale build date tell
- navigator.maxTouchPoints and pointer consistency
- navigator.platform and oscpu on a spoofed OS
- navigator.vendor and productSub: the Firefox tells
- Accept-Language header vs navigator.languages
- window.devicePixelRatio: the pref that spoofs it
- Can you be fingerprinted in incognito mode?
- Is changing the user agent enough to avoid detection?
- Can a website tell you are running on a server?
- Can two devices share a browser fingerprint?
- Does clearing cookies stop fingerprint tracking?
- Color-gamut and HDR media queries as a fingerprint
- Battery API fingerprint: does Firefox expose it?
- Is navigator.connection a fingerprint in Firefox?
- Can the Gamepad API fingerprint or detect a bot?
- Do accelerometer and gyroscope APIs leak on desktop?
- prefers-reduced-motion and other OS-setting tells
- Does storage quota estimate reveal disk size?
- Can scrollbar width reveal my operating system?
-
Canvas, WebGL, Fonts and Audio
- Canvas fingerprint noise: why per-call randomising fails
- Firefox WebGL renderer strings: what ANGLE reports
- WebGL parameters: the numbers are the same on every GPU
- Your renderer string says NVIDIA. Your pixels say software.
- Why headless browsers render different fonts
- How to make Linux and macOS report real Windows fonts
- measureText and TextMetrics as a fingerprinting surface
- AudioContext fingerprinting, and why adding noise backfired
- Canvas and WebGL fingerprints, identical across OSes
- Emoji fingerprinting: why emoji look the same on any OS
- Detecting installed fonts in JavaScript by width
- WebGL shader precision as a fingerprint surface
- AudioContext sampleRate and latency as a fingerprint
- Is WebGPU a browser fingerprint?
-
Network, Proxy and WebRTC
- WebRTC leak with a proxy in Playwright and Selenium
- WebRTC ICE candidate spoofing: the fields that give it away
- Playwright proxy in Python: per-context, and what leaks
- Playwright proxy not working? SOCKS5 auth in Python
- Playwright timezone does not match the proxy IP
- JA3 and JA4: why a TLS fingerprint cannot be patched
- Playwright in Docker: it runs, and still gets blocked
- Web scraping keeps getting blocked with good proxies
- Python web scraping blocked? The TLS fingerprint reason
- SOCKS5 vs HTTP proxy: what each does in the browser
- WebRTC IPv6 leak: why a proxy does not stop it
- HTTP/2 fingerprint: the layer above the TLS handshake
- TLS fingerprint vs User-Agent: the contradiction
- WebRTC has no ICE candidates behind a proxy
- WebRTC IP that matches the proxy exit, by design
- How to check if a proxy leaks your real IP
- about:webrtc: read your real ICE candidates
- Offline timezone resolution from a proxy exit IP
- Residential vs datacenter vs mobile proxies explained
- Sticky vs rotating proxy sessions: which to use
- Does a proxy leak DNS? DoH and DNS leaks explained
- HTTP/3 and QUIC fingerprint: what a site sees
- What is ASN and IP reputation in bot detection?
- What does a mobile carrier IP look like to a site?
- IPv6 vs IPv4: which does your proxy expose?
- Geolocation API vs IP location: keep them consistent
- Does chaining two proxies help avoid detection?
-
The Automation Layer
- Function.prototype.toString and the [native code] check
- The ChromeDriver
cdc_variable, and why renaming it fails - Why an attached debugger makes automation detectable
- Execution context was destroyed, and when it means detection
- Human-like mouse movement: Bezier curves are the easy part
- Why a Playwright upgrade broke 97 of 133 tests overnight
- Playwright persistent profile: what it fixes and breaks
- Why humanized mouse movement can fail on hover()
- Why content_frame() returns None for a cross-origin iframe
- Orphaned Firefox processes on Windows: the killed-runner leak
- Firefox launches but Playwright can't drive it: packaging gap
- Why automating login is riskier than reusing a session
- Playwright new_page vs new_context: the viewport tell
- Playwright dialog and popup handling without a tell
- Playwright download files with Firefox and the tell
- Playwright connect_over_cdp does not work with Firefox
- Playwright mobile emulation on Firefox and isMobile
- Playwright isTrusted: are automated clicks real?
- Playwright set_input_files uploads and the tell
- Can websites detect Playwright? What is actually visible
- Does Playwright Set navigator.webdriver to True?
- Does Playwright Leave Traces a Website Can See?
- Does Playwright Change My Browser Fingerprint?
- Can I Use My Real Browser Profile With Playwright?
- Does Playwright Support Firefox Stealth?
- Is Playwright Firefox Harder to Detect Than Chromium?
- Does Playwright Get Detected on the First Request?
- Why Playwright's bundled Firefox is easy to detect
- ghost-cursor human mouse paths with Playwright
- Stock Playwright, patched Firefox: how they connect
- Intercept and mock network requests with page.route
- Record and replay HTTP traffic with HAR in Playwright
- Record a Playwright trace to debug a failed scrape
- Record a video of a Playwright browser session
- Save and reuse login with storage_state in Playwright
- Read and set cookies in a Playwright context
- Set geolocation and permissions per Playwright context
- Handle HTTP basic auth in Playwright (http_credentials)
- Isolate identities with a browser context per session
- Drag and drop elements in Playwright with drag_to
- When to use an HTTP client vs a real browser
- Migrating from requests + BeautifulSoup to a browser
-
AI Agents and Frameworks
- AI browser agents and stealth: what fits and what does not
- browser-use gets detected: what you can and cannot change
- crawl4ai stealth mode and custom browser engines
- Give a LangChain agent an invisible_playwright browser
- Feed invisible_playwright pages into a RAG index
- Computer-use agents and browser fingerprint detection
- Give an MCP browser server a stealth Firefox engine
- Give each AI agent a reproducible browser identity
- Run parallel browser agents with distinct fingerprints
- Why AI browser agents have their own timing signal
- Running an AI browser agent headless on a server
- Give a browser agent a persistent logged-in session
- smolagents: hand the agent an invisible_playwright tool
- Stagehand and stealth: why a Firefox engine won't drop in
- DOM-reading vs screenshot agents: which stealth helps
- Back a computer-use agent with a real browser engine
- AI agent retry loops trip rate limits, not fingerprints
-
Detectors, Explained
- What bot.sannysoft.com actually checks, row by row
- How CreepJS decides you are lying
- What BotD actually detects, and what it does not
- Why a FingerprintJS visitor ID changes
- reCAPTCHA v3 score: why a fresh browser scores badly
- BrowserLeaks canvas and WebGL hash, explained
- What BrowserLeaks actually tests, surface by surface
- Browser trust scores explained: what the number means
- How do websites detect bots?
- What is a browser fingerprint?
- What data does a website collect about your browser?
- Does a VPN stop browser fingerprinting?
- Do websites know you are using a script?
- How accurate is browser fingerprinting?
- Can a website detect a virtual machine?
- Can websites detect a datacenter or proxy IP?
- getClientRects fingerprinting: subpixel geometry as ID
- Notification.permission as a bot-detection signal
- speechSynthesis voices as a cross-platform fingerprint
- Can a website detect typing by keystroke timing?
- Can a website detect Clipboard API access?
- What are mouse-dynamics behavioural biometrics?
-
Testing and Troubleshooting
- How to test bot detection without a false pass
- Playwright detected as a bot: the checklist to fix it
- Firefox preferences that silently do nothing
- Slow browser launch: a per-request timeout is not a budget
- Playwright screenshot returns noise: readback fix
- Canvas fingerprint changes every run: use a seed
- Playwright TargetClosedError: the causes and the fixes
- Why am I blocked with a clean fingerprint?
- Why Does My Playwright Script Get Blocked?
- Is Playwright headless detectable? What sites check
- Can You Run Playwright Without Being Detected?
- Why Playwright Works Locally but Fails in the Cloud
- Does Playwright Trigger reCAPTCHA More Often?
-
Scraping with Playwright
- How to scrape without getting blocked
- How to scrape a site that blocks headless browsers
- How to scrape infinite scroll pages with Playwright
- How to rotate proxies when scraping with Playwright
- How to scrape data behind a login with Playwright
- How to run Playwright in Docker without getting detected
- How to use invisible_playwright in Docker
- Playwright bot detection: how to avoid it in Python
- How to scrape paginated pages with Playwright
- How to download files with Playwright
- How to upload files with Playwright, and verify it landed
- How to handle cookie consent banners in Playwright
- How to handle popups and modals in Playwright
- How to take full-page screenshots with Playwright
- How to generate a PDF with Playwright and Firefox
- How to wait for content to load in Playwright
- How to retry failed requests when scraping Playwright
- How to scrape pages in parallel with Playwright
- How to rate limit your own Playwright scraper
- How to scrape HTML tables with Playwright
- How to scrape iframe content with Playwright
- How to scrape shadow DOM content with Playwright
- How to capture XHR and API responses in Playwright
- How to scrape geotargeted content with Playwright
- How to scrape real estate listings with Playwright
- How to scrape job postings with Playwright
- How to scrape e-commerce product pages with Playwright
- How to track product prices with Playwright
- How to scrape hotel room prices with Playwright
- How to scrape flight prices with Playwright
- How to scrape classifieds listings with Playwright
- How to scrape vacation rental listings with Playwright
- How to scrape car listings with Playwright
- How to scrape apartment rentals with Playwright
- How to track product stock and restocks with Playwright
- How to scrape location-based store prices with Playwright
- How to scrape flexible-date fare calendars with Playwright
- How to scrape product reviews with Playwright
- How to scrape reviews and ratings with Playwright
- How to scrape news article text with Playwright
- How to scrape business directory listings with Playwright
- How to scrape event and ticket listings with Playwright
- How to scrape restaurant menu data with Playwright
- How to scrape stock and financial data with Playwright
- How to scrape social media profiles with Playwright
- How to scrape forum and community threads with Playwright
- How to scrape image galleries with Playwright
- How to scrape video listings and metadata with Playwright
- How to scrape map-based local results with Playwright
- How to scrape sports scores and stats with Playwright
- How to scrape cryptocurrency prices with Playwright
- How to scrape deals and coupon codes with Playwright
- How to scrape to CSV with Playwright
- How to scrape to JSON Lines with Playwright
- How to scrape into a SQLite database with Playwright
- How to export scraped data to Excel with Playwright
- How to extract JSON-LD structured data with Playwright
- How to extract Open Graph and meta tags with Playwright
- How to extract links and build a crawl frontier in Playwright
- How to scrape RSS and Atom feeds with Playwright
- How to download images in bulk with Playwright
- How to extract clean article text with Playwright
- How to scrape a sitemap.xml with Playwright
- How to scrape into a pandas DataFrame with Playwright
- How to clean scraped prices and dates with Playwright
- Scrape search results by driving a form in Playwright
- Scrape a map-based search with Playwright
- Scrape autocomplete and typeahead inputs with Playwright
- Scrape date-picker calendars with Playwright
- Crawl list pages to detail pages with Playwright
- Scrape lazy-loaded images with Playwright
- Extract data from canvas charts with Playwright
- Scrape a multi-step wizard flow with Playwright
- How to resume an interrupted scrape with Playwright
- Incremental scraping: only new items since last run
- Handle 403 and 429 backoff mid-scrape in Playwright
- Scrape load-more button pages with Playwright
- Scrape nested pagination with Playwright
- Scrape an SPA that changes URL via history API
- Use BeautifulSoup with invisible_playwright
- Run stealth Playwright tests with pytest fixtures
- Run invisible_playwright concurrently with asyncio
- Run invisible_playwright in GitHub Actions CI
- Can you run invisible_playwright serverless?
- Run invisible_playwright in Celery task workers
- Schedule invisible_playwright scrapes with cron
- Run invisible_playwright headful on a server with Xvfb
- Use invisible_playwright in an Airflow DAG
- Combine invisible_playwright with httpx for speed
- Wrap invisible_playwright in a FastAPI service
- Run invisible_playwright in a Jupyter notebook
- Block images to speed up scraping (and when not to)
- Wait for a specific API response in Playwright
- How to scrape course catalogs with Playwright
- How to scrape store locator pages with Playwright
- How to scrape stock levels with Playwright
- How to scrape accordion and tab content with Playwright
- How to scrape size charts with Playwright
- How to scrape delivery slots with Playwright
- How to scrape appointment availability with Playwright
- How to scrape auction listings with Playwright
- How to scrape public transport timetables with Playwright
- How to scrape GraphQL endpoints with Playwright
- How to scrape virtual scrolling tables with Playwright
- How to scrape shipping rates with Playwright
- How to scrape cursor-based pagination with Playwright
- How to scrape multi-select facet filters with Playwright
- How to scrape currency exchange rates with Playwright
- How to scrape WebSocket streams with Playwright
- How to scrape book metadata with Playwright
- How to scrape professional directories with Playwright
- How to scrape range slider filters with Playwright
- How to scrape currency and locale switchers with Playwright
- How to scrape software changelogs and release notes with Playwright
- How to scrape breadcrumb hierarchies with Playwright
- How to scrape microdata and RDFa markup with Playwright
- How to scrape server-sent events with Playwright
- How to scrape open data portals with Playwright
- How to scrape infinite carousels with Playwright
- How to scrape printer-friendly pages with Playwright
- How to handle A/B test variants when scraping with Playwright
- How to scrape recipe data with Playwright
- How to scrape vehicle recall notices with Playwright
- How to scrape public tender notices with Playwright
- How to scrape nutrition labels with Playwright
- How to scrape podcast episode listings with Playwright
- How to scrape weather station data with Playwright
- How to scrape newsletter archives with Playwright
- How to scrape wine and spirits catalogs with Playwright
- How to scrape insurance quotes with Playwright
- How to scrape fitness class schedules with Playwright
- How to scrape flight seat maps with Playwright
- How to scrape concert and tour dates with Playwright
- How to scrape museum and gallery exhibition dates with Playwright
- How to scrape warranty terms with Playwright
- How to scrape sortable data tables with Playwright
- How to scrape salary and pay scale data with Playwright
- How to scrape live sports scores with Playwright
- How to scrape video game prices with Playwright
- How to scrape domain WHOIS records with Playwright
- How to scrape podcast transcripts with Playwright
- How to scrape patent listings with Playwright
- How to scrape clinical trial listings with Playwright
Comparisons
- Playwright stealth in Python: three levels that work
- Firefox or Chromium for anti-detect automation
- Chromium is not Chrome, and detectors know the difference
- Playwright stealth vs Camoufox: two patched Firefoxes
- Playwright stealth vs Patchright: driver vs engine
- Playwright stealth vs undetected-chromedriver and nodriver
- playwright-stealth vs a patched engine: page vs browser
- puppeteer-extra-plugin-stealth: unmaintained since 2023
- selenium-stealth hasn't been updated since November 2020
- pyppeteer's own maintainer says to switch to Playwright
- invisible_playwright vs rebrowser-patches: the same CDP fix
- invisible_playwright vs fingerprint-suite: injection vs engine
- invisible_playwright vs playwright-with-fingerprints
- invisible_playwright vs Scrapling
- invisible_playwright vs Ulixee Hero
- invisible_playwright vs SeleniumBase UC Mode
- Splash is unmaintained, and it was never a real browser
- invisible_playwright vs DrissionPage
- WebDriver BiDi vs CDP: does the new protocol hide you
- invisible_playwright vs hrequests
- zendriver vs invisible_playwright: Chrome CDP vs Firefox
- botasaurus vs invisible_playwright: framework vs library
- curl_cffi vs invisible_playwright: TLS client vs browser
- pydoll vs invisible_playwright: CDP without a driver
- selenium-driverless vs invisible_playwright stealth
- puppeteer-real-browser vs invisible_playwright
- Migrating from Selenium to Playwright for stealth
- Migrating from Puppeteer to Playwright for stealth
- undetected-chromedriver vs a patched Firefox browser
- scrapy-playwright vs a patched Firefox for stealth
- playwright-extra stealth plugins vs a patched browser
- tls-client vs a real browser: when TLS is enough
- Anti-detect browser or Playwright stealth: which you need
- undetected-playwright vs a patched Firefox binary
Integrations
- Using invisible_playwright with CodeceptJS
- Using invisible_playwright with Crawlee for Python
- Using invisible_playwright with Crawlee for JavaScript
- Using invisible_playwright with scrapy-playwright
- Using invisible_playwright with Robot Framework Browser
- Cypress, WebdriverIO, TestCafe and Nightwatch integration
- Using invisible_playwright with Microsoft's Playwright MCP
- Using the engine from Go, Java, C#, Ruby and Rust
docs/ source folder