Repository navigation
Releases: 2scraper/medium-scraper
Release list
v0.1.0 — family-architecture rewrite
If you used the old scripts, every command you have will stop working.
medium_playwright.py,medium_selenium.pyandmedium_pyppeteer.pyare
gone, and so are--mode publicationand--target. The entry points are
nowplaywright_scraper.py,selenium_scraper.pyand
puppeteer_scraper.py, and the target is a full--url. A publication
home page is now refused rather than scraped — see below for the
measurement that decided that.
First release. A complete rewrite onto the architecture the rest of the
2scraper family uses: four modes, four
engines, one row shape, run metadata, an exit-code contract, and 726 offline
checks. Every number in the README was measured against the live site on
2026-09-16.
Four modes
| mode | URL | one fetch returns |
|---|---|---|
--mode tag |
/tag/{slug} |
53 stories |
--mode archive |
/tag/{slug}/archive/{yyyy}/{mm}/{dd} |
up to 128 stories |
--mode author |
/@{username} |
10 stories |
--mode post |
/p/{id} |
1 story, with its full text |
Only the day archive has per-page addresses, so only it walks with --pages
and accepts --concurrency. ?page=2 on a tag feed is ignored rather than
failing, which is how a silent single-page run happens.
The one thing to know
Cloudflare hard-blocks the HeadlessChrome token in the User-Agent, not
headless browsing — 11 of 11 refused with it, 0 of 11 without. These
engines build the UA from the browser's own version and never send it, which
is why they get data from an address curl gets nothing from.
Two payloads, and why columns differ between modes
Medium serves the same catalogue through two renderers. The modern site ships
__APOLLO_STATE__; a tag's day archive ships the legacy
window["obvInit"] with 84 fields per story. That is why
reading_time_min, word_count and language are null on a tag-feed row
and populated on an archive row — it is the site, not a failed read, and
data_source on every row says which view built it.
What the paid 2Captcha products actually buy
Nothing is required to get data: every figure above came from a datacentre
address with no key and no proxy.
- Scraper API — all four modes, HTTP 200, rows identical to a local
browser's, $0.0005 a request, no browser installed anywhere. One request
returns the whole 2.3 MB archive-day payload. - Scraping Browser — 254 rows over two archive days, and 55 rows from
an author page against a local browser's 10: that feed extends by
scrolling, and the scroll is refused from an ordinary address and not from
its exit. - Captcha solving buys nothing here. Medium's refusal is Cloudflare's
managed challenge, which publishes no sitekey, so no solve is ever
attempted and nothing is ever charged.
Refused, with the reason
A publication home page carries 167 (or 51) post ids and no post data —
no title, no author, no date. Rows built from it would hold a sku and 26
nulls while the run reported success, so the URL is refused and the two
alternatives that work are named. towardsdatascience.com is refused with
the reason it actually left (WordPress).
Full detail in CHANGELOG.md and
TROUBLESHOOTING.md.