Skip to content

Releases: 2scraper/medium-scraper

v0.1.0 — family-architecture rewrite

Choose a tag to compare

@jehrr jehrr released this 16 Sep 14:44
84efe5c

If you used the old scripts, every command you have will stop working.
medium_playwright.py, medium_selenium.py and medium_pyppeteer.py are
gone, and so are --mode publication and --target. The entry points are
now playwright_scraper.py, selenium_scraper.py and
puppeteer_scraper.py, and the target is a full --url. A publication
home page is now refused rather than scraped — see below for the
measurement that decided that.

First release. A complete rewrite onto the architecture the rest of the
2scraper family uses: four modes, four
engines, one row shape, run metadata, an exit-code contract, and 726 offline
checks. Every number in the README was measured against the live site on
2026-09-16.

Four modes

mode URL one fetch returns
--mode tag /tag/{slug} 53 stories
--mode archive /tag/{slug}/archive/{yyyy}/{mm}/{dd} up to 128 stories
--mode author /@{username} 10 stories
--mode post /p/{id} 1 story, with its full text

Only the day archive has per-page addresses, so only it walks with --pages
and accepts --concurrency. ?page=2 on a tag feed is ignored rather than
failing, which is how a silent single-page run happens.

The one thing to know

Cloudflare hard-blocks the HeadlessChrome token in the User-Agent, not
headless browsing
— 11 of 11 refused with it, 0 of 11 without. These
engines build the UA from the browser's own version and never send it, which
is why they get data from an address curl gets nothing from.

Two payloads, and why columns differ between modes

Medium serves the same catalogue through two renderers. The modern site ships
__APOLLO_STATE__; a tag's day archive ships the legacy
window["obvInit"] with 84 fields per story. That is why
reading_time_min, word_count and language are null on a tag-feed row
and populated on an archive row — it is the site, not a failed read, and
data_source on every row says which view built it.

What the paid 2Captcha products actually buy

Nothing is required to get data: every figure above came from a datacentre
address with no key and no proxy.

  • Scraper API — all four modes, HTTP 200, rows identical to a local
    browser's, $0.0005 a request, no browser installed anywhere. One request
    returns the whole 2.3 MB archive-day payload.
  • Scraping Browser — 254 rows over two archive days, and 55 rows from
    an author page against a local browser's 10
    : that feed extends by
    scrolling, and the scroll is refused from an ordinary address and not from
    its exit.
  • Captcha solving buys nothing here. Medium's refusal is Cloudflare's
    managed challenge, which publishes no sitekey, so no solve is ever
    attempted and nothing is ever charged.

Refused, with the reason

A publication home page carries 167 (or 51) post ids and no post data —
no title, no author, no date. Rows built from it would hold a sku and 26
nulls while the run reported success, so the URL is refused and the two
alternatives that work are named. towardsdatascience.com is refused with
the reason it actually left (WordPress).

Full detail in CHANGELOG.md and
TROUBLESHOOTING.md.