Skip to content

v0.1.0 — mercor.com, ungated

Latest

Choose a tag to compare

@jehrr jehrr released this 18 Sep 10:36
· 9 commits to main since this release

First release. Scrapes mercor.com with Playwright, Selenium, Puppeteer, or the 2Captcha Scraping Browser API over CDP.

The headline is that you need none of it. Every route this scraper reads was served in full to a bare datacentre address with no key, no proxy and no account — measured 2026-09-18 from a Hetzner VPS (AS24940, Helsinki), to curl, to python-requests and to an empty User-Agent. Eighteen candidate bot-challenge markers were counted across seven captures and every one was zero. python3 playwright_scraper.py is a working command.

The canary is therefore ungated: it runs a real scrape daily from a plain GitHub runner and is expected green, which is what keeps that claim honest rather than merely asserted.

Three modes, because the site publishes three different things

Mode Route What it gives you
listings (default) work.mercor.com/explore 390 contract roles in ONE response — rate ranges, the site's own domain category, work/residence eligibility, slot counters
job work.mercor.com/jobs/{id}/{slug} One listing, adding the company website, description, spot counts and a currency the index does not publish. The paginated mode: a page is one job, enumerated from sitemap ∪ index
careers www.mercor.com/careers 108 of Mercor's own salaried roles via Ashby — a separate population with its own id space

What the site turned out to be

  • The obvious structured source is the wrong one. /explore ships an ItemList JSON-LD of 326 URLs beside a React Query cache of 390 full records; the 64 it omits are all 55 evergreen listings plus 9 standard. Anchoring on the JSON-LD loses 16% of the catalogue in silence.
  • No single route enumerates the catalogue. sitemap 447, index 390, shared 375, union 462 — neither contains the other, and the sitemap-only listings checked were all live.
  • /explore has no second page. ?page=, ?limit=, ?offset= and ?domain= each returned a byte-identical payload.
  • Currency is read, never defaulted: null on every index row, USD on 40 of 40 sampled detail pages, USD or GBP on careers.

Captchas: configured, never encountered

work.mercor.com runs invisible reCAPTCHA Enterprise on every page including its 404s; www.mercor.com and /careers run none at all. It renders no challenge frame — it scores the session in the background — so it never blocks and is never paid for. Its sitekey is in no served HTML and no eagerly-loaded bundle, so a static grep finds nothing; only a real browser reveals it.

This repo does implement enterprise reCAPTCHA (RecaptchaV2EnterpriseTaskProxyless), ordinary v2/v3, and Cloudflare Turnstile (TurnstileTaskProxyless). It has simply never needed them here.

Verification

  • All three engines run live across all three modes, exit 0, identical rows.
  • 422 offline checks, green with no engine library installed and in each single-engine venv.
  • Fixtures cut from real captures, each verified to parse identically to its untrimmed original.