Skip to content

Releases: 2scraper/wellfound-scraper

v0.1.1 — audit fixes

Choose a tag to compare

@jehrr jehrr released this 17 Sep 18:14

Correction to v0.1.0's documentation. v0.1.0 described the captcha path as wired in and bounded, and did not say what buying a solve actually achieves here. Measured now: on this site, from a datacenter address, a paid Turnstile token is refused — the challenge is captured and solved correctly and Cloudflare does not accept it (3 tokens, 3 refusals). If you read v0.1.0's README and bought solving to get past the block, that is not what will get you in. A residential exit is, and it costs less.

An audit against this project family's own checklist, section by section, executing each claim rather than re-reading it. Three defects, all found by running a path that had never been run, and none of them visible to the 356 green offline checks v0.1.0 shipped with.

Fixed

  • SOLVES_PER_PAGE was not enforced, and it is a money limit. Every engine calls the solver twice per attempt — once before the page is classified, once after — and only the second was counted. From a datacenter address, where a real Cloudflare challenge renders on every fetch: one page bought three Turnstile solves. Both call sites are now counted; the same run buys one.
  • A dead proxy was reported as a plain "gave up loading". The reason was computed and logged on rotation with a pool, then dropped without one — the single---proxy case, which is the common one. --proxy http://127.0.0.1:9 said nothing about the proxy. All three engines now name it.
  • Two inherited claims that were false about this site. scraper_api_client.py printed a sibling repo's measurement as this repo's; replaced with what was measured here (the Scraper API's own exits get upstream 403 and 11,825 bytes; the same task through a Scraping Browser session returns 200, 610,897 bytes and 37 job rows). captcha_solver.py's docstring described another site's defences and called itself load-bearing here.
  • The README said angel.co's old routes no longer exist. They do — angel.co/role/r/software-engineer returns the same page. The host is still refused so one job never appears under two spellings, but the stated reason is now the true one.

Verified, by running it

The fingerprint path applies its user agent, locale (en-US, not en-{country}) and timezone to the real browser · a failing API call does not leak the key · credentials never reach the browser's argv · the flag set covers all 28 family-contract flags · pip check passes per engine venv · the Scraping Browser path serves this site · --dump-html writes on success · exit 3 and exit 4 are distinct · no column is null on every row of every run · all three engines return 52 identical rows after the edits.

373 offline checks, CI green on Python 3.9 and 3.12, all three engine-smoke jobs and the Docker build.

v0.1.0 — first release

Choose a tag to compare

@jehrr jehrr released this 17 Sep 16:22

First release of wellfound-scraper. Scrapes wellfound.com (formerly AngelList Talent): role landing pages, the /jobs discovery feed, and individual job pages.

Every number below was measured on 2026-09-17.

What you need

A residential exit, and nothing else. Wellfound refuses a datacenter address on every route, robots.txt included — HTTP 403 to plain HTTP, to headless Chromium and to headful Chromium alike. From a 2Captcha residential US exit the same machine got HTTP 200 on /, /jobs, every /role/… shape and every job page, three times out of three.

What it does

  • Three modes. role (a /role/… landing page — full descriptions, company badges, size and tagline, salary and equity; paginates), jobs (the discovery feed, one page), job (one job from its schema.org JobPosting).
  • Four fetchers. Playwright (primary, with --concurrency), Selenium, Puppeteer, and a Scraper API client. All three browser engines were run live against the same listing and returned byte-identical rows — the same 52 skus, 38 columns, and the same position on every row.
  • 38-column JobPosting schema, JSON and CSV, with a .meta.json sidecar per run.
  • 355 offline checks, fixtures cut from real captures and verified to parse identically to their untrimmed originals.

What the site turned out to be

  • Gating is per ROUTE. /role/…, /jobs and /jobs/{id}-{slug} are served to a residential exit; /company/{slug} is refused to one, to plain HTTP and to a real browser alike, 3 of 3 each. There is no --mode company, because there is nothing this repo could honestly offer for it.
  • The obvious challenge marker is inverted here. challenges.cloudflare.com appears once on every page Wellfound serves — the site loads Turnstile as part of its own application, with its sitekey in its own page config — and zero times on either refusal. cf-turnstile appears nowhere at all. The marker set uses _cf_chl_opt and its siblings instead, and the suite fails the build if the inverted one comes back.
  • Walking past the last page does not fail — it repeats. ?page=48 of a 47-page listing answers HTTP 200 with page 1's rows again. Runs plan against the site's own pageCount and read back the page number the server states in its Apollo cache key.
  • A "page" is twenty COMPANIES, not twenty jobs. perPage: 20, pageCount: 47, totalStartupCount: 923, totalJobCount: 1881 — and 32 to 45 job rows per page.
  • /jobs has no addressable pages. ?page=2 returns the identical 46 job ids, so that mode reports single_page_listing.
  • Two structured sources, different applications. A landing page carries zero JSON-LD and a Next.js Apollo cache; a job page carries no __NEXT_DATA__ and one JobPosting. Where both publish a salary they agree.
  • Pay is one rendered string with nine measured shapes: the bullet splits salary from equity and either side may be absent, L is the Indian lakh (1e5), and No equity is a statement rather than a missing value.