Releases: 2scraper/wellfound-scraper
Release list
v0.1.1 — audit fixes
Correction to v0.1.0's documentation. v0.1.0 described the captcha path as wired in and bounded, and did not say what buying a solve actually achieves here. Measured now: on this site, from a datacenter address, a paid Turnstile token is refused — the challenge is captured and solved correctly and Cloudflare does not accept it (3 tokens, 3 refusals). If you read v0.1.0's README and bought solving to get past the block, that is not what will get you in. A residential exit is, and it costs less.
An audit against this project family's own checklist, section by section, executing each claim rather than re-reading it. Three defects, all found by running a path that had never been run, and none of them visible to the 356 green offline checks v0.1.0 shipped with.
Fixed
SOLVES_PER_PAGEwas not enforced, and it is a money limit. Every engine calls the solver twice per attempt — once before the page is classified, once after — and only the second was counted. From a datacenter address, where a real Cloudflare challenge renders on every fetch: one page bought three Turnstile solves. Both call sites are now counted; the same run buys one.- A dead proxy was reported as a plain "gave up loading". The reason was computed and logged on rotation with a pool, then dropped without one — the single-
--proxycase, which is the common one.--proxy http://127.0.0.1:9said nothing about the proxy. All three engines now name it. - Two inherited claims that were false about this site.
scraper_api_client.pyprinted a sibling repo's measurement as this repo's; replaced with what was measured here (the Scraper API's own exits get upstream 403 and 11,825 bytes; the same task through a Scraping Browser session returns 200, 610,897 bytes and 37 job rows).captcha_solver.py's docstring described another site's defences and called itself load-bearing here. - The README said
angel.co's old routes no longer exist. They do —angel.co/role/r/software-engineerreturns the same page. The host is still refused so one job never appears under two spellings, but the stated reason is now the true one.
Verified, by running it
The fingerprint path applies its user agent, locale (en-US, not en-{country}) and timezone to the real browser · a failing API call does not leak the key · credentials never reach the browser's argv · the flag set covers all 28 family-contract flags · pip check passes per engine venv · the Scraping Browser path serves this site · --dump-html writes on success · exit 3 and exit 4 are distinct · no column is null on every row of every run · all three engines return 52 identical rows after the edits.
373 offline checks, CI green on Python 3.9 and 3.12, all three engine-smoke jobs and the Docker build.
v0.1.0 — first release
First release of wellfound-scraper. Scrapes wellfound.com (formerly AngelList Talent): role landing pages, the /jobs discovery feed, and individual job pages.
Every number below was measured on 2026-09-17.
What you need
A residential exit, and nothing else. Wellfound refuses a datacenter address on every route, robots.txt included — HTTP 403 to plain HTTP, to headless Chromium and to headful Chromium alike. From a 2Captcha residential US exit the same machine got HTTP 200 on /, /jobs, every /role/… shape and every job page, three times out of three.
What it does
- Three modes.
role(a/role/…landing page — full descriptions, company badges, size and tagline, salary and equity; paginates),jobs(the discovery feed, one page),job(one job from its schema.orgJobPosting). - Four fetchers. Playwright (primary, with
--concurrency), Selenium, Puppeteer, and a Scraper API client. All three browser engines were run live against the same listing and returned byte-identical rows — the same 52 skus, 38 columns, and the samepositionon every row. - 38-column
JobPostingschema, JSON and CSV, with a.meta.jsonsidecar per run. - 355 offline checks, fixtures cut from real captures and verified to parse identically to their untrimmed originals.
What the site turned out to be
- Gating is per ROUTE.
/role/…,/jobsand/jobs/{id}-{slug}are served to a residential exit;/company/{slug}is refused to one, to plain HTTP and to a real browser alike, 3 of 3 each. There is no--mode company, because there is nothing this repo could honestly offer for it. - The obvious challenge marker is inverted here.
challenges.cloudflare.comappears once on every page Wellfound serves — the site loads Turnstile as part of its own application, with its sitekey in its own page config — and zero times on either refusal.cf-turnstileappears nowhere at all. The marker set uses_cf_chl_optand its siblings instead, and the suite fails the build if the inverted one comes back. - Walking past the last page does not fail — it repeats.
?page=48of a 47-page listing answers HTTP 200 with page 1's rows again. Runs plan against the site's ownpageCountand read back the page number the server states in its Apollo cache key. - A "page" is twenty COMPANIES, not twenty jobs.
perPage: 20,pageCount: 47,totalStartupCount: 923,totalJobCount: 1881— and 32 to 45 job rows per page. /jobshas no addressable pages.?page=2returns the identical 46 job ids, so that mode reportssingle_page_listing.- Two structured sources, different applications. A landing page carries zero JSON-LD and a Next.js Apollo cache; a job page carries no
__NEXT_DATA__and oneJobPosting. Where both publish a salary they agree. - Pay is one rendered string with nine measured shapes: the bullet splits salary from equity and either side may be absent,
Lis the Indian lakh (1e5), andNo equityis a statement rather than a missing value.