Repository navigation
Releases: 2scraper/potterybarn-scraper
Releases · 2scraper/potterybarn-scraper
Release list
v1.0.0
First release. Measurements were made on 2026-09-24 and re-made on
2026-10-01; each one below says which.
Transports
api_scraper.py: category listings through the site's own listing
API (Constructor.io's browse API onac.cnstrc.com, with the
storefront's public key). It needs no proxy, browser or key, and answers
a residential address in Moscow and a GitHub Actions runner alike. The
whole sofa category: 434 of 434 (09-24) and 433 of 433 (10-01). The API
host is a third party that potterybarn.com's robots.txt does not govern;
the README says so in its own section.- Unstable page order is recovered, not hidden. The API sometimes
repeats a product on a later page and so skips another: 415 distinct of
434 (09-24), 403 of 433 (10-01), and 1 and 4 repeats on two daily canary
runs. With--all-pagesa second, price-sorted sweep recovers the rest
(all 30 on 10-01). A listing still short is reported partial, never
complete. A--pages Nrun reportsrepeated_resultsand is not compared
with the category total. - Rate limit, measured to zero (10-01):
x-ratelimit-remainingreached 0
after 240 uncached calls, and 194 further calls at zero all answered 200
with full results. The budget is per client address. The scraper slows to
one page per 5 seconds below 10 remaining, and does not stop. http_scraper.py: product pages over system curl, one row per SKU.
On fresh US exits: curl 29 of 31 (09-24) and 4 of 4 (10-01);
python-requests0 of 18 and 2 of 4; a local Chromium 0 of 15 and 0 of 4.
The proxy reaches curl on stdin, never argv. Redirects are followed only
inside www.potterybarn.com, and a hop to potterybarn.co.uk is a refusal
that rotates to a freshly minted session.- Browser engines:
playwright_scraper.py,puppeteer_scraper.py,
selenium_scraper.py, sharing onebrowser_bridge.py, all three run live.
Listings are identical toapi_scraper.pyfield for field through all
three (72 of 72). Product pages over the Scraping Browser API: refused with
HTTP 403 on 09-24, served on 14 of 14 sessions on 10-01 and identical to
curl field for field (4 of 4 and 1,380 of 1,380 SKUs).Captcha.setAutoSolve
was accepted every time, with zero Captcha events. One connection per run,
closed infinally, never preceded byGET /json/version. - Navigation never waits for the page to render. Playwright returns on
commitand reads the body;domcontentloadedhad timed out in 3 of 20
instrumented runs, and a timeout can hold the Scraping Browser profile. On
the new wait: 0 of 10. Selenium usespage_load_strategy = "eager". - Browsers read the response body, not the DOM. The storefront deletes
its state<script>on hydration, so a served product page's DOM reads
as empty. The engines read the navigation's body first and fall back to
the livewindow.__INITIAL_STATE__object, recordingstate_source. --browser-pathfor every engine: pyppeteer's downloaded Chromium is an
x86_64 build that did not start under Rosetta. Measured engine limits are
reported as exit 5 with the reason: Selenium cannot authenticate a proxy
(Chrome then shows an empty 39-byte document) and cannot use an
authenticated CDP endpoint, and pyppeteer cannot authenticate a proxy on
Chrome 154 (Network.setRequestInterceptionis gone).catalog_walk.py: the category and product sitemaps (795 and 16,966
URLs on 09-24), gzip-aware, robots-filtered, no proxy needed. The files
refuse the defaultpython-requestsuser agent, so the browser header set
is sent.
Parsing
- Product pages are read from
window.__INITIAL_STATE__.product.productDetails
with a string-aware scanner. There is noProductJSON-LD on this site. - Guided products' SKUs are decoded from
subsetsCompressedValue
(base64 Brotli). On a guided product, which covers every sofa measured,
the servedsubsetslist is empty. Each product's SKU count is checked
against the page's ownskuCount(rows − NLA). Product(listing) andSkuRow(product page) share the family's
ten-column prefix, and both carryproduct_idfor the join. A listing
result can be a sub-group of a product, a slice by finish (15 of 144
results):skuis the result's id,product_idthe URL path's slug, and
sub_group_idthe slice.--from-listingfetches each product once.- Listing price ranges equal the product page's own aggregate range.
- Currency is read from what a page states. The listing API states none, so
a listing row's currency comes from the listing page, or from the currency
recorded for that exact key. - Exit codes follow the family contract everywhere: a refused URL (another
host, a robots-disallowed path, the wrong page kind) is exit 2, never the
crash code 1, and a top-level/shop/<department>/is a hub (exit 4) even
when its page cannot be read, instead of being guessed into a broad API
group (furniture: 8,278 results). - The branded 403 is recognised by status and its own text. An empty document
is a transport failure, not a refusal. The Scraping Browser's auto-solve
extension scripts are stripped before classification, because on its own
403 page they namecf-turnstileand would otherwise route a refusal to
the paid solver. The page's reCAPTCHA sitekey (v2 checkbox, by anchor
probe) is ignored unless a widget is actually rendered.
Completeness is never claimed for a subset
- A
--max-productscap stops withmax_products_reachedand is a partial
run (exit 6), nevercomplete. The sidecar's scope records the cap and
is_full_catalog, and the recovery sweep does not run after a deliberate cap. - A product whose SKU count disagrees with the page's own
skuCountmakes
the run partial (sku_count_mismatch;--accept-count-mismatchto
override). The counting rule is "displayable, or not NLA". It matched on
40 of 40 product pages (20 fitted, 20 held out), where the first rule,
"rows − NLA", missed 3 of 20. diff_runs.pyfails closed: a missing or unreadable sidecar, or two runs
differing in mode, schema, source, listing group or currency, is refused
without--force.potterybarn-browseron a base install (no[playwright]extra) says what
to install and exits 2, instead of a traceback on--help.
Safety
tools/scan_secrets.pywith--staged,--worktree,--rangeand
--history, run by versionedpre-commitandpre-pushhooks and by CI.
--rangesplits the pre-push hook's<sha> --not --remotesinto separate
arguments and fails closed on any git error. Without that, a new branch's
first push is never scanned.- robots.txt is embedded as well as committed (
robots.snapshot.txt), so an
installed wheel cannot lose it. The matcher fails closed on an empty rule
set. - CI actions on their Node 24 majors, read-only workflow permissions, and
Dependabot for pip, actions and the Docker base. The image runs as an
unprivileged user.
Not included, and why
- No Scraper API transport. The 2Captcha Scraper API answered HTTP 200
for two different product URLs with the same prerendered generic page and
no product in it (3 of 3 calls, 09-24). - No rating or review columns. Zero rating keys appeared in 3 product
pages and 24 listing results.