Repository navigation
v0.2.0 — audit pass; sidecar parity across the three engines
An audit pass over every section of this family's notes, plus a live re-run of
all three engines and both modes. Two real defects, both engine divergences
that no single-engine run could show.
Changes behaviour for an existing consumer
The Selenium and pyppeteer engines were writing a different sidecar from
the Playwright one. They carried three keys belonging to a sibling repo —
scroll,result_header,pages_still_growing, of whichscrollis
meaningless on a site that serves its whole page at once — and they were
missing the cap arithmetic.On this site that arithmetic is what keeps
status: completehonest:
Rakuten served 6,750 results of a query that matched 7,460,738, so a
sidecar saying only "complete" is lying by omission. All three engines now
write the same 20 keys, verified by a live run of each, and the suite
compares the key sets so it cannot drift again.
If you consume <out>.meta.json from the Selenium or pyppeteer engine, the
keys have changed — and now include total_results, reachable_max,
capped_by_site, page_size and pages_beyond_cap.
The other defect: a guard that was only a warning
Two of three engines logged --fingerprint is ignored with --cdp-endpoint
while leaving the flag True. They were correct only because each remote
branch happens to return before the fingerprint is applied — a claim
enforced by where a return sits rather than by the stated gate. The day
someone moves the fingerprint into shared setup, those two engines would
silently start stacking a second identity onto a browser that already has
one, which on this site is the one thing measured to get a client
refused.
All three now force the flag off, and the suite asserts the behaviour —
every call that fetches or applies a fingerprint must sit behind the flag,
followed through the _apply_fingerprint indirection. The first version of
that check reported two engines as ungated, which was the check being wrong
rather than the code.
Also fixed
- A dead
scrollfield on the per-page outcome, and a scroll measurement
ported verbatim from a sibling. Replaced with this site's own: the same URL
fetched by plain HTTP with no JavaScript and by a real browser parses to
the same 45 rows. env_config.pydocumented the Scraping Browser API endpoint with
country-id— a leftover from another repo in this family.
Verified
- 627 offline checks, green in all three engine virtualenvs.
ci_checks --alland--history-checkclean: 47 blobs across 66 objects
that have ever existed, nothing credential-shaped.- A fresh clone with a venv inside it runs the suite green and the
credential scan clean. - All three engines re-run live: 90 rows each over two pages, identical
skus in identical order, every column equal except one review count that
ticked up between runs. Product mode still returns its four variants and
its verified was-price. - Every check added in this release was controlled — the code broken
deliberately, the suite confirmed red, and the expected check confirmed to
be the one that named the failure.
Full detail in
CHANGELOG.md.