Repository navigation
Releases: 2scraper/rakuten-scraper
Release list
v0.2.1 — output is replaced atomically
The output files are now replaced atomically. Before this release an interrupted run could leave a shorter file where a complete one had been — measured, a 2,084-byte good
out.jsoncame back 0 bytes and invalid JSON after a crash mid-write. If you have ever seen a truncated output or sidecar, that was this.CSV cells beginning
=,+,-,@, tab, CR or LF now carry a leading apostrophe, so a spreadsheet reads them as text rather than as a formula. The JSON is unchanged and still carries the site's own bytes;csv_cells_escapedin<out>.meta.jsonsays how many cells the two files differ in. On this site that number is 0 today, measured across 2,662 stored rows, so no existing CSV changes.
Fixed
- Atomic writes.
write_json,write_csvandwrite_run_metausedopen(path, "w"), which truncates before a byte is written. Each now writes to a temporary file in the same directory (os.replaceis only atomic within one filesystem),fsyncs it and renames over the target. The sidecar matters most: it is the file a consumer branches on, so a truncated one beside good rows reads as a broken run over data that is fine. - CSV formula neutralisation, strings only — escaping a number would turn
-5into text and break every sum written over the column. - Scraper API:
waitForis sent as a JSON object. The string form is refused with HTTP 422 and still billed, so every--wait-*run was a paid exit 5. - Scraper API: the target's HTTP status is read from
http_code, not from the API's own verdict string, so the page classifier now sees a target 403/503. - The end-of-listing check compared against the wrong page number, so
--url '...?p=2' --pages 1threw away 45 good rows as "the end of the listing". - The credential scan was blind to JSON-escaped fixtures — the key-shaped-field rule used bare quotes and matched zero times in the largest file in the repository.
- Donor-repo leftovers, none of which changes behaviour.
Changed
- CI tests Python 3.13. The badge has claimed
3.9 | 3.13while the offline matrix ran 3.9 and 3.12 — the ceiling a reader sees first was the one version nobody ran.
Full notes: CHANGELOG.md
v0.2.0 — audit pass; sidecar parity across the three engines
An audit pass over every section of this family's notes, plus a live re-run of
all three engines and both modes. Two real defects, both engine divergences
that no single-engine run could show.
Changes behaviour for an existing consumer
The Selenium and pyppeteer engines were writing a different sidecar from
the Playwright one. They carried three keys belonging to a sibling repo —
scroll,result_header,pages_still_growing, of whichscrollis
meaningless on a site that serves its whole page at once — and they were
missing the cap arithmetic.On this site that arithmetic is what keeps
status: completehonest:
Rakuten served 6,750 results of a query that matched 7,460,738, so a
sidecar saying only "complete" is lying by omission. All three engines now
write the same 20 keys, verified by a live run of each, and the suite
compares the key sets so it cannot drift again.
If you consume <out>.meta.json from the Selenium or pyppeteer engine, the
keys have changed — and now include total_results, reachable_max,
capped_by_site, page_size and pages_beyond_cap.
The other defect: a guard that was only a warning
Two of three engines logged --fingerprint is ignored with --cdp-endpoint
while leaving the flag True. They were correct only because each remote
branch happens to return before the fingerprint is applied — a claim
enforced by where a return sits rather than by the stated gate. The day
someone moves the fingerprint into shared setup, those two engines would
silently start stacking a second identity onto a browser that already has
one, which on this site is the one thing measured to get a client
refused.
All three now force the flag off, and the suite asserts the behaviour —
every call that fetches or applies a fingerprint must sit behind the flag,
followed through the _apply_fingerprint indirection. The first version of
that check reported two engines as ungated, which was the check being wrong
rather than the code.
Also fixed
- A dead
scrollfield on the per-page outcome, and a scroll measurement
ported verbatim from a sibling. Replaced with this site's own: the same URL
fetched by plain HTTP with no JavaScript and by a real browser parses to
the same 45 rows. env_config.pydocumented the Scraping Browser API endpoint with
country-id— a leftover from another repo in this family.
Verified
- 627 offline checks, green in all three engine virtualenvs.
ci_checks --alland--history-checkclean: 47 blobs across 66 objects
that have ever existed, nothing credential-shaped.- A fresh clone with a venv inside it runs the suite green and the
credential scan clean. - All three engines re-run live: 90 rows each over two pages, identical
skus in identical order, every column equal except one review count that
ticked up between runs. Product mode still returns its four variants and
its verified was-price. - Every check added in this release was controlled — the code broken
deliberately, the suite confirmed red, and the expected check confirmed to
be the one that named the failure.
Full detail in
CHANGELOG.md.
v0.1.0 — Rakuten Ichiba to JSON/CSV
First release. Reads Rakuten Ichiba listings and
product pages to JSON or CSV, with four interchangeable back ends and the row
schema shared across the 2scraper family.
You need nothing to run this
No key, no proxy, no account. Measured 2026-09-21 from a datacentre address:
a keyword search returned 45 products a page, and a three-page run
returned 135 rows, 100% with a price, status: complete.
The canary in this repository is that claim under test — a real three-page
scrape, daily, from a bare GitHub runner, with no secrets. Its first
dispatch, from GitHub's own address rather than ours:
canary OK: 135 rows, 100% priced, complete, 6,750 of 142,106 results reachable
What this site turned out to be
- The gate is the CLIENT, not the address. From one address, unchanged:
curlwith curl's own User-Agent was served 92 KB and 45 products; the
samecurlclaiming a Chrome User-Agent got a 43-byte Akamai deny;
headless Chromium was served 855 KB. What Akamai refuses here is a claimed
identity that disagrees with the TLS fingerprint under it — so if you are
refused, look for a disguise you added before you buy an exit. - A refusal arrives under HTTP 200. The status code is not the signal.
- HTTP 503 is a throttle, and it renders the identical body to the 403
refusal — only the status separates them. A 503 costs a wait at the same
exit and does not count as blocked. - The listing JSON-LD is a trap: one
ItemListholding ten
SEO-carousel items beside a page of 45 products. It is read for one thing
only — the currency, which the payload never states, and which Rakuten
publishes on page 1 only. - Detail pages are EUC-JP while listing pages are UTF-8.
- Pagination lies past the end.
?p=151of a 150-page query redirects to
page 1 and serves it with HTTP 200;/category/{id}/?p=2ignores the
parameter entirely. Both look like success, so every page is checked
against the offset the server states. - Complete is not exhaustive. Every query is capped at 6,750 results (150
pages of 45) however many it matched — one query reportednumFound: 3,053,682against that same 6,750 — so the sidecar records both numbers. - No captcha is configured anywhere, across 13 captures, served and
refused alike.
Modes
--mode listing a keyword search, a genre listing or a genre landing
page — 45 products a page
--mode product one item page, adding per-variant prices, a variant
count and a verified was-price
There is deliberately no --mode shop: a merchant storefront carries an
empty payload and no product grid.
Verification
565 offline checks, green with no engine library installed and in each
engine's own virtualenv. Fixtures cut from 13 real captures, each proven to
parse identically to the untrimmed original and scrubbed of two front-end
API keys, a customer's review nickname and text, and ad session ids. All
three browser engines produced 90 identical rows over two pages, every
column equal. Every new check was controlled — the code broken deliberately,
the suite confirmed red, and the expected check confirmed to be the one that
named it.
Five defects were found by running the code rather than reading it, three of
them inherited from this family's shared core — including an uncounted
captcha solve budget, and a user-agent override that got one engine refused
from its second navigation onward while page 1 looked fine. Full list in
CHANGELOG.md.