Repository navigation
v0.1.0 — Rakuten Ichiba to JSON/CSV
First release. Reads Rakuten Ichiba listings and
product pages to JSON or CSV, with four interchangeable back ends and the row
schema shared across the 2scraper family.
You need nothing to run this
No key, no proxy, no account. Measured 2026-09-21 from a datacentre address:
a keyword search returned 45 products a page, and a three-page run
returned 135 rows, 100% with a price, status: complete.
The canary in this repository is that claim under test — a real three-page
scrape, daily, from a bare GitHub runner, with no secrets. Its first
dispatch, from GitHub's own address rather than ours:
canary OK: 135 rows, 100% priced, complete, 6,750 of 142,106 results reachable
What this site turned out to be
- The gate is the CLIENT, not the address. From one address, unchanged:
curlwith curl's own User-Agent was served 92 KB and 45 products; the
samecurlclaiming a Chrome User-Agent got a 43-byte Akamai deny;
headless Chromium was served 855 KB. What Akamai refuses here is a claimed
identity that disagrees with the TLS fingerprint under it — so if you are
refused, look for a disguise you added before you buy an exit. - A refusal arrives under HTTP 200. The status code is not the signal.
- HTTP 503 is a throttle, and it renders the identical body to the 403
refusal — only the status separates them. A 503 costs a wait at the same
exit and does not count as blocked. - The listing JSON-LD is a trap: one
ItemListholding ten
SEO-carousel items beside a page of 45 products. It is read for one thing
only — the currency, which the payload never states, and which Rakuten
publishes on page 1 only. - Detail pages are EUC-JP while listing pages are UTF-8.
- Pagination lies past the end.
?p=151of a 150-page query redirects to
page 1 and serves it with HTTP 200;/category/{id}/?p=2ignores the
parameter entirely. Both look like success, so every page is checked
against the offset the server states. - Complete is not exhaustive. Every query is capped at 6,750 results (150
pages of 45) however many it matched — one query reportednumFound: 3,053,682against that same 6,750 — so the sidecar records both numbers. - No captcha is configured anywhere, across 13 captures, served and
refused alike.
Modes
--mode listing a keyword search, a genre listing or a genre landing
page — 45 products a page
--mode product one item page, adding per-variant prices, a variant
count and a verified was-price
There is deliberately no --mode shop: a merchant storefront carries an
empty payload and no product grid.
Verification
565 offline checks, green with no engine library installed and in each
engine's own virtualenv. Fixtures cut from 13 real captures, each proven to
parse identically to the untrimmed original and scrubbed of two front-end
API keys, a customer's review nickname and text, and ad session ids. All
three browser engines produced 90 identical rows over two pages, every
column equal. Every new check was controlled — the code broken deliberately,
the suite confirmed red, and the expected check confirmed to be the one that
named it.
Five defects were found by running the code rather than reading it, three of
them inherited from this family's shared core — including an uncounted
captcha solve budget, and a user-agent override that got one engine refused
from its second navigation onward while page 1 looked fine. Full list in
CHANGELOG.md.