Skip to content

v0.1.0 — Rakuten Ichiba to JSON/CSV

Choose a tag to compare

@jehrr jehrr released this 21 Sep 09:48
· 13 commits to main since this release

First release. Reads Rakuten Ichiba listings and
product pages to JSON or CSV, with four interchangeable back ends and the row
schema shared across the 2scraper family.

You need nothing to run this

No key, no proxy, no account. Measured 2026-09-21 from a datacentre address:
a keyword search returned 45 products a page, and a three-page run
returned 135 rows, 100% with a price, status: complete.

The canary in this repository is that claim under test — a real three-page
scrape, daily, from a bare GitHub runner, with no secrets. Its first
dispatch, from GitHub's own address rather than ours:

canary OK: 135 rows, 100% priced, complete, 6,750 of 142,106 results reachable

What this site turned out to be

  • The gate is the CLIENT, not the address. From one address, unchanged:
    curl with curl's own User-Agent was served 92 KB and 45 products; the
    same curl claiming a Chrome User-Agent got a 43-byte Akamai deny;
    headless Chromium was served 855 KB. What Akamai refuses here is a claimed
    identity that disagrees with the TLS fingerprint under it — so if you are
    refused, look for a disguise you added before you buy an exit.
  • A refusal arrives under HTTP 200. The status code is not the signal.
  • HTTP 503 is a throttle, and it renders the identical body to the 403
    refusal — only the status separates them. A 503 costs a wait at the same
    exit and does not count as blocked.
  • The listing JSON-LD is a trap: one ItemList holding ten
    SEO-carousel items beside a page of 45 products. It is read for one thing
    only — the currency, which the payload never states, and which Rakuten
    publishes on page 1 only.
  • Detail pages are EUC-JP while listing pages are UTF-8.
  • Pagination lies past the end. ?p=151 of a 150-page query redirects to
    page 1 and serves it with HTTP 200; /category/{id}/?p=2 ignores the
    parameter entirely. Both look like success, so every page is checked
    against the offset the server states.
  • Complete is not exhaustive. Every query is capped at 6,750 results (150
    pages of 45) however many it matched — one query reported numFound: 3,053,682 against that same 6,750 — so the sidecar records both numbers.
  • No captcha is configured anywhere, across 13 captures, served and
    refused alike.

Modes

--mode listing   a keyword search, a genre listing or a genre landing
                 page — 45 products a page
--mode product   one item page, adding per-variant prices, a variant
                 count and a verified was-price

There is deliberately no --mode shop: a merchant storefront carries an
empty payload and no product grid.

Verification

565 offline checks, green with no engine library installed and in each
engine's own virtualenv. Fixtures cut from 13 real captures, each proven to
parse identically to the untrimmed original and scrubbed of two front-end
API keys, a customer's review nickname and text, and ad session ids. All
three browser engines produced 90 identical rows over two pages, every
column equal. Every new check was controlled — the code broken deliberately,
the suite confirmed red, and the expected check confirmed to be the one that
named it.

Five defects were found by running the code rather than reading it, three of
them inherited from this family's shared core — including an uncounted
captcha solve budget, and a user-agent override that got one engine refused
from its second navigation onward while page 1 looked fine. Full list in
CHANGELOG.md.