Skip to content

Releases: 2scraper/rakuten-scraper

v0.2.1 — output is replaced atomically

Choose a tag to compare

@jehrr jehrr released this 30 Sep 07:57

The output files are now replaced atomically. Before this release an interrupted run could leave a shorter file where a complete one had been — measured, a 2,084-byte good out.json came back 0 bytes and invalid JSON after a crash mid-write. If you have ever seen a truncated output or sidecar, that was this.

CSV cells beginning =, +, -, @, tab, CR or LF now carry a leading apostrophe, so a spreadsheet reads them as text rather than as a formula. The JSON is unchanged and still carries the site's own bytes; csv_cells_escaped in <out>.meta.json says how many cells the two files differ in. On this site that number is 0 today, measured across 2,662 stored rows, so no existing CSV changes.

Fixed

  • Atomic writes. write_json, write_csv and write_run_meta used open(path, "w"), which truncates before a byte is written. Each now writes to a temporary file in the same directory (os.replace is only atomic within one filesystem), fsyncs it and renames over the target. The sidecar matters most: it is the file a consumer branches on, so a truncated one beside good rows reads as a broken run over data that is fine.
  • CSV formula neutralisation, strings only — escaping a number would turn -5 into text and break every sum written over the column.
  • Scraper API: waitFor is sent as a JSON object. The string form is refused with HTTP 422 and still billed, so every --wait-* run was a paid exit 5.
  • Scraper API: the target's HTTP status is read from http_code, not from the API's own verdict string, so the page classifier now sees a target 403/503.
  • The end-of-listing check compared against the wrong page number, so --url '...?p=2' --pages 1 threw away 45 good rows as "the end of the listing".
  • The credential scan was blind to JSON-escaped fixtures — the key-shaped-field rule used bare quotes and matched zero times in the largest file in the repository.
  • Donor-repo leftovers, none of which changes behaviour.

Changed

  • CI tests Python 3.13. The badge has claimed 3.9 | 3.13 while the offline matrix ran 3.9 and 3.12 — the ceiling a reader sees first was the one version nobody ran.

Full notes: CHANGELOG.md

v0.2.0 — audit pass; sidecar parity across the three engines

Choose a tag to compare

@jehrr jehrr released this 21 Sep 13:17

An audit pass over every section of this family's notes, plus a live re-run of
all three engines and both modes. Two real defects, both engine divergences
that no single-engine run could show.

Changes behaviour for an existing consumer

The Selenium and pyppeteer engines were writing a different sidecar from
the Playwright one. They carried three keys belonging to a sibling repo —
scroll, result_header, pages_still_growing, of which scroll is
meaningless on a site that serves its whole page at once — and they were
missing the cap arithmetic.

On this site that arithmetic is what keeps status: complete honest:
Rakuten served 6,750 results of a query that matched 7,460,738, so a
sidecar saying only "complete" is lying by omission. All three engines now
write the same 20 keys, verified by a live run of each, and the suite
compares the key sets so it cannot drift again.

If you consume <out>.meta.json from the Selenium or pyppeteer engine, the
keys have changed — and now include total_results, reachable_max,
capped_by_site, page_size and pages_beyond_cap.

The other defect: a guard that was only a warning

Two of three engines logged --fingerprint is ignored with --cdp-endpoint
while leaving the flag True. They were correct only because each remote
branch happens to return before the fingerprint is applied — a claim
enforced by where a return sits rather than by the stated gate. The day
someone moves the fingerprint into shared setup, those two engines would
silently start stacking a second identity onto a browser that already has
one, which on this site is the one thing measured to get a client
refused
.

All three now force the flag off, and the suite asserts the behaviour —
every call that fetches or applies a fingerprint must sit behind the flag,
followed through the _apply_fingerprint indirection. The first version of
that check reported two engines as ungated, which was the check being wrong
rather than the code.

Also fixed

  • A dead scroll field on the per-page outcome, and a scroll measurement
    ported verbatim from a sibling. Replaced with this site's own: the same URL
    fetched by plain HTTP with no JavaScript and by a real browser parses to
    the same 45 rows.
  • env_config.py documented the Scraping Browser API endpoint with
    country-id — a leftover from another repo in this family.

Verified

  • 627 offline checks, green in all three engine virtualenvs.
  • ci_checks --all and --history-check clean: 47 blobs across 66 objects
    that have ever existed, nothing credential-shaped.
  • A fresh clone with a venv inside it runs the suite green and the
    credential scan clean.
  • All three engines re-run live: 90 rows each over two pages, identical
    skus in identical order, every column equal except one review count that
    ticked up between runs. Product mode still returns its four variants and
    its verified was-price.
  • Every check added in this release was controlled — the code broken
    deliberately, the suite confirmed red, and the expected check confirmed to
    be the one that named the failure.

Full detail in
CHANGELOG.md.

v0.1.0 — Rakuten Ichiba to JSON/CSV

Choose a tag to compare

@jehrr jehrr released this 21 Sep 09:48

First release. Reads Rakuten Ichiba listings and
product pages to JSON or CSV, with four interchangeable back ends and the row
schema shared across the 2scraper family.

You need nothing to run this

No key, no proxy, no account. Measured 2026-09-21 from a datacentre address:
a keyword search returned 45 products a page, and a three-page run
returned 135 rows, 100% with a price, status: complete.

The canary in this repository is that claim under test — a real three-page
scrape, daily, from a bare GitHub runner, with no secrets. Its first
dispatch, from GitHub's own address rather than ours:

canary OK: 135 rows, 100% priced, complete, 6,750 of 142,106 results reachable

What this site turned out to be

  • The gate is the CLIENT, not the address. From one address, unchanged:
    curl with curl's own User-Agent was served 92 KB and 45 products; the
    same curl claiming a Chrome User-Agent got a 43-byte Akamai deny;
    headless Chromium was served 855 KB. What Akamai refuses here is a claimed
    identity that disagrees with the TLS fingerprint under it — so if you are
    refused, look for a disguise you added before you buy an exit.
  • A refusal arrives under HTTP 200. The status code is not the signal.
  • HTTP 503 is a throttle, and it renders the identical body to the 403
    refusal — only the status separates them. A 503 costs a wait at the same
    exit and does not count as blocked.
  • The listing JSON-LD is a trap: one ItemList holding ten
    SEO-carousel items beside a page of 45 products. It is read for one thing
    only — the currency, which the payload never states, and which Rakuten
    publishes on page 1 only.
  • Detail pages are EUC-JP while listing pages are UTF-8.
  • Pagination lies past the end. ?p=151 of a 150-page query redirects to
    page 1 and serves it with HTTP 200; /category/{id}/?p=2 ignores the
    parameter entirely. Both look like success, so every page is checked
    against the offset the server states.
  • Complete is not exhaustive. Every query is capped at 6,750 results (150
    pages of 45) however many it matched — one query reported numFound: 3,053,682 against that same 6,750 — so the sidecar records both numbers.
  • No captcha is configured anywhere, across 13 captures, served and
    refused alike.

Modes

--mode listing   a keyword search, a genre listing or a genre landing
                 page — 45 products a page
--mode product   one item page, adding per-variant prices, a variant
                 count and a verified was-price

There is deliberately no --mode shop: a merchant storefront carries an
empty payload and no product grid.

Verification

565 offline checks, green with no engine library installed and in each
engine's own virtualenv. Fixtures cut from 13 real captures, each proven to
parse identically to the untrimmed original and scrubbed of two front-end
API keys, a customer's review nickname and text, and ad session ids. All
three browser engines produced 90 identical rows over two pages, every
column equal. Every new check was controlled — the code broken deliberately,
the suite confirmed red, and the expected check confirmed to be the one that
named it.

Five defects were found by running the code rather than reading it, three of
them inherited from this family's shared core — including an uncounted
captcha solve budget, and a user-agent override that got one engine refused
from its second navigation onward while page 1 looked fine. Full list in
CHANGELOG.md.