Skip to content

Releases: 2scraper/googleplay-scraper

v0.1.0 — five modes, four back ends, and a canary that needs no key

Choose a tag to compare

@jehrr jehrr released this 22 Sep 14:01

First release. Google Play to JSON or CSV, with one row schema shared across the 2scraper family.

pip install -r requirements.txt -r requirements-playwright.txt
playwright install chromium
python3 playwright_scraper.py --mode listing --category GAME --pages 3

The thing to know first

Google Play computes a rating per COUNTRY, and the count is global. One app, one datacentre address, only gl changed on 2026-09-22:

--gl rating currency ratings counted
US 4.08 USD 6,312,298
DE 3.90 EUR 6,312,298
JP 3.96 JPY 6,312,298
IN 4.13 INR 6,312,298
BR 4.49 BRL 6,312,298

The score moves by more than half a star; the count is identical. So gl and hl are columns on every row here, never run metadata, and no market's rating is averaged into another's. All five rows came from one address — you do not need an exit per country.

You need no key, no proxy and no account

Measured from a datacentre address, and re-measured every morning from a bare GitHub runner: every route answers plain curl with no cookies, under curl's own User-Agent and a Chrome one alike, and no captcha marker of any spelling appears in twelve captures. The canary runs a real three-page listing scrape and a real reviews run daily with no secrets and is expected to be green — which is the only honest way to keep a sentence like that true.

Five modes

--mode listing · search · app · developer · reviews. The four app modes share one row class keyed on the package id, so their files join on sku.

What only the app page publishes

The rating histogram, the split between ratings (244,140,056) and written reviews (1,952,258) on one app, the version, the update date, the developer's contact details, Play's own chart standing — and the exact install count: 12,322,703,000 beside a displayed "10,000,000,000+". Measured across 330 listing tiles, zero carry the exact figure.

Reviews

Page 1 is embedded in the app page and costs no extra request. Everything after it comes from Play's own batchexecute, which is not capped at the 40 its front end asks for: 2000 in one call returned 2000, with a continuation token, and no ceiling was found.

Reviews are about identifiable people, so --no-authors drops a reviewer's name, profile id and avatar together, and the committed captures are scrubbed structurally with the result guarded by pattern.

Stated plainly, because none of it is a limit of the store

  • No --mode books or --mode movies: both storefronts answer and neither has been captured or parsed, so those URLs are refused with that reason.
  • Only --sort newest: the other two sort keys return an empty payload in this request shape.
  • No review score filter: the payload's slot for one changes nothing.
  • --pages above 1 warns on --mode search: Play answers a query once (22 to 50 results, unchanged over four scroll rounds).

Engines

Playwright, Selenium, pyppeteer and the 2Captcha Scraping Browser API over --cdp-endpoint. Verified live to agree: one category through three of them gave 52 rows and three identical sets of package ids; a second gave 37, 37 and 37. The CLI, the mode loops and main() live in engine_core.py, so a flag cannot exist on one engine and not its twins.

Install exactly one engine per virtualenv — Playwright and pyppeteer pin incompatible pyee versions, and pyppeteer and Selenium collide on urllib3.