Releases: 2scraper/googleplay-scraper
Release list
v0.1.0 — five modes, four back ends, and a canary that needs no key
First release. Google Play to JSON or CSV, with one row schema shared across the 2scraper family.
pip install -r requirements.txt -r requirements-playwright.txt
playwright install chromium
python3 playwright_scraper.py --mode listing --category GAME --pages 3The thing to know first
Google Play computes a rating per COUNTRY, and the count is global. One app, one datacentre address, only gl changed on 2026-09-22:
--gl |
rating | currency | ratings counted |
|---|---|---|---|
| US | 4.08 | USD | 6,312,298 |
| DE | 3.90 | EUR | 6,312,298 |
| JP | 3.96 | JPY | 6,312,298 |
| IN | 4.13 | INR | 6,312,298 |
| BR | 4.49 | BRL | 6,312,298 |
The score moves by more than half a star; the count is identical. So gl and hl are columns on every row here, never run metadata, and no market's rating is averaged into another's. All five rows came from one address — you do not need an exit per country.
You need no key, no proxy and no account
Measured from a datacentre address, and re-measured every morning from a bare GitHub runner: every route answers plain curl with no cookies, under curl's own User-Agent and a Chrome one alike, and no captcha marker of any spelling appears in twelve captures. The canary runs a real three-page listing scrape and a real reviews run daily with no secrets and is expected to be green — which is the only honest way to keep a sentence like that true.
Five modes
--mode listing · search · app · developer · reviews. The four app modes share one row class keyed on the package id, so their files join on sku.
What only the app page publishes
The rating histogram, the split between ratings (244,140,056) and written reviews (1,952,258) on one app, the version, the update date, the developer's contact details, Play's own chart standing — and the exact install count: 12,322,703,000 beside a displayed "10,000,000,000+". Measured across 330 listing tiles, zero carry the exact figure.
Reviews
Page 1 is embedded in the app page and costs no extra request. Everything after it comes from Play's own batchexecute, which is not capped at the 40 its front end asks for: 2000 in one call returned 2000, with a continuation token, and no ceiling was found.
Reviews are about identifiable people, so --no-authors drops a reviewer's name, profile id and avatar together, and the committed captures are scrubbed structurally with the result guarded by pattern.
Stated plainly, because none of it is a limit of the store
- No
--mode booksor--mode movies: both storefronts answer and neither has been captured or parsed, so those URLs are refused with that reason. - Only
--sort newest: the other two sort keys return an empty payload in this request shape. - No review score filter: the payload's slot for one changes nothing.
--pagesabove 1 warns on--mode search: Play answers a query once (22 to 50 results, unchanged over four scroll rounds).
Engines
Playwright, Selenium, pyppeteer and the 2Captcha Scraping Browser API over --cdp-endpoint. Verified live to agree: one category through three of them gave 52 rows and three identical sets of package ids; a second gave 37, 37 and 37. The CLI, the mode loops and main() live in engine_core.py, so a flag cannot exist on one engine and not its twins.
Install exactly one engine per virtualenv — Playwright and pyppeteer pin incompatible pyee versions, and pyppeteer and Selenium collide on urllib3.