Releases: 2scraper/binance-scraper
Release list
v0.2.0
Added
http_scraper.py: the three modes with no browser, through the same
fetch loop, output and exit codes. Same flags as the browser engines
minus the ones that describe a browser. It sends requests' own
User-Agent rather than a Chrome one over a non-Chrome TLS handshake.
Measured 2026-09-30: identical rows to Playwright in all three modes, in
about a third to a half of the time. If AWS WAF answers instead of an
endpoint it reports exit 3 straight away, without waiting for a
challenge script that nothing here can run.- The canary runs P2P through it daily, with no browser installed.
v0.1.1
A moved payload shape is no longer reported as an empty listing. A
page the endpoint served, whose own total says it holds rows, and from
which no row could be read, used to end the run as exit 4 ("the listing
has nothing in it"). It is now exit 5 on page 1 and exit 6 on a later
page, withstop_reason: parser_found_nothing. Measured by renaming the
row container in a real P2P capture that still counted 187 adverts.
Changed
- The sidecar carries
core_field_shortfall: for each page, the core
columns filled on fewer than 99% of its rows.{}on a healthy run. A
renamed field used to write a complete-looking file with the column null
on every row and one warning in the log; the daily canary now fails on it. - The canary names a moved shape on exit 5 and 6.
- CI tests Python 3.9 and 3.14 (was 3.9 and 3.12). 3.9 stays as the floor.
- The AWS WAF token cookie is set on the domain the challenge page lists in
awsWafCookieDomainList, or on the page's own host when it lists none,
rather than always on the registrable domain. engine-smokechecks that the engine imports, rather than grepping the
suite's output for a phrase it never printed.
v0.1.0 — P2P, copy-trading and announcements
No key, no proxy and no account needed for any mode: measured 2026-09-24 from a datacentre VPS and from a GitHub-hosted runner.
First release. Three modes over binance.com's own JSON endpoints, three
browser engines over one shared fetch loop, and the 2Captcha Scraper API for
the one mode it can reach.
Added
--mode p2p: the P2P order book for an asset/fiat pair. One row per
advert: price, per-order limits, payment methods, time limit, and the
advertiser's 30-day orders, completion and feedback. Both sides of the
trade are kept (sideis what was asked,advertiser_sideis what the
advert says, and they are always opposite).--mode copytrading: Futures copy-trading lead portfolios, one row per
portfolio: ROI, PnL, max drawdown, win rate, AUM, copier PnL, Sharpe,
copiers and seats, badge, the period and the ordering.--mode announcements: one announcement catalogue (new listings,
delistings, news, activities, maintenance, API updates, airdrops), one row
per article, with the site's canonical/detail/{code}address.--urlreads the query from a P2P trade page, the copy-trading page or an
announcement catalogue address; the page itself is never fetched.- Every query parameter is allowlisted, and
--pay-typeis checked against
the site's own list for the fiat, because the API answers several wrong
values with a plausible response instead of an error. - Pages are planned from the total page 1 states; the sidecar records
total_results,pages_availableand the query. - AWS WAF: its CAPTCHA is solved with AmazonTask / AmazonTaskProxyless, and
the solution'sexisting_tokenis set asaws-waf-tokenon the
registrable domain, the arrangement measured to clear it. diff_runs.pydiffs two runs of one mode bysku, over columns derived
from the row class rather than listed by hand.- An offline suite over real, scrubbed API responses, including an
end-to-end run of the shared fetch loop with a fake browser, and a daily
canary of all three modes with no secrets.
Live-verified on 2026-09-24: all three engines through the same seven scenarios, the Scraping Browser API (country-de) on Playwright and pyppeteer (8 of 8 runs), an AWS WAF AmazonTask solve on a gated page, the Scraper API for announcements, and the canary from GitHub in all three modes.