v2.0.0
Breaking changes
This release removes public names, so it is a major version. It also changes what video
intake does when YouTube asks for a bot check. What changed, and what to do:
- The
stealthextra is gone.pip install 'gather-engine[stealth]'still installs Gather;
pip warns that the extra does not exist and installs the core. Drop[stealth]from
requirement files and install commands.gather-engine[all]now installslxmland
playwrightonly. - The module
gather.backends_stealthis gone, and importing it raisesImportError.
Remove the import. Nothing replaces it. - The constant
gather.backends.CAP_STEALTHis gone, and importing it raisesImportError.
gather capsno longer listsstealth, even wherecurl_cffiis installed. Remove any
reference to it. - A YouTube bot check ("confirm you're not a bot") is never retried. Builds from main after
1.9.1 retried it with backoff like HTTP 429; 1.9.1 did not retry it either. Gather records
the failure with the codebot-check, and the record starts "YouTube asked for a bot
check, and gather stopped", then gives yt-dlp's own line.gather channelstops the pass
at the first bot check, whatever--max-throttledsays, and exits 1. It settles the entry
that met the check, so a resumed run does not ask for it again. A bot check while listing
a tab stops the run before any entry is gathered. What to do: when you choose to try a
video again, ask for it by name withgather video URL --store DIRand the pass's flags.
--retriesand the backoff flags now cover HTTP 429 and YouTube's session rate limit
only.gather.ytdlp.THROTTLE_CODESno longer holdsbot-check;TERMINAL_CODESdoes.
Removed: the stealth extra and its backend
- The
stealthextra, itscurl_cffidependency, and thestealthcapability backend
(gather.backends_stealth) are removed. That backend impersonated a browser's TLS
fingerprint to get past bot detection. Gather no longer ships bot-detection evasion of any
kind, and nothing replaces it. Gather's own HTTP requests go out with Gather's own
User-Agent. pip install 'gather-engine[stealth]'still installs Gather. pip warns
gather-engine 2.0.0 does not provide the extra 'stealth'and installs the core
without it. Drop[stealth]from requirement files and install commands.gather-engine[all]now installslxmlandplaywrightonly. An upgrade leaves an
installedcurl_cffiin place. Gather's own code no longer uses it, but a yt-dlp in
the same environment can (see below), so runpip uninstall curl_cffiif nothing else
needs it.- Caption downloads no longer pose as a browser. yt-dlp marks every YouTube caption track
for impersonation and keeps the mark in the info JSON Gather saves and hands back for the
caption download. Where yt-dlp could importcurl_cffi, the track went out with a
browser's TLS fingerprint and headers. Gather now removes the mark first. A test on the
saved info runs everywhere, and a test against real yt-dlp, run where yt-dlp is
installed, checks that the caption request carries yt-dlp's own headers. - What yt-dlp still decides: it sends its own default headers, including a desktop Chrome
User-Agent whose version it picks each run, and for some sites other than YouTube its
extractors ask for impersonation while they extract. Gather passes no flag that asks for
either, and yt-dlp has no flag that turns off an extractor's request. That request takes
effect only where yt-dlp can importcurl_cffi. gather capsno longer listsstealth, including on a machine wherecurl_cffiis
installed.- Code that imports
gather.backends_stealthorgather.backends.CAP_STEALTHnow raises
ImportError. Remove those imports. - Tests fail if a dependency or extra names a known fingerprint-impersonation,
patched-browser, or challenge-solving package, if a source file imports or looks one up
or passes yt-dlp--impersonate, or if an installedcurl_cffiregisters a capability.
The package list matches by name, so a new tool under another name still needs review. - The credential strip on a same-host redirect from https to http has its own test for both
redirect handlers, the onehttp_getuses and the one the accountablefetchuses. The
removed stealth tests were the only ones that covered that branch.
Video intake pacing and channel runs
gather channel URL --store DIRlists a channel'svideos,shorts, andstreams
tabs (or one playlist) with--flat-playlistand gathers each entry into the corpus with
bounded concurrency (default 2) and paced entry starts. A per-pass ledger under
DIR/intake/makes the run resumable, andsummary-<pass>.jsoncounts entries per tab,
captions (manual, auto, missing by reason), comments, failures by reason, and retries.
A run killed mid-write leaves an unfinished last row; the next run drops it and gathers
that entry again. Any other unreadable row stops the run, before it calls yt-dlp, with
the line number and exit status 1.- Separate passes:
--no-captionsgathers metadata and comments without touching the
caption endpoint;--captions-onlystores only the transcript item. - A
gather runconfig or an MCPgather.runvideo job takes
"captions": "with" | "skip" | "only"(defaultwith). Any other value is a config error,
raised before any job runs. On MCP the job still needs thevideonetwork grant,
whichever pass it asks for. - The run summary names its files relative to
--store(a--summaryoutside the store by
file name) and records the yt-dlp program by file name and a JS runtime without its
path, so a summary you pass on carries no local path. --timeoutmust be above 0, and--sleep-requestsand--sleep-subtitlesmust be 0 or
more. Any other value exits 2 before yt-dlp starts.- Caption intake downloads exactly one track per video, chosen from the info JSON: manual
first, then the original-language auto-caption (en-orig). The olden.*pattern fetched
every English variant and could pick a machine translation; a translation-only video is
now recorded as missing with the reasontranslation-only. - HTTP 429 and YouTube's session rate limit are retried with exponential backoff and
jitter, bounded by attempts and by total wait. Every retry and final failure is logged and
recorded. A channel run stops starting new entries once an entry spends its whole budget
still throttled, and records the rest as stopped. A bot check is never retried and stops
the run at once (see Breaking changes). - The extraction runs with
--ignore-no-formats-error, so a video whose formats are missing
still yields its metadata and caption tracks. With that flag yt-dlp reports YouTube's
playability reason as a warning and exits 0. When the extraction lists no formats, Gather
reads that warning: a session rate limit is retried like an HTTP 429, a bot check ends
the entry without a retry, and a private, members-only, age-restricted or removed video is
recorded as failed with that reason and settled. A geo-blocked or upcoming video is
recorded as failed and tried again on the next run. None of them stores a metadata item
or a "no captions offered" outcome. - yt-dlp runs with
--js-runtimes nodewhen it can startnode(--js-runtimeoverrides),
and--sleep-requests/--sleep-subtitlespass through. The check uses the same PATH
lookup as every child Gather starts, so anodeonly the working folder holds does not
count. - Failure messages report yt-dlp's
ERRORlines instead of the first 160 characters of
stderr, which was often a version warning. - A timeout, a missing yt-dlp binary, or a refused start is recorded as a failed call
(timeout,tool-missing,tool-refused), not an exception. - Every yt-dlp call (tab listing, extraction, caption download) starts the way every
other tool does, from an absolute path in a private empty folder with an environment
allowlist, and carries--ignore-config, so noyt-dlp.confchanges what it runs.