Skip to content

Releases: hamza140202/agent-video-downloader

v1.2.2 — remove all specific test content references from public docs

Choose a tag to compare

@hamza140202 hamza140202 released this 04 Oct 00:41

What's new in v1.2.2

This release removes all specific test content references (artist names, post IDs, author handles, specific video URLs) from public-facing documentation, PyPI metadata, and GitHub release notes.

Why

Specific test video identities appearing in public PyPI metadata could be misinterpreted as the project's purpose or content focus. The actual purpose — a generic video downloader for AI agents — should not be associated with any specific content domain.

What was cleaned

Files cleaned (specific test URLs/names replaced with generic placeholders):

  • README.md (PyPI long_description)
  • CLAUDE.md
  • PHASES.md (Phase 6 renamed from 'Babymonster batch test' → 'Real-world batch test')
  • PLAN.md
  • AGENTS.md, SKILLS.md, TECHSTACK.md
  • docs/research-blog.md, docs/research-report.md, docs/endpoint-matrix.md
  • src/avd/cli.py, src/avd/tester.py
  • src/tests/test_tiktok.py, src/tests/test_agents.py

Files removed (test fixtures with specific content references):

  • babymonster_batch.txt
  • src/tests/babymonster_urls_*.json
  • src/tests/{reddit,instagram,douyin,rednote,real_tweet}_methods.json
  • src/tests/reddit_methods_github_actions_workflow.yaml
  • tests/results/ (test run output with specific URLs)

Files added:

  • src/tests/sample_urls.json.example — template with placeholder URLs
  • tests/sample_urls.json.example — same template

Gitignore updates:

  • tests/sample_urls.json and src/tests/sample_urls.json are now gitignored (users copy from .example and populate with their own test URLs)
  • *_batch_urls.txt pattern added (any future batch test files)
  • src/agent_video_downloader.egg-info/ added
  • tests/results/ added

GitHub release notes for v1.2.0 and v1.2.1 also updated

The release notes for v1.2.0 and v1.2.1 were edited via the GitHub API to remove all specific test content references. The release notes now describe the real-world batch test in generic terms (counts and sizes only, no content descriptions).

PyPI

  • v1.2.2 published: https://pypi.org/project/agent-video-downloader/1.2.2/
  • Wheel verified CLEAN — automated scan confirmed no test content references in any .py or METADATA file inside the wheel
  • 66 KB wheel + 50 KB sdist
  • Install: pip install agent-video-downloader (now installs v1.2.2)

Test results

  • 30/30 unit tests pass
  • No code behavior changes — the extractors, Verifier, Truth Agent, Tester, and CLI all work identically to v1.2.1
  • The smoke test (avd test --smoke) now uses placeholder URLs by default; users must populate sample_urls.json from the .example template to enable functional smoke testing

Migration notes

Existing v1.2.0 or v1.2.1 users have no required upgrade — code behavior is identical. The v1.2.2 release is purely a documentation and metadata cleanup.

If you have a local sample_urls.json with real test URLs, it will continue to work — the file is just no longer tracked in git. If you don't have one, copy from sample_urls.json.example and populate with your own test URLs.

v1.2.1 — reference-architecture attribution + contributor addition

Choose a tag to compare

@hamza140202 hamza140202 released this 03 Oct 23:55

What's new in v1.2.1

This is a documentation and attribution release. No code behavior changes.

Research blog expanded (705 → 839 lines)

New subsection §3.1 "The Reference Architecture — How Existing Agent Repos Shaped This Build" with detailed per-repo breakdown of architectural patterns adopted from the Bilal140202/* family of reference repos:

  • ttagent (TikTok) — 4-slot chain (TikWM → embed/v2 → tiklydown → oEmbed) adopted verbatim with CDN allowlist extension; honest-empty pattern adopted across all platforms
  • igagent (Instagram) — critical contribution: documented the facebookexternalhit/1.1 User-Agent bypass for Instagram's datacenter-IP wall. Without igagent's README, the v1.1 Instagram fix would not have been attempted.
  • xthread-agent (Twitter/X) — fxtwitter→vxtwitter decoder chain, thread walker with replying_to_status filtering, living endpoint matrix format adopted verbatim as docs/endpoint-matrix.md
  • ytagent (YouTube, doctrine source) — 6-layer Verifier (adopted verbatim with adaptive-size + moov-on-failure modifications), 13-method tiered bypass doctrine (adopted as conceptual framework), Truth Agent pattern (simplified to advisory per-call), MCP wrapper (adopted verbatim), "slots, not brands" principle (foundational)

A new "What was NOT adopted" section documents research-integrity limitations: single-file architecture, per-method success ranking, 13-method chain depth, browser-fallback extractor.

2 new "Tech in a Minute" yellow-banner boxes added (endpoint matrix, innertube). Total: 21 boxes throughout the blog.

New CONTRIBUTORS.md file

Research-grade attribution:

  • ansaribilal1402 listed as author/maintainer (PyPI publisher)
  • Bilal140202/ytagent, ttagent, igagent, xthread-agent listed as research sources with specific architectural contributions
  • The repos are NOT dependencies (not imported, not vendored) — attribution is intellectual, not technical

pyproject.toml updates

  • authors field updated to ansaribilal1402 (PyPI publisher)
  • maintainers field added
  • Comment added pointing to CONTRIBUTORS.md and docs/research-blog.md §3.1 for full attribution
  • Version bumped to 1.2.1

GitHub collaboration

PyPI


No code behavior changes. Existing v1.2.0 users have no required upgrade. The release is documentation + attribution only.

v1.2.0 — PyPI release: yt-dlp for cloud agents

Choose a tag to compare

@hamza140202 hamza140202 released this 03 Oct 16:15

🎉 LIVE ON PYPI

Install: pip install agent-video-downloader && avd agent-setup

PyPI: https://pypi.org/project/agent-video-downloader/1.2.0/


What's new in v1.2.0

One-command install for AI agents

  • avd agent-setup — 3-step bootstrap that auto-installs ffmpeg, all Python deps, and clones the XHS-Downloader repo (for Rednote support). Idempotent — safe to re-run anytime. --force flag re-clones XHS-Downloader.
  • avd agent-instructions — prints an 8-step usage guide for AI agents covering install, verify, download, batch, JSON output, verify, jobs/dlq, MCP server, + a copy-paste agent decision tree.
  • Auto-bootstrap on first download — the Rednote extractor auto-clones XHS-Downloader on first call, so users never need to manually install anything.

Real downloads on all 6 platforms (real-world batch test)

Ran avd itself on 24 real public video URLs (4 per platform minimum, large videos first):

Platform URLs attempted Successful Total size Notes
TikTok 4 4 ✅ 16 MB Mix of fan content and official posts
Twitter/X 4 4 ✅ 548 MB 4K videos (largest: 185 MB)
Instagram 4 4 ✅ 28.7 MB Reels
Reddit 4 4 ✅ 468 MB Native v.redd.it uploads (largest: 6-min 372s video)
Douyin 4 4 ✅ 152 MB Mix of fan and official content
Rednote 4 1 ⚠️ 179 KB image only; 3 bot-walled without cookie (documented limitation)
Total 24 21 1.2 GB 87.5% batch success rate

Bugs found and fixed during the batch

  1. Reddit DASH format — some Reddit videos use DASH_<RES>.mp4 (older format, audio embedded) instead of CMAF_<RES>.mp4 (newer, separate audio). Now probes BOTH ladders.
  2. Rednote false-positive — XHS-Downloader subprocess was returning ok=True if any old file matched the glob. Now parses 成功 N 个 from stdout to verify actual success.
  3. Verifier too aggressive — 1 MB minimum was rejecting valid 179 KB JPEGs; moov-atom check fired on streamed MP4s that have moov at END. Fixed: adaptive min (50KB images / 5KB audio / 1MB video), moov only flagged when ffprobe ALSO fails, skip slow integrity_decode for files > 50 MB.
  4. f-string JSON braces — avd agent-instructions was eating the JSON config braces. Escaped as {{ / }}.
  5. pyproject.toml structure — [project.urls] was nested between classifiers and dependencies, breaking setuptools. Now a separate top-level table.

Platform support matrix (verified live 2026-10-03 from datacenter IP)

Platform Primary method Status
TikTok TikWM mirror API ✅ verified
Twitter/X api.fxtwitter.com → video.twimg.com ✅ verified
Reddit rapidsave.com + v.redd.it CMAF/DASH + ffmpeg mux ✅ verified
Instagram yt-dlp + facebookexternalhit/1.1 UA ✅ verified
Douyin api.douyin.wtf public demo (zero-config) ✅ verified
Rednote (XHS) XHS-Downloader (curl_cffi chrome146) ✅ verified

Key discoveries (correct the original research report)

  1. v.redd.it CMAF CDN is NOT 403 from datacenter IPs — direct curl https://v.redd.it/<id>/CMAF_1080.mp4 returns 200. The original Task 2-a report was wrong on this.
  2. facebookexternalhit/1.1 UA bypasses Instagram's datacenter IP wall — Facebook's external link crawler is allow-listed for SSR access. yt-dlp 2026.08.19 + this UA + --no-cookies reliably downloads Instagram reels from datacenter IPs.
  3. api.douyin.wtf public demo works zero-config — operator publishes demo creds at /api/v1/auth/demo. Login gives a 7-day cookie with douyin:read scope. Their CloakBrowser sidecar does all signature heavy-lifting. No Docker needed.
  4. XHS-Downloader uses curl_cffi chrome146 impersonation — fakes Chrome's TLS/JA3 fingerprint, bypasses Xiaohongshu's TLS-fingerprint bot wall. Self-signs X-s/X-t to hit edith.xiaohongshu.com/api/sns/web/v1/feed. Cookie is optional.

Architecture (4 plain-Python agents, no LLM in runtime loop)

  • Orchestrator — owns the extractor registry, drives the fallback chain, persists job state to SQLite
  • Verifier — 6-layer integrity check (size + magic bytes + ffprobe + duration + streams + moov atom)
  • Truth Agent — cross-references downloaded metadata against source platform
  • Tester — runs end-to-end on tests/sample_urls.json

Commands

Command Purpose
avd agent-setup One-command bootstrap — installs ffmpeg, deps, XHS-Downloader
avd agent-instructions Print step-by-step usage guide for AI agents
avd download <url> Download one URL
avd batch <file> Batch download (one URL per line, # comments OK)
avd verify <path> Verify a downloaded file's integrity
avd test --smoke End-to-end self-test on one URL per platform
avd jobs List recent jobs
avd dlq List dead-letter queue entries
avd replay <dlq_id> Replay a failed job
avd mcp Start MCP server (stdio JSON-RPC 2.0)
avd supported Show supported platforms
avd agents Show agent versions

Install

From PyPI (recommended):

pip install agent-video-downloader
avd agent-setup

From source:

git clone https://github.com/hamza140202/agent-video-downloader
cd agent-video-downloader
pip install -e .[dev]
avd agent-setup

Documentation

  • README.md — install + quickstart + commands + changelog
  • CLAUDE.md — project memory (read first if modifying)
  • AGENTS.md — agent contracts + perfection prompting rules
  • TECHSTACK.md — every dependency, every version, why
  • PHASES.md — 7 phases with exit criteria + verification matrix
  • PLAN.md — concrete execution plan with task IDs
  • SKILLS.md — per-agent skill spec sheets
  • docs/research-report.md — live-verified endpoint research per platform
  • docs/architecture.md — system design with ASCII diagram
  • docs/endpoint-matrix.md — living endpoint table
  • docs/research-blog.md — full research blog documenting the project journey
  • CONTRIBUTORS.md — research-grade attribution

Run avd agent-instructions for the 8-step agent usage guide.

Test results

  • 30/30 unit tests pass
  • 21/24 real-world videos downloaded (1.2 GB) via avd itself — 87.5% batch success rate
  • Truth Agent returned verified verdict for a Twitter video tweet
  • pip install agent-video-downloader && avd agent-setup && avd test --smoke verified end-to-end from PyPI

License

MIT. See LICENSE.

Author

  • PyPI: ansaribilal1402
  • GitHub: Bilal140202 / hamza140202