Skip to content

Releases: toniher/substack-saved-mcp

0.3.2

Choose a tag to compare

@toniher toniher released this 14 Sep 21:20

Highlights

  • SSRF hardening on authenticated browser navigation. save_post, unsave_post, and fetch_post_content (reachable from both the CLI and the MCP tools exposed to LLM clients) now reject any non-HTTPS URL via a new _require_https() guard before Playwright's authenticated browser context navigates to it. This closes a path where a caller-supplied URL — including one reachable via prompt injection into an LLM client — could reach file:// paths or internal/cloud-metadata HTTP endpoints (e.g. 169.254.169.254).
  • Session credential file hardening. The storage_state.json session file written during login is now pre-created with owner-only 0o600 permissions before Playwright populates it, closing a brief window where it could be created world/group-readable under a permissive umask. chmod failures on the credentials file and its parent directories are now logged instead of silently swallowed.
  • Dedup cleanup. Extracted _with_playwright() (shared Playwright instance dispatch), _resolve_row() (posts/notes URL-or-ID lookup), and _last_sync_info() (per-entity sync-status query) helpers, plus module-level STEALTH_LAUNCH_ARGS/STEALTH_CONTEXT_KWARGS constants, removing repeated logic across substack_client.py, database.py, and cli.py.
  • Dropped the unused pytest-asyncio dev dependency.

Full changelog

0.3.1...0.3.2

0.3.1

Choose a tag to compare

@toniher toniher released this 22 Aug 16:01

Highlights

  • Reading-progress tracking & read_state filtering. Substack exposes per-post reading progress (read_progress/max_read_progress/is_viewed) on both unified and legacy reader APIs. They are now stored as columns and surfaced via read_state filters (unread/in_progress/finished/started) on list_posts/search_posts and MCP tools, with sorting by read_progress and minutes_remaining (word_count * (1 - max_read_progress)). is_fully_read and minutes_remaining are derived at read time from the high-water mark max_read_progress against a configurable SUBSTACK_SAVED_FULLY_READ_THRESHOLD (default 0.95, env SUBSTACK_SAVED_FULLY_READ_THRESHOLD). get_status now reports posts_unread/posts_in_progress/posts_fully_read/minutes_remaining_total. upsert_post now preserves 0.0 progress values (previously lost to or-coalescing).

  • inspect-network stability fixes. Removes the context.route("**/*") → route.fetch() → route.fulfill() interception that caused cross-tab stalls, CORS breakage, and content-encoding: gzip double-decompression (React feed failed to render). Replaced with a passive context.on("response") listener so every page asset and API call loads natively. Scopes the browser launch with --disable-blink-features=AutomationControlled, a desktop Chrome user-agent and 1280×800 viewport to pass Cloudflare Turnstile / cookie-banner checks.

  • Rate-limit resilience. Raises reader-API max_retries from 3→5 (with Retry-After-aware capped exponential backoff, 2–30 s) across all three cursor-paginated fetchers (legacy posts, unified posts, notes). Adds a 15 s cooldown between posts and notes phases of sync when the posts phase was truncated by rate-limiting, so the notes phase no longer immediately re-trips the same bucket.

  • Sync warning hygiene. Explicit api_context.dispose() in finally blocks for all Playwright request contexts (_fetch_saved_posts_page_impl, _fetch_saved_notes_page_impl, probe_api) suppresses Task was destroyed but it is pending! / TargetClosedError noise on sync completion after a mid-pagination 429.

Full changelog

0.3.0...0.3.1

0.3.0

Choose a tag to compare

@toniher toniher released this 16 Aug 13:23

Highlights

  • Saved posts now use Substack's unified reader API by default. A live parity comparison found it a strict superset of the older posts-only reader API (985 vs 1079 posts on a real account, with the legacy API's count also unstable run-to-run), with saved_at/ordering/wordcount all confirmed at 100%. The legacy API and DOM scraping remain automatic fallbacks; SUBSTACK_SAVED_POSTS_SOURCE (auto/unified/legacy/dom) can force a specific source.
  • Fixed a silent-truncation bug shared by all three cursor-paginated fetchers (posts, unified posts, notes): a persistent 429 mid-pagination used to return a partial list indistinguishable from a complete one. A --force sync now detects this and skips reconciliation for that run instead of risking soft-deleting posts/notes that were merely unreachable, not actually unsaved. Reported via a new partial sync status.
  • New diagnostic commands: probe-api (headless GET of any known API URL) and compare-saved-apis (parity report between the legacy and unified endpoints) for verifying Substack API behavior without guesswork.
  • Fixed inspect-network's response-body capture, which was silently recording zero bodies due to a sync-API deadlock risk; now uses request interception.
  • Fixed a version-string drift: --version and the package's __version__ had been stuck at 0.1.0 since a previous release.

Full changelog

0.2.0...0.3.0

Support Saved Notes

Choose a tag to compare

@toniher toniher released this 16 Aug 11:02

Adds full support for Substack's saved notes alongside posts — sync, search, list, save/unsave, and full-content fetch, backed by a dedicated notes table and FTS5 index. All notes operations use direct authenticated API calls, no browser required. New CLI commands (list-notes, search-notes, note-authors, save-note, unsave-note, get-note) plus matching MCP tools and resource. Also fixes a few latent bugs in the posts code: a crash on URL-changing upserts, filters silently dropped by search's fallback path, and inaccurate sync timestamps.

Initial release

Choose a tag to compare

@toniher toniher released this 05 Aug 09:46

Initial release of Substack saved favourites manager CLI and MCP