Skip to content

EN Developers

drew.po28@gmail.com edited this page Jun 20, 2026 · 3 revisions

For developers

Developer-facing internals of pico-spec (not part of the on-screen menu). First topic: the pico-spec-catalog that powers the Web Archives browser.

pico-spec-catalog

A separate repository — drewpo28/pico-spec-catalog — that feeds the device's Web Archives browser (Network → F5 → Web Archives). It hides every per-site difference (HTML scraping, JSON APIs, HTTPS, unzipping) behind one trivial tab-separated line protocol, so the firmware stays thin and new sources are added in the catalog repo — no reflashing.

How it connects to pico-spec

  • Device side: src/HttpCatalogFs.cpp (a RemoteFs implementation) + Config::catalog_host. The Web Archives browser lists and downloads through it, reusing the same file-browser UI as FTP/SFTP.
  • The device does HTTPS itself — TLS 1.2 via mbedTLS (TlsSock) runs on the RP2350; the ESP-01S is just a plain-TCP bridge (the same host-crypto / dumb-ESP split as SSH).
  • catalog_host (stored in wifi.cfg) selects the source — HttpCatalogFs::useStaticTree():
    • empty → the built-in default https://drewpo28.github.io/pico-spec-catalog (CATALOG_DEFAULT_URL) — static tree.
    • a full http(s)://… base URL → static tree at that base.
    • a bare host / host:port → the dynamic /v1 server (see below).

Two serving modes

  1. Serverless static (default, live). A GitHub Action (cron, 04:17 UTC) runs gen_static.py, which pre-renders the whole catalog into a static .tsv tree with mirrored file bytes, and publishes it to GitHub Pages. No always-on server. The crawl runs on GitHub's runners (their IP, a real browser UA), so the device never touches the live archive sites — it only GETs plain static URLs.
  2. Dynamic server (optional). A FastAPI service (docker compose up, port 8080) acts as an always-fresh, cached HTTP proxy speaking /v1/sites, /v1/list, /v1/get. The firmware still speaks /v1 when catalog_host is a bare host:port — point it at the server. (pico-spec's own always-on catalog-server code was removed from the firmware; the catalog repo keeps the server for local/dev use.)

Static tree layout (Pages root)

sites.tsv                 "<id>\t<display>\n"  — one line per source
<site>/_root.tsv          root directory listing of a site
<site>/<slug>.tsv         listing of directory <path>   (slug: "" → _root, '/' → '~')
<site>/files/<slug>/<fn>  mirrored file bytes (the download targets)

Listing line format (TAB-separated)

D <TAB> <name> <TAB> 0      <TAB> <child-slug>   sub-dir → GET <site>/<child-slug>.tsv
F <TAB> <name> <TAB> <size> <TAB> <url>          file   → GET <url>  (relative to the
                                                 Pages root, or absolute if it starts http)

The 4th locator column is what lets a static client resolve a download with no server. HttpCatalogFs mirrors gen_static.py's slug() exactly to build the .tsv URLs, reads the locator from the F line for get(), pulls each .tsv in 16 KB Range chunks (so each TLS read stays small), and caches it to /tmp/.catv_*.tsv on SD so a later download doesn't re-fetch over HTTPS.

Sources (adapters — app/adapters/)

id source how
vtrd vtrd.in HTML scrape — no API, and it 403s non-browser UAs, so the crawl must run server-side
sc Spectrum Computing listing built from the ZXDB MySQL dump; files served from spectrumcomputing.co.uk (device TLS handles its cert via mbedTLS SHA384_C)
zxart zxart.ee JSON API (Games + Demoscene)

The build workflow defaults to SITES="vtrd sc zxart", MAX_FILES=400, MAX_DEPTH=4 (workflow_dispatch lets you override). Add a new archive: implement Adapter.list() / Adapter.fetch() and register it in app/adapters/__init__.py — the firmware needs no change. gen_static.py reuses the same adapters as the dynamic server, so there is no second scraper to maintain.

Why a catalog at all (not on-device)

  • vtrd.in has no API (plain HTML) and 403s bots — scraping + a browser UA belong on a server, not in firmware.
  • Upstream downloads are HTTPS and often zipped; building the tree server-side and mirroring pre-unzipped .trd/.tap means the device just GETs a plain static URL instead of fighting a 403 + archive unpack.
  • Listings are pre-rendered, so the device is fast and the archives aren't hammered.

💡 On hardware, Network → HTTP test (curl) can GET any of these static URLs to validate TLS-over-ESP and inspect the raw .tsv.

Links

Clone this wiki locally