Skip to content

Releases: GolyBidoof/mokuro-bridge

mokuro-bridge v0.7.4 - installable from PyPI

Choose a tag to compare

@GolyBidoof GolyBidoof released this 03 Oct 23:35

The 0.7 line in one note: the bridge became an installable package with a check that says when a newer release exists, plus the fixes that followed it. v0.7.0 through v0.7.3 carry tags but never got a release page of their own, so their notes are collected here.

Added

  • pipx install mokuro-bridge (or uv tool install). Two console scripts: mokuro-bridge, mokuro-bridge-ocr. server.py and ocr_folder.py stay as wrappers, so the checkout install is unchanged.
  • --check-update prints the newest published release and the upgrade command for the way this copy was installed (pipx, pip or a git checkout). Exit codes: 0 up to date, 1 update available, 2 the check could not be completed. --version prints the running version.
  • /health gained update_check, latest_version, update_available, update_url and update_error, read from a cache so the request never waits on a socket. The check runs in the background at startup and is refreshed at most every six hours (MOKURO_BRIDGE_UPDATE_TTL_S). A failure is cached too, so an offline machine does not retry in a loop. MOKURO_BRIDGE_UPDATE_CHECK=0 turns it off.
  • python -m mokuro_bridge runs the bridge, for an install whose script directory is not on PATH.
  • mokuro_bridge/update.py, with tests for version comparison, the injected fetcher, TTL caching, and the offline and disabled paths.
  • tests/test_packaging.py holds the packaging contract: pyproject.toml and requirements*.txt cannot drift apart, every package directory is listed for the wheel, the installed-vs-checkout output directory rule is pinned, and the console script targets must resolve.
  • CI: .github/workflows/tests.yml runs the suite on Linux for Python 3.10 to 3.13, plus macOS and Windows on 3.13, and builds the wheel to assert the vendored ./mokuro fork never ships inside it. .github/workflows/publish.yml publishes a vX.Y.Z tag to PyPI with Trusted Publishing.
  • MOKURO_BRIDGE_INGEST_ROOTS: comma-separated absolute paths added to the local-ingest allowlist, which is otherwise the home directory plus the POSIX temp paths. Those resolve to C:\tmp on Windows and never match, so a library on a second drive had no way in. Relative entries are ignored.

Fixed

  • The local-ingest root list was defined twice in config.py, so every MOKURO_BRIDGE_INGEST_ROOTS entry was appended twice. One definition now; the duplicate shipped in 0.7.3.
  • test_metadata_file_is_private is skipped on Windows, which has no POSIX mode bits: st_mode reports 0o666 for any writable file and chmod only toggles read-only. The boundary there is the user profile ACL.
  • A base install can read a generated .mokuro again. v0.7.0 shipped a crash here, found by CI immediately after the first publish: _validate_mokuro_artifact() read the file through _mokuro_submodule("utils").load_json, and with no OCR engine installed the helper raised AttributeError from its None placeholder. The validator now parses with the standard library, since it reads a file the bridge itself generated and has no business depending on PyTorch. _mokuro_submodule() raises an ImportError naming the missing dependency instead.
  • /session/resume answered 500 on every call. _safe_component() was called without being imported. Fixed, and locked by tests/test_endpoints.py.
  • OCR is about 1.8x faster again. v0.6.0 raised num_beams to 4, costing ~43 ms/crop against ~23.5 at one beam (measured on MPS). Back to one beam.
  • A page resent under a new name supersedes the old one instead of sitting beside it. A 249-page volume was once finalized as 498 and shipped that way.
  • A volume wedged by the bridge's own staging can be finalized again. Remote uploads stage in <vol>/_mega_upload and the ingest walk rejects images in subdirectories, so that debris made every later attempt fail too. It is cleared on the next attempt.
  • A volume could never be reused. The collision check ran after the work directory was created, so it was always true and every fresh volume was given a random <title>_<6 hex> suffix. A later reuse_existing start looked the session up by the plain title, found nothing, and created yet another directory, so the volume and its OCR cache were thrown away on every run. The check now happens before the directory is made.
  • Imprint and edition tags no longer split one series across folders. サンプル作品(1) (サンプルコミックス) now derives a single series rather than one per volume.

Changed

  • The recommended OCR engine is now the mokuro fork rather than the stock package, since it adds batched recognition and is worth roughly 1.8x on a measured workload. It installs in one line, pip install "mokuro @ git+https://github.com/GolyBidoof/mokuro", and the bridge detects its batch API by introspection, so there is nothing to configure. Stock mokuro from PyPI still works and remains the fallback in requirements-ocr.txt.
  • The default output directory now depends on how the bridge was installed. A git checkout keeps using <repo>/output; an installed wheel uses ~/mokuro-bridge/output, because the old default resolved to the parent of site-packages and would have put the user's volumes inside the virtualenv, where a pipx upgrade rebuild can strand them. MOKURO_BRIDGE_OUTPUT_DIR still overrides both.
  • The CLI moved from server.py into mokuro_bridge/cli.py. server.py is now a wrapper that also re-exports the ASGI app, so python server.py and uvicorn server:app behave as before. ocr_folder.py moved into the package for the same reason, with a wrapper left behind.
  • ./mokuro is excluded from the distribution explicitly. A wheel shipping a top-level mokuro package would shadow the real OCR engine that users install from PyPI.
  • The fetch proxy serves /_bwdd/<host>/<path> and advertises it as pathUpstream in /health (fetchPathUpstream), for clients that need a fixed origin.
  • MOKURO_BRIDGE_FETCH_UPSTREAM and MOKURO_BRIDGE_FETCH_ALLOWED_HOSTS are documented in the README and .env.example. The allow-list still defaults to *; _target_is_safe() is the gate that rejects loopback, private, link-local, multicast and reserved targets.
  • page_num is optional on /page and /page-local. Clients that omit it are unaffected; it was already documented as advisory.
  • The publish workflow's build job runs on pushes to main as well as tags, so the sdist and wheel are built and twine checked on every commit. Uploading still requires a v* tag or a manual dispatch.
  • The version check uses httpx rather than urllib. httpx is already a base dependency and verifies TLS against certifi, where a python.org macOS build ships no CA store of its own and fails the request with CERTIFICATE_VERIFY_FAILED.
  • README quickstart now leads with pipx and keeps the checkout path beside it, and the update check is documented alongside it (--check-update, /health's update status, and the MOKURO_BRIDGE_UPDATE_* knobs in .env.example).

Notes

  • Nothing is updated automatically. --check-update reports, and prints the command; applying it, and restarting a service, stays your call, because a restart mid-OCR would lose work.
  • The first PyPI release is v0.7.0. v0.6.0 was git-only, so there is no version collision, and the update check compares against the v0.7.0 tag. publish.yml fails the build if the tag and the packaged version disagree.
  • Publishing goes to PyPI through Trusted Publishing, set up once against owner GolyBidoof, repository mokuro-bridge, workflow publish.yml, environment pypi. Before that publisher existed the publish job failed at its OIDC step, while the build job still produced the wheel as an artifact.

mokuro-bridge v0.6.0 - named upload accounts

Choose a tag to compare

@GolyBidoof GolyBidoof released this 25 Sep 02:05

Several accounts per upload provider, and a fix for the keychain lookup that made a re-run of the MEGA wizard appear to do nothing.

Added

  • Named upload accounts. Every provider can now hold more than one account: python server.py --setup-upload mega --name work adds a second MEGA account, and it is addressed as mega:work (upload_method=mega:work, ocr_folder.py --upload-method mega:work, MOKURO_BRIDGE_UPLOAD_DEFAULT=mega:work). A bare provider id still means the default account, so existing clients and configs are untouched.
  • Each account has its own remote root (--root), display label (--label) and credential store: one keychain item per MEGA email, a 0600 credential or token file per Drive/OneDrive account. Non-secret metadata lives in ~/.config/mokuro-bridge/accounts/<provider>__<name>.json (MOKURO_BRIDGE_ACCOUNTS_DIR).
  • python server.py --list-uploads prints every account with its readiness, credential source and remote root; --remove-upload mega:work forgets one.
  • /upload-methods and /health list one entry per account, each with provider, account, its own root and current_folder.
  • mokuro_bridge/accounts.py: the instance registry (id parsing, metadata, legacy default resolution), plus tests for id parsing, account isolation and the keychain cleanup rules.

Fixed

  • The macOS keychain lookup paired the email from one mega.nz item with the password from another, and an account-less find-internet-password kept returning the older item. A stale entry for a previous address therefore shadowed the working one and every upload failed with API call 'us' failed: Server returned error ENOENT even after a successful wizard run. The lookup now reads both halves from a single item, prefers the email recorded for the account, and the wizard removes only the entries it knows are orphaned, never a sibling account's.

Notes

  • accounts_dir() no longer creates its directory on read, so importing the bridge cannot fail on an unwritable $HOME.

mokuro-bridge v0.5.2 - clearer install failures

Choose a tag to compare

@GolyBidoof GolyBidoof released this 24 Sep 00:29

mokuro-bridge v0.5.2 - clearer install failures

Two install failures reported from the field, both now diagnosed in the README and one of them handled by the server itself.

Added

  • server.py prints an actionable message when a base dependency is missing, rather than a bare traceback. It names the interpreter that failed, prints the matching -m pip install line, and points at the usual cause.
  mokuro-bridge cannot start: fastapi is not installed for this Python.

  interpreter: /usr/bin/python3
  install with: /usr/bin/python3 -m pip install -r requirements.txt

  If you already installed the requirements, pip targeted a different
  Python. Prefer `python3 -m pip` over a bare `pip`, and activate the
  virtualenv in every new terminal.

Documented

  • No space left on device while pip builds a wheel is normally $TMPDIR, not the disk. Many distros mount /tmp as a small tmpfs, so df -h / reports a different filesystem and the machine really does have free space. Building unidic-lite, which ships an sdist and no wheel, unpacks around 45 MB of dictionary into it. Fix: mkdir -p ~/.cache/pip-tmp && TMPDIR=~/.cache/pip-tmp pip install -r requirements-ocr.txt.
  • ModuleNotFoundError: No module named 'fastapi' means pip targeted a different Python than the one running server.py. Compare python3 -c "import sys; print(sys.executable)" against python3 -m pip -V, and install with python3 -m pip rather than a bare pip.

Notes

  • Both failures come from the OCR engine's dependency tree. The base install has not contained that since v0.5.1, so installing the bridge on its own builds nothing from source.

mokuro-bridge v0.5.1 - optional OCR engine

Choose a tag to compare

@GolyBidoof GolyBidoof released this 24 Sep 00:27

mokuro-bridge v0.5.1 - the OCR engine is optional now

mokuro depends on PyTorch, so pip install -r requirements.txt used to pull several GB onto machines that only ever wanted page capture, the fetch accelerator or uploads.

The bridge already ran fine without an engine, with /health reporting mokuro_installed: false, so this is a packaging fix rather than a behavioural one.

Changed

  • requirements.txt is the server only now: fastapi, uvicorn, python-multipart, httpx and keyring.
  • mokuro moved to the new requirements-ocr.txt. Install it when you want OCR.
  • A MOKURO_REPO checkout still works and needs neither file.
  • README quickstart and troubleshooting explain that mokuro_installed: false is the expected state on a base install, not a fault to chase.

Added

  • tests/test_imports.py walks the AST of every module and fails if anything outside the base requirements is imported at module level. That is what stops the OCR engine and the cloud SDKs from creeping back into the startup path.
  • tests/test_optional_ocr.py imports the API, the OCR module, the providers and /health with mokuro blocked on the meta path, and checks the fetch accelerator still answers.

Upgrading

Only the install changes. Reinstalling requirements is enough, and nothing needs uninstalling: if you already have mokuro, the bridge keeps using it, whether it came from requirements-ocr.txt, a MOKURO_REPO checkout or your own install.

mokuro-bridge v0.5.0 - page-fetch accelerator

Choose a tag to compare

@GolyBidoof GolyBidoof released this 24 Sep 00:17

mokuro-bridge v0.5.0 - page-fetch accelerator

Chrome allows only 6 concurrent HTTP/1.1 connections per origin, and an origin is scheme + host + port. BookWalker's page CDN is a single host and refuses to negotiate HTTP/2, so a viewer download was pinned to 6 sockets however fast the connection was.

Because the port is part of the origin, the bridge now serves the same small proxy on a range of extra localhost ports. The browser treats each as a fresh origin, so every port is worth 6 more sockets, while the bridge does the fetching under no browser limit at all. With the default 48 ports that is 288 extra sockets for the downloader.

Highlights

  • Multi-port fetch proxy (mokuro_bridge/fetchproxy.py), 48 ports by default. Set MOKURO_BRIDGE_FETCH_PORTS=0 to turn it off.
  • No configuration anywhere. The bridge advertises the ports it bound on /health as fetchProxyPorts, and the userscript picks them up automatically. The startup banner reports the range.
  • File-descriptor headroom. Every proxied page holds a descriptor on each side of the bridge, so 48 ports at 6 sockets is roughly 600 at once. The bridge raises RLIMIT_NOFILE at startup, which covers launchd, run.sh and manual runs alike. macOS gives a launchd agent a soft limit of 256, which previously showed up as OSError: Too many open files on accept() and ECONNRESET in the browser, with nothing pointing at the ulimit.
  • Test suite under tests/, run with python3 -m pytest. 16 checks over port binding, proxy forwarding and streaming, the host allow-list guard, upstream failure handling and the environment defaults. No mokuro, torch or network access needed.

Notes

  • Downloading never depends on the bridge. With it stopped, the userscript falls back to its own page and background-context lanes.
  • UVICORN_RELOAD=1 disables the fetch proxy: the reloader spawns a second process that would contend for the same ports.
  • The proxy refuses any host outside its allow-list, so it cannot be used as an open proxy.

mokuro-bridge v0.4.0

Choose a tag to compare

@GolyBidoof GolyBidoof released this 06 Sep 00:50

mokuro-bridge v0.4.0 - first release

A local OCR bridge for manga pages to reader.mokuro.app. Capture pages from any storefront into a session, OCR them with mokuro, and get the .cbz / .mokuro / .webp trio that reader.mokuro.app reads - kept locally or pushed to your own cloud.

Highlights

  • OCR pages as they arrive in chunked batches (8 pages default); sessions survive restarts; one bridge serves many capture clients in parallel.
  • Four cloud destinations plus local: MEGA, Google Drive, OneDrive and WebDAV (Nextcloud/ownCloud/...), all storing under the mokuro-reader// layout reader.mokuro.app scans. Cloud upload is opt-in and off by default.
  • Works with stock pip install mokuro or the faster GolyBidoof/mokuro fork (batched recognition, auto-detected at runtime).

Uploads

  • Per-request destination: upload_method=local|mega|drive|onedrive|webdav on /session/{id}/finalize.
  • Sticky defaults: the last explicitly-requested method (and local output dir) becomes the default, persisted across restarts.
  • overwrite=fail|skip|overwrite: decide what happens when a file already exists at the destination (uniform across providers).
  • Live progress: real-time NDJSON upload_progress events (bytes / percent / speed) for every remote method, plus a per-file shareable URL on completion.
  • GET /upload-methods: list configured methods and their current folder before uploading.

Operability

  • GET /health: engine and credential status, upload-method registry, and a busy / busy_stage / busy_detail signal so scripts can wait for the bridge to go idle.
  • Clean console logging: in-place OCR/upload progress lines; model-load noise suppressed.
  • Auth wizards: python server.py --setup-upload <mega|drive|onedrive|webdav> installs the provider's dependencies and verifies credentials before storing (keychain / 0600 file / env).

For developers

  • Source split into a clean mokuro_bridge/ package (single-file server.py kept as a thin entrypoint).
  • Everything is env-configurable (MOKURO_BRIDGE_, MEGA_, DRIVE_, ONEDRIVE_, WEBDAV_*).

MIT licensed - see LICENSE.