Has Zstandard (ZIP method 93) ever been explored for the OCF container? — some measurements, and open questions #3025
bariskayadelen
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hello — and apologies in advance if this has been discussed before and I missed it (pointers very welcome).
I help maintain a small open-source EPUB tool (ePubLift — AGPL, pure-Rust) that modernizes EPUB 2 files to EPUB 3 and re-encodes their images. While working on it I got curious about the OCF ZIP container itself: it mandates Stored + Deflate, while ZIP has long registered Zstandard as compression method 93 (PKWARE APPNOTE). I couldn’t find prior discussion of Zstd for OCF, so rather than argue for anything, I ran some measurements and would love to hear how the WG and implementers think about it.
To be clear about scope: I understand EPUB 3.4 is effectively frozen, and I’m explicitly not proposing a change to it. I’m also fully aware of the elephant in the room — the installed base. A second compression method is worthless to a publisher until essentially every reading system supports it, and “the publishing industry is a large tanker.” So please read this as “here’s some data, is this interesting to think about for the long term?”, not as a proposal.
What I measured
A corpus of 170 real EPUBs (a personal library; mixed genres and sizes). For each book I re-packed its already-uncompressed entries two ways — Deflate (today’s conformant packaging) and Zstandard (level 19) — and compared sizes. Because already-compressed images and fonts dominate most EPUBs and don’t benefit from any re-compression, the whole-archive difference is small (~3%). So the more meaningful number isolates the text/markup (XHTML, CSS, OPF, NCX, SVG…), which is what Deflate vs Zstd actually acts on:
Text-only (images & fonts excluded), bucketed by raw text size — personal library of 170 mixed EPUBs
The same, on a small public-domain sample anyone can reproduce
So this isn’t just my private files, here are the same measurements on 16 Project Gutenberg titles (spanning a flash-fiction-length piece up to the complete works of Shakespeare). These are predominantly text, so they also show where the shared-dictionary win really lives — large, many-chapter works:
The small/medium Gutenberg books are mostly a single content file, so the shared dictionary has no cross-file redundancy to exploit and correctly falls back to per-entry; the large multi-chapter works (War and Peace, Don Quixote, the Shakespeare collection…) are where it pays off (−15%). To replicate (Gutenberg IDs 11 35 84 100 174 345 996 1260 1342 1400 1661 1952 2554 2600 2701 5200):
Two flavours, because they ask different things of a future spec:
A couple of secondary observations that might matter to implementers:
Where this honestly lands
So I’m not claiming this clears the bar — only that the numbers exist now, and they seem worth a conversation.
Open questions for the group
I’m happy to share the measurement methodology and the (open-source, reproducible) tooling, and to run additional numbers if a particular cut would be useful — e.g. a public-domain corpus (Standard Ebooks / Gutenberg) so anyone can replicate, or specific book profiles you’d find more representative.
Thanks for reading, and again — this is meant as “here’s what I found, what do you all think?”, not a push for any outcome.
All reactions