doc-html-translate 26.0718.0252
What's new
- Comic archives - new input format (CBZ / CBR / CB7 / CBT). Every page becomes a translatable image: the OCR overlay turns speech bubbles into text your browser can translate. CBR/CB7 use 7-Zip on the desktop app.
- Large PDFs stop freezing. One-pass image extraction takes a heavy scanned PDF from ~2h20m down to ~6s; the extension renders in chunks and extracts page images lazily.
- Better OCR on scans. DPI-aware upscaling for real page scans, plate text now fits its box and re-fits after translation, and recognition runs across a process pool (~6x faster on multi-page books).
- Standalone images fixed. TIFF transcodes to PNG (multi-page -> one page per frame); PDFs pick the right raster instead of a thumbnail or a duplicate.
- Smarter text files. TXT now decodes UTF-16/BOM and legacy Cyrillic code pages (cp1251 / koi8-r / cp866 / iso-8859-5) instead of mojibake.
- Honest handling. Unsupported binaries are refused with a clear message instead of being turned into a garbage document; page and OCR counts no longer over-report.
- Reader polish. Scanned pages fill the window, larger reader text (~28px), a tab icon on results, calmer progress, and the GUI can re-open a finished result.
- Under the hood. pdf.js 4->6, tesseract.js 5->7, marked 12->18, pdfcpu 0.13, goldmark 1.8.
Downloads
doc-html-translate-setup-26.0718.0252.exe- universal installer (x86/x64, per-user, no admin)
SHA256:BBE8BE7C0321AF6332B370B7E8D8A8EB7CF969CD7120B67C4A4D1A4F2BAA287Ddoc-html-translate-26.0718.0252-windows-x64.exe- command-line tool
SHA256:E52EF730C995A1AFC97A90AB0F5EA1C15FDE1E1F96DE652D8B5F73F7205E2793doc-html-ui-26.0718.0252-windows-x64.exe- GUI desktop app
SHA256:BD8D53DEC0AFEC81B67CCC04DDC63B56104DFD32A4D2C2644644DB976F051AF5doc-html-translate-26.0718.0252-windows-x64.zip- full archive (both binaries + LICENSE + README)
SHA256:6866B5A13F2023D0797456B0EC03FBA0155772DC0F1C33E38ED94DE8551372B0