Skip to content

PDF OCR Action v1

Latest

Choose a tag to compare

@muhac muhac released this 06 Oct 06:27

English | 中文

English

Current version: v1.8.0 · updated 2026-10-06. This page follows the v1 tag and is updated when a minor version adds something; fixes in between move v1 without changing this page.

Turn scanned PDFs into searchable PDFs — text you can select, copy and search — with the OCR engine built into macOS, running on GitHub's macOS runners.

Quick start

jobs:
  ocr:
    runs-on: macos-15
    steps:
      - uses: actions/checkout@v7
      - uses: muhac/pdf-ocr-action@v1
        with:
          input: scans
          output: searchable
          language: eng

What it does

  • Files or folders. One PDF or a folder of PDFs; in a folder, a file that cannot be processed does not stop the others.
  • About 30 languages, including Simplified and Traditional Chinese, Japanese and Korean. Latin text inside a Chinese, Japanese or Korean page is still recognized.
  • Language subfolders. In a folder, PDFs inside a subfolder named after a language code (chi_sim/, chi_tra/, …) are read in that language.
  • Clean text layer. The invisible text is placed over the page without the debug boxes the AppleOCR plugin draws; bookmarks are kept.
  • Leaves existing text alone by default; mode: force runs OCR on every page.
  • Quiet mode keeps file names and OCR output out of the log.

Also included

  • OCR service workflow: takes PDFs from a private repository's inbox folder and commits searchable results back. Works on the default branch or on any branch (so a squash-merged pull request keeps originals out of the default branch), and takes files over 100 MB through releases. Logs never contain file, folder, branch or repository names.
  • Trigger action, muhac/pdf-ocr-action/trigger@v1: starts the service from the private repository.
  • Check action, muhac/pdf-ocr-action/check@v1: compares each batch of results with the originals (pages, bookmarks, text per page) and flags pages that lost text; optionally shows each flagged page to Claude next to its OCR text, using a Claude subscription token, never an API key; Claude says whether OCR captured or missed the text and adds a one-line note. Reports in English or Chinese.

See the README for setup.

Requirements

  • A macOS runner; macos-15 is recommended. On macos-26, Traditional Chinese pages lose much of their text; Simplified Chinese is unaffected. macos-14 is not supported.
  • Nothing to install: OCRmyPDF 17.13.0 and OCRmyPDF-AppleOCR 0.4.0 are pinned and their dependencies frozen by date, so an upstream release cannot silently change results.

Changes by minor version

  • 1.8 Claude compares each flagged page with the text OCR found on it (captured or missed), reviewing pages in batches; max-pages: 0 reviews them all.
  • 1.7 Claude adds a note for each page it reviews; reports in Chinese with language: zh.
  • 1.6 Check action; the service tells the storage repository when results are saved.
  • 1.5 The service works on a chosen branch; per-branch queues; choice of macOS version and recognition mode per run; back to macos-15 by default.
  • 1.4 The plugin's debug boxes are removed from the text layer.
  • 1.3 The service downloads only its working folders; six-hour runs; results kept when the inbox is emptied meanwhile; damaged images no longer fail a whole file.
  • 1.2 Files over 100 MB go through releases.
  • 1.1 Language subfolders.
  • 1.0 OCR action, OCR service, trigger action.

Built on OCRmyPDF and OCRmyPDF-AppleOCR.


中文

当前版本:v1.8.0 · 更新于 2026-10-06。本页跟随 v1 标签,在小版本增加新内容时更新;其间的修复只移动 v1,不改本页。

把扫描版 PDF 变成可搜索的 PDF——文字可以选中、复制、搜索。使用 macOS 自带的文字识别引擎,在 GitHub 的 macOS runner 上运行。

快速开始

jobs:
  ocr:
    runs-on: macos-15
    steps:
      - uses: actions/checkout@v7
      - uses: muhac/pdf-ocr-action@v1
        with:
          input: scans
          output: searchable
          language: chi_sim

功能

  • 文件或文件夹。 处理单个 PDF 或整个文件夹;某个文件失败不影响其他文件。
  • 约 30 种语言,包括简体中文、繁体中文、日文和韩文。中日韩文页面里夹杂的英文和数字也能识别。
  • 语言子文件夹。 放在以语言代码命名的子文件夹(chi_sim/、chi_tra/ 等)里的 PDF 按该语言识别。
  • 干净的文字层。 隐藏文字层去掉了 AppleOCR 插件画的调试红框;书签保留。
  • 默认不动已有文字;mode: force 对每一页重新识别。
  • 安静模式:日志里不出现文件名和识别过程的输出。

同时提供

  • OCR 服务 workflow:从私有仓库的收件文件夹取 PDF,把可搜索的结果提交回去。可以在默认分支或任意分支上工作(用 squash 合并 PR,原件就不会进入默认分支的历史),超过 100 MB 的文件走 Release。日志里不出现文件名、文件夹名、分支名和仓库名。
  • 触发用的 Action,muhac/pdf-ocr-action/trigger@v1:从私有仓库启动服务。
  • 检查用的 Action,muhac/pdf-ocr-action/check@v1:把每批结果和原件对比(页数、书签、逐页文字量),标出可能丢字的页面;可选地用 Claude 订阅 token(不用 API key)把这些页面连同识别文字交给 Claude 对照,判断文字是已识别还是漏识别,并写一句说明。报告可用英文或中文。

配置方法见 README。

运行要求

  • 需要 macOS runner,推荐 macos-15。在 macos-26 上繁体中文页面会丢失大量文字,简体中文不受影响。不支持 macos-14。
  • 不需要预先安装:OCRmyPDF 17.13.0 和 OCRmyPDF-AppleOCR 0.4.0 版本固定,依赖按日期冻结,上游更新不会悄悄改变识别结果。

各小版本的变化

  • 1.8 Claude 把每个被标记的页面和识别出的文字对照,判断已识别还是漏识别,分批审阅;max-pages: 0 审阅全部。
  • 1.7 Claude 为审阅的每一页写一句说明;language: zh 输出中文报告。
  • 1.6 新增检查 Action;服务存完结果后通知文档仓库。
  • 1.5 服务可以在指定分支上工作;按分支排队;每次运行可选 macOS 版本和识别模式;默认改回 macos-15。
  • 1.4 去掉文字层里插件画的调试红框。
  • 1.3 服务只下载工作用的文件夹;单次运行最长 6 小时;识别期间清空收件箱也不丢结果;图片损坏不再导致整个文件失败。
  • 1.2 超过 100 MB 的文件走 Release。
  • 1.1 语言子文件夹。
  • 1.0 OCR Action、OCR 服务、触发用的 Action。

基于 OCRmyPDF 和 OCRmyPDF-AppleOCR。