`omniocr`

Python packge for using OmniOcr: https://omniocr.ai

pip install omniocr

Usage

Get your API key from: https://omniocr.ai/

Then you can start to OCR documents with:

export OMNIOCR_API_KEY=<OMNIOCR_API_KEY>

omniocr examples/resources/sample.pdf \
    --model=lightonocr-2-1b \
    --format=markdown \
    --pages "1-3" > output.md

Alternatively, you can run it programmatically:

from omniocr import OmniOcr


client = OmniOcr()

document = client.process(
    "examples/resources/sample.pdf",
    model="lightonocr-2-1b",
    format="markdown",
    pages="1-3"
)

print(document)

Formats

There are two types of formats that omniocr supports:

markdown conversion -- this is the simplest, the document is just converted to markdown, typically with placeholders for images
block-based output -- if you need bounding boxes for where the text comes from, you should use a model that supports bounding box outputs

Name		Name	Last commit message	Last commit date
Latest commit History 12 Commits
examples		examples
omniocr		omniocr
test		test
.gitignore		.gitignore
README.md		README.md
pyproject.toml		pyproject.toml

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

`omniocr`

Usage

Formats

Supported Models

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

omniocr

Usage

Formats

Supported Models

About

Topics

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

`omniocr`

Packages