Skip to content

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 09 Nov 18:24
· 42 commits to master since this release

New Features

  • PaddleOCR‑VL backend (SigLIP + Ernie 0.9B with FlashAttention) is now selectable alongside DeepSeek‑OCR. Documentation covers model switching and architecture/memory differences.
  • Model-aware prompts with bilingual Markdown feedback when requests omit placeholders.
  • Lazy loading: the server defers weight mmap until the first request, reducing startup time.
  • Server/CLI sampling controls: --do-sample, --temperature, --top-p, --top-k, --repetition-penalty, and --no-repeat-ngram-size are now recognized across both entry points.

Improvements

  • Common utilities (token sampling, embedding gathers, transformer KV cache) moved into deepseek-ocr-core, trimming duplication between DeepSeek and Paddle crates.
  • Documentation clarifies multi-model selection, DeepSeek-only dynamic crop mode, bilingual terminology, and per-flag behavior.
  • CLI/Server prompt builders choose the correct format per model, improving output quality without manual tweaks.

Bug Fixes

  • ModelScope provider now respects arbitrary repo IDs and exact file paths, fixing Paddle asset downloads that previously fetched the wrong config.json.
  • Requests without images no longer throw transport errors; both sync and streaming responses return a structured bilingual warning instead.
  • Prompt/image mismatches surface as normal assistant replies instead of opaque “prompt formatting failed” errors, keeping clients compatible with standard OpenAI flows.