You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
PaddleOCR‑VL backend (SigLIP + Ernie 0.9B with FlashAttention) is now selectable alongside DeepSeek‑OCR. Documentation covers model switching and architecture/memory differences.
Model-aware prompts with bilingual Markdown feedback when requests omit placeholders.
Lazy loading: the server defers weight mmap until the first request, reducing startup time.
Server/CLI sampling controls: --do-sample, --temperature, --top-p, --top-k, --repetition-penalty, and --no-repeat-ngram-size are now recognized across both entry points.
Improvements
Common utilities (token sampling, embedding gathers, transformer KV cache) moved into deepseek-ocr-core, trimming duplication between DeepSeek and Paddle crates.
CLI/Server prompt builders choose the correct format per model, improving output quality without manual tweaks.
Bug Fixes
ModelScope provider now respects arbitrary repo IDs and exact file paths, fixing Paddle asset downloads that previously fetched the wrong config.json.
Requests without images no longer throw transport errors; both sync and streaming responses return a structured bilingual warning instead.
Prompt/image mismatches surface as normal assistant replies instead of opaque “prompt formatting failed” errors, keeping clients compatible with standard OpenAI flows.