Skip to content

v0.8.1

Latest

Choose a tag to compare

@github-actions github-actions released this 10 Aug 00:30

0.8.1

  • Added Kimi image description, follow-up questions, OCR with structured coordinates, and AI Agent support through the OpenAI-compatible Kimi API.
  • Improved image-description prompts across supported engines with localized defaults.
  • Improved mathematical image descriptions by converting visible formulas to LaTeX and returning formula-only images as LaTeX formulas.
  • Improved automatic image recognition with concise 20 to 30 word descriptions without Markdown.
  • Improved browsable image-description output by using standard Markdown tables for tabular content and avoiding code fences.
  • Removed model and model-provider names from the follow-up dialog.
  • Added Copy and Close buttons to browsable recognition and follow-up result dialogs.

0.7.0

  • Added the Apple Vision (OCR Server) engine for local-network OCR through the open-source OCR Server iOS app.

0.6.5

  • Improved settings and follow-up dialog layouts, including high-DPI scaling.
  • Improved editing of long custom and automatic-recognition prompts.
  • Improved follow-up questions so initial descriptions and latest answers can be viewed as formatted content.

0.6.4

  • Improved formula rendering in HTML and PaddleOCR results.
  • Fixed AI Agent text input and automatic recognition on English interfaces.
  • Improved English and Simplified Chinese UI text and documentation.

0.6.2

  • Updated the available Gemini models, with Gemini 3.6 Flash as the new default and Gemini 3.5 Flash-Lite as a low-cost option.

0.6.1

  • Removed the separate AI Agent start/stop command from Input Gestures; use the main Vis Aware command in Agent mode.
  • Updated the English and Simplified Chinese documentation for commands, automatic recognition, settings, data handling, and the recommended Gemma 4 setup through Ollama.

0.6.0

  • Added PaddleOCR / PaddleOCR-VL and Google Gemma engines.
  • Added Markdown rendering for supported recognition and image description results.
  • Added follow-up question support for conversational image description engines.
  • Added per-engine enable / disable controls so unused engines can be hidden from normal use and engine cycling.
  • Improved engine settings handling and result presentation reliability.

SHA256:
98b571c3e815a67be79faa97be2a3d38fc8c2eccbfee6ffa5417a99a9be83d2c visAware-0.8.1.nvda-addon