Skip to content

Multimodal EN

martialzinsou edited this page Sep 21, 2026 · 3 revisions

Multimodal (images, videos, audio, voice)

🇫🇷 French version: Multimodal

KarenOS lets you send more than text: attachments (images, videos, audio, files) and direct voice input.

Attachments

In the chat input bar, the paperclip button gives access to four choices:

  • Image… — photos and image files.
  • Video… — video files.
  • Audio… — audio files.
  • File… — any other file type.

Selected files appear as thumbnails above the input field (removable chips) before sending. Once the message is sent, they are shown in the message bubble; a right-click lets you reveal the file in Finder.

Voice input (talking to the app)

The mic button starts dictation:

  1. macOS asks for the Microphone and Speech Recognition permissions (described in Info.plist).
  2. The app listens and shows the live transcription (a red dot + the recognized text).
  3. Clicking stop (or at the end of the phrase) inserts the text into the input field: it can be edited before sending.

Technically, dictation uses the macOS Speech framework (SFSpeechRecognizer) in local mode (requiresOnDeviceRecognition = true) when the voice model is installed on the machine.

When the model truly « sees » an image

A language model only understands an image if it is a vision model (Llava / Qwen-VL family…), paired with a *.mmproj file placed next to the *.gguf file.

  • When a model loads, the app auto-detects the presence of a .mmproj in the model's folder.
  • If present, attached images are sent in the OpenAI vision format:
    {"type":"image_url","image_url":{"url":"data:image/png;base64,…"}}
    
  • Without a .mmproj, the image remains an attachment displayed in the conversation; only the text part goes to the model.

The default Store models (Qwen, SmolLM, Phi) are text-only: vision is therefore inactive until a vision model is installed with its .mmproj.

Recap

Input Tool Where
Text input field chat bar
Image paperclip → Image… chat bar
Video paperclip → Video… chat bar
Audio paperclip → Audio… chat bar
File paperclip → File… chat bar
Voice (dictation) mic button chat bar

See also: Inference of a message, Autonomous agents, Limits & roadmap.

Clone this wiki locally