Repository navigation
Multimodal EN
🇫🇷 French version: Multimodal
KarenOS lets you send more than text: attachments (images, videos, audio, files) and direct voice input.
In the chat input bar, the paperclip button gives access to four choices:
- Image… — photos and image files.
- Video… — video files.
- Audio… — audio files.
- File… — any other file type.
Selected files appear as thumbnails above the input field (removable chips) before sending. Once the message is sent, they are shown in the message bubble; a right-click lets you reveal the file in Finder.
The mic button starts dictation:
- macOS asks for the Microphone and Speech Recognition permissions (described in
Info.plist). - The app listens and shows the live transcription (a red dot + the recognized text).
- Clicking stop (or at the end of the phrase) inserts the text into the input field: it can be edited before sending.
Technically, dictation uses the macOS Speech framework (SFSpeechRecognizer) in local mode (requiresOnDeviceRecognition = true) when the voice model is installed on the machine.
A language model only understands an image if it is a vision model (Llava / Qwen-VL family…), paired with a *.mmproj file placed next to the *.gguf file.
- When a model loads, the app auto-detects the presence of a
.mmprojin the model's folder. - If present, attached images are sent in the OpenAI vision format:
{"type":"image_url","image_url":{"url":"data:image/png;base64,…"}} - Without a
.mmproj, the image remains an attachment displayed in the conversation; only the text part goes to the model.
The default Store models (Qwen, SmolLM, Phi) are text-only: vision is therefore inactive until a vision model is installed with its
.mmproj.
| Input | Tool | Where |
|---|---|---|
| Text | input field | chat bar |
| Image | paperclip → Image… | chat bar |
| Video | paperclip → Video… | chat bar |
| Audio | paperclip → Audio… | chat bar |
| File | paperclip → File… | chat bar |
| Voice (dictation) | mic button | chat bar |
See also: Inference of a message, Autonomous agents, Limits & roadmap.
- Accueil
- Architecture
- Inférence d'un message
- Modèles et Boutique
- Agents autonomes
- Multimodal
- Compétences
- Persistance
- Code source
- Construction & lancement
- Dépannage
- Limites & évolutions
- Home
- Architecture
- Inference of a message
- Models and Store
- Autonomous agents
- Multimodal
- Skills
- Persistence
- Source code
- Build & launch
- Troubleshooting
- Limits & roadmap
Main repository — created by Martial Zinsou