Skip to content

Releases: IstiN/flutter_local_models

qwen3-8b-4bit @ main

Choose a tag to compare

Automated model bundle release for qwen3-8b-4bit.

Description: General-purpose local Qwen text model for chat, automation, and coding.

Use cases

  • Local assistant chat
  • Coding and automation
  • Tool-enabled Flutter app experiments

Limitations

  • Quality and speed depend on Apple Silicon memory pressure
  • Tool calling requires a compatible runtime parser

Release metadata

  • Runtime: mlx_lm
  • Tasks: chat, code
  • Source: mlx-community/Qwen3-8B-4bit@main
  • Archive size: 4623841280 bytes
  • Languages: multilingual
  • Default parameters: {"max_tokens": 512, "temperature": 0.7, "top_p": 0.95}

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

qwen2-audio-7b-instruct-4bit @ main

Choose a tag to compare

Automated model bundle release for qwen2-audio-7b-instruct-4bit.

Description: Audio-first local Qwen model that accepts spoken input and returns text output.

Use cases

  • Direct audio-question answering
  • Voice assistant prototypes
  • Audio understanding experiments

Limitations

  • Audio chat support depends on model template and mlx runtime coverage
  • May be slower than ASR plus text LLM pipeline

Release metadata

  • Runtime: mlx_audio
  • Tasks: chat, audio_input
  • Source: mlx-community/Qwen2-Audio-7B-Instruct-4bit@main
  • Archive size: 6574714880 bytes
  • Languages: multilingual

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

qwen3-tts-12hz-1.7b-customvoice-4bit @ main

Choose a tag to compare

Automated model bundle release for qwen3-tts-12hz-1.7b-customvoice-4bit.

Description: Higher-quality Qwen3 TTS CustomVoice model with built-in multilingual speaker presets.

Use cases

  • Stable preset voices for assistant playback
  • Fast voice UX prototyping
  • Multilingual local TTS experiments

Limitations

  • Speaker presets work best in each speaker native language
  • Optional instruction quality affects emotion and prosody

Release metadata

  • Runtime: mlx_audio
  • Tasks: text_to_speech, audio_output
  • Source: mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-4bit@main
  • Archive size: 2312130560 bytes
  • Languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian
  • Output: {"media_type": "audio/wav", "extension": "wav", "supports_inline_playback": true}
  • Default parameters: {"audio_format": "wav", "join_audio": true, "lang_code": "auto", "voice": "Ryan", "speed": 1.0}
  • Voices: Vivian, Serena, Uncle Fu, Dylan, Eric, Ryan

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

qwen3-tts-12hz-0.6b-customvoice-4bit @ main

Choose a tag to compare

Automated model bundle release for qwen3-tts-12hz-0.6b-customvoice-4bit.

Description: Smaller Qwen3 TTS CustomVoice model with built-in multilingual speaker presets.

Use cases

  • Stable preset voices for assistant playback
  • Fast voice UX prototyping
  • Multilingual local TTS experiments

Limitations

  • Speaker presets work best in each speaker native language
  • Optional instruction quality affects emotion and prosody

Release metadata

  • Runtime: mlx_audio
  • Tasks: text_to_speech, audio_output
  • Source: mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-4bit@main
  • Archive size: 1693675520 bytes
  • Languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian
  • Output: {"media_type": "audio/wav", "extension": "wav", "supports_inline_playback": true}
  • Default parameters: {"audio_format": "wav", "join_audio": true, "lang_code": "auto", "voice": "Ryan", "speed": 1.0}
  • Voices: Vivian, Serena, Uncle Fu, Dylan, Eric, Ryan

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

z-image-turbo-mflux-4bit @ main

Choose a tag to compare

Automated model bundle release for z-image-turbo-mflux-4bit.

Description: Fast Z-Image Turbo mflux 4bit model for low-latency local text-to-image generation.

Use cases

  • Fast local text-to-image generation
  • Photorealistic prompt previews
  • Offline creative workflows on Apple Silicon

Limitations

  • License metadata should be reviewed before redistribution/use in products
  • Very low step counts can trade detail for latency

Release metadata

  • Runtime: mflux
  • Tasks: image_generation
  • Source: filipstrand/Z-Image-Turbo-mflux-4bit@main
  • Archive size: 5911244800 bytes
  • Languages: English prompts
  • Output: {"media_type": "image/png", "extension": "png", "default_size": "1024x1024"}
  • Default parameters: {"width": 1024, "height": 1024, "steps": 9, "guidance": 3.5}

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

whisper-tiny-asr-4bit @ main

Choose a tag to compare

Automated model bundle release for whisper-tiny-asr-4bit.

Description: Tiny multilingual Whisper MLX-Audio speech-to-text model for fast Russian and English dictation tests.

Use cases

  • Fast voice input for voice-to-voice pipeline tests
  • Russian and English ASR smoke tests
  • Offline dictation prototypes

Limitations

  • Lowest accuracy Whisper tier
  • No audio generation; pair with a TTS model for voice output

Release metadata

  • Runtime: mlx_audio
  • Tasks: speech_to_text, audio_input
  • Source: mlx-community/whisper-tiny-asr-4bit@main
  • Archive size: 26490880 bytes
  • Languages: English, Russian, multilingual
  • Output: {"media_type": "text/plain", "format": "text"}
  • Default parameters: {"task": "transcribe", "language": "auto", "temperature": 0.0, "word_timestamps": false}

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

whisper-small-asr-4bit @ main

Choose a tag to compare

Automated model bundle release for whisper-small-asr-4bit.

Description: Compact multilingual Whisper MLX-Audio model for better Russian and English transcription under 150 MB.

Use cases

  • Better-quality voice input for local assistants
  • Russian and English ASR validation
  • Offline transcription and translation experiments

Limitations

  • Slower than tiny/base variants
  • No audio generation

Release metadata

  • Runtime: mlx_audio
  • Tasks: speech_to_text, audio_input
  • Source: mlx-community/whisper-small-asr-4bit@main
  • Archive size: 143616000 bytes
  • Languages: English, Russian, multilingual
  • Output: {"media_type": "text/plain", "format": "text"}
  • Default parameters: {"task": "transcribe", "language": "auto", "temperature": 0.0, "word_timestamps": false}

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

whisper-base-asr-4bit @ main

Choose a tag to compare

Automated model bundle release for whisper-base-asr-4bit.

Description: Small multilingual Whisper MLX-Audio model balancing speed and Russian/English transcription quality.

Use cases

  • Russian and English speech-to-text
  • Voice assistant ASR stage
  • Low-memory offline transcription

Limitations

  • Lower accuracy than Whisper Small/Large variants
  • No audio generation

Release metadata

  • Runtime: mlx_audio
  • Tasks: speech_to_text, audio_input
  • Source: mlx-community/whisper-base-asr-4bit@main
  • Archive size: 46694400 bytes
  • Languages: English, Russian, multilingual
  • Output: {"media_type": "text/plain", "format": "text"}
  • Default parameters: {"task": "transcribe", "language": "auto", "temperature": 0.0, "word_timestamps": false}

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

voxtral-4b-tts-2603-4bit @ main

Choose a tag to compare

Automated model bundle release for voxtral-4b-tts-2603-4bit.

Description: Multilingual Voxtral TTS model with 20 built-in voice presets and faster-than-real-time MLX output.

Use cases

  • Stable preset TTS voices for voice-to-voice pipeline output
  • Multilingual assistant speech playback
  • Comparing Qwen3-TTS and Voxtral voice quality

Limitations

  • Non-commercial source license
  • Text-to-speech only; pair with ASR or STS for full voice workflows

Release metadata

  • Runtime: mlx_audio
  • Tasks: text_to_speech, audio_output
  • Source: mlx-community/Voxtral-4B-TTS-2603-mlx-4bit@main
  • Archive size: 2542929920 bytes
  • Languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Arabic, Hindi
  • Output: {"media_type": "audio/wav", "extension": "wav", "sample_rate_hz": 24000, "supports_inline_playback": true}
  • Default parameters: {"voice": "casual_male", "audio_format": "wav", "sample_rate": 24000}
  • Voices: Casual Male, Casual Female, Cheerful Female, Neutral Male, Neutral Female, French Male

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.

voxcpm2-4bit @ main

Choose a tag to compare

@github-actions github-actions released this 10 May 15:22

Automated model bundle release for voxcpm2-4bit.

Description: Multilingual MLX text-to-speech and voice cloning model for local assistant voices.

Use cases

  • Spoken assistant responses
  • Voice cloning with short reference audio and transcript
  • Local TTS playback in Flutter chat
  • Voice UX prototyping

Limitations

  • Voice controls vary by model family
  • Generated speech quality depends on prompt and runtime support
  • Use a clean 3-10 second reference clip and transcript for stable voice cloning
  • Requires a GitHub mlx-audio build because PyPI mlx-audio 0.4.3 does not include VoxCPM2 yet

Release metadata

  • Runtime: mlx_audio
  • Tasks: text_to_speech, audio_output
  • Source: mlx-community/VoxCPM2-4bit@main
  • Archive size: 2302095360 bytes
  • Languages: multilingual
  • Output: {"media_type": "audio/wav", "extension": "wav", "supports_inline_playback": true}
  • Default parameters: {"audio_format": "wav", "join_audio": true, "lang_code": "en", "cfg_scale": 2.0, "ddpm_steps": 7, "max_tokens": 2000}
  • Voices: Default VoxCPM2

This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.