v0.11.0c
v0.11.0c
Patch release focused on Android Whisper ASR GPU split execution and repeated-utterance latency.
Changed
- Added native compiled-model caching for Whisper ASR smoke execution.
- The first utterance compiles the encoder/decoder models.
- Later utterances in the same app process reuse the compiled models instead of recompiling every run.
- ASR JSON/status output now reports
compiledModelCacheasmissorhit.
- Updated the Unity Android AAR with the cached Whisper ASR path.
- Updated ASR smoke tooling to support repeated benchmark runs through
-BenchmarkRuns. - Improved Whisper f32 encoder companion resolution so
whisper_tiny_30s_f32_encoder.tfliteis preferred while the legacy alias is still supported. - Restored the patch-based Unity AAR build flow so the LiteRT-LM submodule can be patched and rebuilt from the Unity repository.
Verified
- Built the patched Android AAR successfully.
- Built the Android ASR smoke APK successfully.
- Ran
whisper_tiny_30s_f32.tflite+whisper_tiny_30s_f32_encoder.tfliteon the physical Qualcomm Android device withGPU_FP16. - 10-run ASR result:
- run 1:
compiledModelCache=miss - runs 2-10:
compiledModelCache=hit - backend:
GPU_ENCODER_CPU_DECODER - average elapsed:
2.372s - average compile:
0.153s - average encode:
0.097s - average decode:
1.959s
- run 1:
Notes
- Whisper ASR GPU remains split execution: GPU encoder, CPU decoder.
whisper_tiny_30s_i8.tfliteremains the smallest CPU default for Korean/English ASR.whisper_tiny_30s_f32.tflitepluswhisper_tiny_30s_f32_encoder.tfliteis the recommended path for validating Whisper GPU ASR.