Skip to content

v0.11.0c

Choose a tag to compare

@Leuconoe Leuconoe released this 17 May 00:47

v0.11.0c

Patch release focused on Android Whisper ASR GPU split execution and repeated-utterance latency.

Changed

  • Added native compiled-model caching for Whisper ASR smoke execution.
    • The first utterance compiles the encoder/decoder models.
    • Later utterances in the same app process reuse the compiled models instead of recompiling every run.
    • ASR JSON/status output now reports compiledModelCache as miss or hit.
  • Updated the Unity Android AAR with the cached Whisper ASR path.
  • Updated ASR smoke tooling to support repeated benchmark runs through -BenchmarkRuns.
  • Improved Whisper f32 encoder companion resolution so whisper_tiny_30s_f32_encoder.tflite is preferred while the legacy alias is still supported.
  • Restored the patch-based Unity AAR build flow so the LiteRT-LM submodule can be patched and rebuilt from the Unity repository.

Verified

  • Built the patched Android AAR successfully.
  • Built the Android ASR smoke APK successfully.
  • Ran whisper_tiny_30s_f32.tflite + whisper_tiny_30s_f32_encoder.tflite on the physical Qualcomm Android device with GPU_FP16.
  • 10-run ASR result:
    • run 1: compiledModelCache=miss
    • runs 2-10: compiledModelCache=hit
    • backend: GPU_ENCODER_CPU_DECODER
    • average elapsed: 2.372s
    • average compile: 0.153s
    • average encode: 0.097s
    • average decode: 1.959s

Notes

  • Whisper ASR GPU remains split execution: GPU encoder, CPU decoder.
  • whisper_tiny_30s_i8.tflite remains the smallest CPU default for Korean/English ASR.
  • whisper_tiny_30s_f32.tflite plus whisper_tiny_30s_f32_encoder.tflite is the recommended path for validating Whisper GPU ASR.