Repository navigation
Releases: codemyriad/cassini-android
Release list
0.1.0
Transcription now tells speakers apart, remembers the people you name, and survives interruptions.
- One action. Transcribe recognizes the words and labels each speaker in a single job. The first time, the app asks once to download the models (about 730 MiB: Parakeet v3, the speaker model and the voice model). Earlier installations download only what is missing.
- Name speakers. Tap a speaker label, or use ⋯ → Speakers…, to give it a name. The name is written into the note's Cassini document; words, timing and audio stay unchanged.
- Remembered voices. When you name a speaker, the app can keep a voice fingerprint on this phone. Later transcriptions apply the name automatically when the match is strong (marked, easy to change) and suggest it with one tap when it is weaker. Not this person undoes a wrong match. Settings → People renames or forgets saved voices. Fingerprints never leave the phone and are never written into shared files.
- Resume processing. If transcription is interrupted, the note keeps its progress. Resume processing continues from where it stopped instead of starting over.
- Simpler note screen. Transcribe again, transcript choice, Save Cassini, Info and Settings are now in the ⋯ menu.
Speaker labels and automatic names are estimates. Overlapping voices are not separated, and short turns can be attributed to the wrong person. Check names against the recording.
Install cassini-android-0.1.0.apk below over an earlier beta to keep notes and downloaded models. Installation needs Android 8+ and 64-bit Android. See the README for memory and storage recommendations.
Validation: 204 JVM tests, the published Cassini format vectors and Android lint. CI runs the emulator checks on Android 8, 10 and 14. Transcription with speakers, naming, voice recognition and resumed processing were exercised on a Pixel during development, including private multi-speaker meeting excerpts. The voice-match thresholds were tuned on a small set of real meetings and may need adjustment.
Native dependency sources and license notices are attached; model weights download separately.
Native runtime: sherpa-onnx 1.13.7 + Nemotron streaming diarization
Native runtime for Cassini Android, not an app release.
sherpa-onnx 1.13.7 with Sortformer speaker diarization (Nemotron-3-Diarization) and the streaming diarization extension, built for arm64-v8a and x86_64 with Android NDK r28c and ONNX Runtime 1.27.1.
This build replaces native-sherpa-onnx-1.13.7-nemotron.1. That earlier asset predates scripts/streaming-diarization.patch. It lacks processStreamingChunk, saveStreamingState and restoreStreamingState, and current Cassini does not compile against it.
Source:
- https://github.com/codemyriad/sherpa-onnx/tree/feat/nemotron3-diarization-android at commit 1f23907a9d345fca044d015ff4f4b3fbd58de9dc: upstream tag v1.13.7 plus two commits adding the Sortformer runtime and its Kotlin/JNI configuration.
streaming-diarization.patch, attached and identical toscripts/streaming-diarization.patchin this repository (SHA-25691b3f2c55e2d5188a71a3602a56f94d3ae128e0b04635838c5f70800412700d5). It adds a stateful Sortformer chunk call through C++ and Kotlin/JNI.
Rebuild with scripts/build-native-aar.sh; see docs/native-runtime.md. The pinned source archives and their hashes are listed in scripts/native-dependencies.json.
SHA-256 of sherpa-onnx-1.13.7-nemotron.aar: d376f7e25707cf6834208db297e5955c8e02117131740df7ff7fd6ca17216068
🤖 Generated with Claude Code
Native runtime: sherpa-onnx 1.13.7 + Nemotron diarization
Superseded by native-sherpa-onnx-1.13.7-nemotron.2. This build predates the streaming diarization patch, and current Cassini does not compile against it.
Native runtime for Cassini Android, not an app release.
sherpa-onnx 1.13.7 with Sortformer speaker diarization (Nemotron-3-Diarization), built for arm64-v8a and x86_64 with Android NDK r28c and ONNX Runtime 1.27.1.
Source: https://github.com/codemyriad/sherpa-onnx/tree/feat/nemotron3-diarization-android (commit 1f23907a9d345fca044d015ff4f4b3fbd58de9dc): upstream tag v1.13.7 plus two commits adding the Sortformer runtime and its Kotlin/JNI configuration. Rebuild with scripts/build-native-aar.sh; see docs/native-runtime.md.
SHA-256 of sherpa-onnx-1.13.7-nemotron.aar: d03a06f8cad8f9d761c97974ee8042d713e166aa1c14df09101303fda7a65b08 (pinned in scripts/setup.sh).
🤖 Generated with Claude Code
0.0.6-beta
You can now add anonymous speaker labels to a completed transcript, including notes you already have. This is an optional pass after transcription; it does not repeat speech recognition or run while recording.
- Tap Identify speakers, then choose automatic detection or a known count (1–8). The separate speaker models download once, with your consent (about 40 MiB).
- The result adds a transcript variant. Your original transcript, audio and imported speaker names stay available under Choose transcript.
- Cancel returns to the library and keeps the original result. The native computation may take time to finish in the background before another transcription can start.
Labels are estimates. Automatic detection can merge or split voices, and mixed audio does not separate people speaking at the same time. Check labels against the recording. In the public four-speaker fixture, automatic clustering produced seven speaker IDs; specifying four produced four. That count check does not measure attribution accuracy. Processing runs on CPU; keep the app open until it finishes.
Install cassini-android-0.0.6-beta.apk below over an earlier beta to keep notes and downloaded models. Installation needs Android 8+ and 64-bit Android. Creating Cassini files needs Android 10+ and a working platform Opus encoder. See the README for memory and storage recommendations.
Validation: 143 JVM tests, all 26 published Cassini format vectors and Android lint. CI runs 31 playback, recording, settings and speaker-consent checks across Android 8, 10 and 14 emulators, without model inference. Native speaker identification and cancellation were checked on an Android 14 emulator with a two-voice fixture, including preservation of words, timing, previous variants and the Opus audio digest. This is a clean test fixture, not a conversational accuracy benchmark. The Pixel was unavailable for this feature's device checks.
Native dependency sources and license notices are attached; model weights download separately.
0.0.5-beta
You can now opt in to transcription while recording, and see draft words sooner during batch transcription.
- Turn on Transcribe while recording in Settings after downloading the selected model. It is off by default. Words appear after each short audio chunk is decoded; the phone may still be catching up when you tap Done.
- Done saves the recording first, then finishes transcription in the same library note. If transcription fails or you leave the recording screen, the audio stays available for another attempt.
- Batch transcription shows draft words after each decoded chunk. The draft is marked while processing and may change as overlaps are joined.
Live transcription uses more battery, and its shorter audio context can reduce accuracy. Use Transcribe again for the usual batch pass. Recording and transcription run in the foreground; keep the app open.
Install cassini-android-0.0.5-beta.apk below over an earlier beta to keep your notes and downloaded models. Installation needs Android 8+ and 64-bit Android (ARM64 phones or x86_64 emulators). Creating Cassini files needs Android 10+ and a working platform Opus encoder. See the README for memory and storage recommendations.
Validation: 127 JVM tests, all 26 published Cassini format vectors, Android lint, and 22 device checks across Android 8, 10 and 14 emulators. CI's emulator checks do not run model inference. Live transcription and saving were also checked with the real INT8 decoder on an Android 14 emulator. Native dependency sources and license notices are attached; model weights are downloaded separately.
0.0.4-beta
You can now read the first words while the rest of the recording is still being transcribed.
- Transcription shows progress, pace and an estimate of the time left. Words appear as each stretch of speech is decoded; they can still change when overlapping chunks are joined.
- Silero detects speech spans using the desktop pipeline’s settings. Long spans are cut at quiet points, with recorded context on both sides of each cut.
- Words from overlapping windows are merged, and padding and quiet-audio filtering follow the desktop word gate. A boundary fix keeps the copy with duration when another copy was clipped to zero length.
The small Silero model (about 630 KB) downloads separately. Existing Parakeet downloads are reused; if the detector is unavailable, the app cuts at quiet points instead. Android still uses stock sherpa-onnx on CPU, rather than the desktop’s modified runtime.
Install cassini-android-0.0.4-beta.apk below over an earlier beta to keep your notes and downloaded models. Installation needs Android 8+ and 64-bit Android (ARM64 phones or x86_64 emulators). Creating Cassini files needs Android 10+ and a working platform Opus encoder. See the README for memory and storage recommendations.
Validation: 101 JVM tests, all 26 published Cassini format vectors, Android lint, and 16 playback/settings/Opus checks on Android 8, 10 and 14 emulators. CI’s emulator checks do not run model inference. Earlier speech-cutting and progress checks on a Pixel 8 and an Android 14 emulator are documented in the device results. Native dependency sources and license notices are attached; model weights are downloaded separately.
0.0.3-beta
This beta fixes two playback problems and adds checks on older Android versions.
- Leaving a note before its first playback no longer makes its audio unplayable.
- Opening a note keeps the keyboard closed. On older Android versions, the search field could take focus and pull the page away from the word being played.
- CI now tests the built APK on Android 8, 10 and 14. All three must pass before a release is published.
Install cassini-android-0.0.3-beta.apk below over an earlier beta to keep your notes and downloaded models. Installation needs Android 8+ and 64-bit Android (ARM64 phones or x86_64 emulators). Creating Cassini files needs Android 10+ and a working platform Opus encoder; some emulator images lack one. See the README for memory and storage recommendations.
Validation: 23 JVM tests, all 26 published Cassini format vectors, Android lint, and 16 playback/settings/Opus checks across the three emulators. These checks do not run model inference; recording and transcription results from earlier device runs are linked in the README. This release has not been retested on a physical phone. Native dependency sources and license notices are attached; model weights are downloaded separately.
0.0.2-beta
This beta fixes what a code review of 0.0.1-beta turned up. The most important one is in the audio itself.
- Recordings lost one sample every 64 ms. Microphone recordings (48 kHz AAC) were decoded with a silent sample every 3072 samples. That glitch went into the transcription and into the Opus audio saved in new documents. Decoding is now gapless. Documents made with 0.0.1-beta keep the audio they were saved with.
- Opening the same file twice no longer duplicates it. A file opened from another app was copied again every time the screen was recreated (a theme or language change was enough), each time with a new note. It now returns to the same note, with its playback position and selected transcript.
- Long transcripts play without stalling. The whole transcript was rebuilt for every spoken word. Now only the highlight moves.
- Cancelling “Import audio” goes back to Notes instead of leaving you inside the last note you opened.
Install cassini-android-0.0.2-beta.apk below over 0.0.1-beta: your notes and downloaded models stay. Android 10+ is recommended; ARM64 phones and x86_64 emulators are packaged. The limits are the same as before: recordings up to 2:59, imports up to 3 minutes, and the app has to stay open while recording and transcribing.
Validation: 23 JVM tests, Android lint, all 26 published Cassini format vectors, and 20 recording/library/interface/portable-document checks, including real Parakeet INT8 transcription. The device checks for this release ran on an Android 14 x86_64 emulator, not on a physical phone. The fixes were written by Claude and reviewed by Codex over several rounds. The native dependency sources and license notices are attached; model weights are downloaded separately.
0.0.1-beta
Record a voice note, then find it again in a searchable library. Transcription runs on the phone with Parakeet v3; words link back to the audio. Save a complete Cassini .opus file to keep both together.
- Install cassini-android-0.0.1-beta.apk below. Android 10+ is recommended; ARM64 phones and x86_64 emulators are packaged.
- Download a transcription model in Settings once: INT8 is about 640 MiB, FP32 about 2.37 GiB. Audio processing is offline afterward.
- This beta handles recordings up to 2:59 and imports up to 3 minutes. Keep the app open while recording and transcribing.
- GitHub Actions builds the release APK with a stable beta signing key. Install future beta updates over it to keep your notes and downloaded models.
Validation: 20 JVM tests, Android lint, all 26 published Cassini format vectors, and all 13 recording/library/interface/portable-document checks on the published release APK, installed as an upgrade on a Pixel 8. No device tests were skipped. The native dependency sources and license notices are attached; model weights are downloaded separately.