Pianissimo-sv โ Core ML
A Core ML conversion of KlangAI/pianissimo-sv, Klang AI's Swedish speech-to-text model fine-tuned from NVIDIA Parakeet TDT 0.6B v3. It is packaged for FluidAudio and uses the same component contract as FluidInference/parakeet-tdt-0.6b-v3-coreml, with a fixed 15 s window. It is used for on-device dictation in Aloud.
Attribution
- Model: Pianissimo by Klang AI, CC BY 4.0.
- Base model: NVIDIA Parakeet TDT 0.6B v3, CC BY 4.0.
- Conversion pipeline: adapted from FluidInference's mobius parakeet-redux scripts.
Changes from the original: the weights were converted to Core ML, and the encoder weights were compressed with
8-bit k-means palettization. Preprocessor.mlmodelc and parakeet_vocab.json are the unchanged v3 files, because
Pianissimo keeps v3's mel front-end and tokenizer.
Files
| File | Notes |
|---|---|
Preprocessor.mlmodelc |
Mel front-end (from v3, unchanged) |
Encoder.mlmodelc |
FastConformer encoder, 8-bit LUT, 15 s window โ 188 frames |
Decoder.mlmodelc |
RNNT prediction network |
JointDecisionv3.mlmodelc |
Single-step joint + TDT duration + top-K 64 |
parakeet_vocab.json |
8192-token vocabulary (from v3, unchanged) |
Notes
- Pianissimo uses local attention (256/256 frames). Inside the fixed 15 s window, which produces 188 encoder frames, local attention is exactly equivalent to full attention. The encoder is therefore exported with full relative-position attention, and the output is bit-identical to the original on 15 s inputs.
- On an M1 Pro GPU, 257 short Swedish clips took 0.28 s per clip (p50).
- 6-bit palettization, as used for stock v3, noticeably degrades this fine-tune. Use 8-bit.
- Downloads last month
- 39