diarization support?

#13
by evewashere - opened

we need diarization/speaker detection support.
Is that planned?

Mistral AI_ org

Yes next version!

Yes next version!

🀩🀩

And when is the next version planned???

Also if diarization is planned, you need to provide the word and segment timestamps in the json, else this is will be very messy for sure. It is in the API version so why not in the vLLM version?

Any idea when the next release will be coming?
Community would appreciate it..

@patrickvonplaten – thanks for confirming back in February that diarization is planned for the next version. Eight months on, with several follow-up questions in this thread still unanswered, could you share where things stand?

  1. Is there a target window (month or quarter) for the next open-weights Voxtral Realtime release?
  2. Will diarization be part of the open weights, or will it stay API-only as with Voxtral Mini Transcribe V2?
  3. Will word/segment timestamps be exposed in the vLLM path as well (as @aberthil asked)? Without them, speaker labels are hard to use downstream.

Even a rough "not before X" would help those of us planning self-hosted deployments decide whether to wait or to add a separate diarization pipeline in the meantime.

Thanks for the great work on this model!

Sign up or log in to comment