feat(voice): match local and macOS voices to the language of the text
Text-to-speech picked one voice regardless of what language a reply was in. A dependency-free language detector (script, marker letters, function words) now decides the language of the whole message once; with the new "Match the voice to the language of the text" setting the local provider switches to a catalog model for that language (Kokoro zh/en and Piper models for 12 languages, downloaded on first use like the existing model) and macOS say switches to an installed voice whose locale matches. The local voice picker lists voices of every installed model, and the settings show which language models are on disk. The Ukrainian Piper medium build is a character-level model that sherpa-onnx turns into noise, so the espeak-based Lada build is used instead. Claude-Session: https://claude.ai/code/session_017TK5JAYDfT3Fotc23UEg98
This commit is contained in:
@@ -11,12 +11,25 @@ live transcript costs O(n^2) work for a result the final decode replaces. The
|
||||
composer shows no text while recording and inserts the full transcript on
|
||||
stop.
|
||||
|
||||
Local TTS (Kokoro via sherpa-onnx OfflineTts) runs in the same worker process
|
||||
and is exposed as `POST /api/dictation/tts/speak` (JSON `{text, speakerId?,
|
||||
speed?, model?}` → WAV bytes; 503 with `reasonCode` while the model is
|
||||
downloading). TTS models live in the same catalog/downloader as STT models
|
||||
(`local/model-catalog.js` `LOCAL_TTS_MODEL_CATALOG`) and are managed by the
|
||||
same status/download/delete routes.
|
||||
Local TTS (Kokoro and Piper/VITS via sherpa-onnx OfflineTts) runs in the same
|
||||
worker process and is exposed as `POST /api/dictation/tts/speak` (JSON
|
||||
`{text, speakerId?, speed?, model?, language?, languageSample?}` → WAV bytes; 503 with
|
||||
`reasonCode` while the model is downloading). TTS models live in the same
|
||||
catalog/downloader as STT models (`local/model-catalog.js`
|
||||
`LOCAL_TTS_MODEL_CATALOG`) and are managed by the same status/download/delete
|
||||
routes.
|
||||
|
||||
Each TTS catalog entry declares the `languages` it speaks. With
|
||||
`language: 'auto'` the service detects the language of `languageSample` — the
|
||||
whole message the chunk belongs to, sent by the client with every chunk — or
|
||||
of `text` when no sample is given
|
||||
(`../tts/language-detect.js`, script plus function-word scoring, no
|
||||
dependencies) and keeps the caller's model when it speaks that language;
|
||||
otherwise it switches to the catalog model for the language, downloading it on
|
||||
first use like any other model, and starts from that model's default speaker
|
||||
(`defaultSpeakerByLanguage`) instead of the caller's speaker id. A language no
|
||||
catalog model covers keeps the caller's model, so text is always spoken. The
|
||||
response carries `X-Speech-Model` and `X-Speech-Language`.
|
||||
|
||||
## Ownership
|
||||
|
||||
|
||||
Reference in New Issue
Block a user