> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
I learned about all of these projects on HN at one point or another.
On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?
defaulting to wasm on ios devices now
IMHO, its at about the same quality level of historic TTS tools.
here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!