Inflect-Micro-v2: complete voice in 9.36M parameters

https://huggingface.co/owensong/Inflect-Micro-v2

Comments

modinfoJul 26, 2026, 5:48 AM
This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd

thanks for shearing!

yjftsjthsd-hJul 26, 2026, 3:21 AM
Couple highlights:

> Complete local text-to-waveform speech synthesis under 10M parameters.

In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

> English only, with one fixed male voice. This is not zero-shot voice cloning.

(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

semiquaverJul 26, 2026, 4:33 PM
When would “text-to-waveform speech synthesis” ever imply speech to text?
yjftsjthsd-hJul 26, 2026, 6:42 PM
The HN title is "Inflect-Micro-v2: complete voice in 9.36M parameters".
NetOpWibbyJul 26, 2026, 7:43 AM
The inflections are weird but this doesn't sound like a robot. Not bad!
billdueberJul 26, 2026, 2:06 PM
I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?
SamPattJul 27, 2026, 1:59 AM
I just built my own voice assistant with my Pebble Time 2 watch and it uses a VPS hosted TTS (Piper) and Hermes agent.

I learned about all of these projects on HN at one point or another.

eightysixfourJul 26, 2026, 3:34 PM
I use STT/TTS to interface with a local LLM for Home Assistant in my house.
_davide_Jul 26, 2026, 6:41 PM
I'm do so as well, i tried qwen3 omni 3 but it was ridiculously stupid, and i ended up with stt thinker and tts. kokoro for now
tmalyJul 26, 2026, 2:01 AM
This is impressive. I wish there were a voice clone option.
fastballJul 26, 2026, 4:19 AM
With so few parameters, I imagine a voice fine-tune might be readily tractable.
K0baltJul 26, 2026, 6:43 PM
How heavy in inference on this? The model would easily fit on many microcontroller modules, I wonder if they could run it?
sudbJul 26, 2026, 4:39 PM
this is extremely encouraging for individuals/small companies being able to train pareto-frontier TTS models (specifically compute required to run vs quality of model output)
da-xJul 26, 2026, 5:40 PM
I think we need more neurons in the human brain than parameters in this model for speech. I wonder what it says about the human brain vs LLM efficiency.
StilesCrisisJul 26, 2026, 12:48 PM
I'd love to hear it but it seems your quota is exhausted.
g58892881Jul 26, 2026, 2:46 PM
StilesCrisisJul 26, 2026, 4:44 PM
Nice! Strangely, "Nano" sounds a lot better than "Micro" to me.

On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?

g58892881Jul 26, 2026, 9:21 PM
also, i double checked, nano is nano and micro is micro. didnt fuck that up
g58892881Jul 26, 2026, 6:18 PM
right. happens on my 13 too. memory leak confirmed, not sure yet what's causing it
g58892881Jul 26, 2026, 9:20 PM
unable to fix it. latest dev ort, dispoing tensors, restarting the session/worker. nothing, crashes every time.

defaulting to wasm on ios devices now

jsomedonJul 26, 2026, 2:55 AM
amazing quality for such small size!
itakeJul 26, 2026, 6:29 AM
Amazing quality for small size, but definitely not that enjoyable to listen to.

IMHO, its at about the same quality level of historic TTS tools.

stavrosJul 26, 2026, 7:18 AM
I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.
itakeJul 26, 2026, 8:17 AM
I compared the macos Samantha just now and I guess the inflect-micro is marginally better...
leobgJul 26, 2026, 7:53 AM
Ivona „Joey“, „Amy“
phoenixrangerJul 26, 2026, 4:22 PM
amazing! was looking for something similar
mcbetzJul 26, 2026, 8:55 AM
Alternative title: Text to speech in 9.36M, English only.
fintunerJul 26, 2026, 4:05 PM
[flagged]
ameliusJul 26, 2026, 5:49 PM
[dead]
zenith605Jul 26, 2026, 4:48 PM
[flagged]
afdsaifdoiJul 26, 2026, 3:37 PM
[dead]