Gemini 3.8 text-to-speech

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/

Comments

rcr-antiSep 23, 2026, 6:22 PM
Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have, consumer, prosumer, cloud. Scroll to the end of every release, including this one, and you'll see different availabilities. The fun part is the models don't even have the same capabilities across platforms! Omni Flash, last I tried and read the docs, is video and text out on consumer and prosumer but video out only on GCP. So if your org disables consumer and prosumer, like mine, it's a coin flip whether you can use the fancy new models or what they can do.
buredorannaSep 23, 2026, 8:27 PM
> Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have...

just remember: despite they would have you believe they are are a united front... they are better thought of as a "loose federation of warring tribes".

BarbingSep 23, 2026, 10:10 PM
I hear much of what happens there is from PMs who want promotions, which makes me wonder if a goal (like interorg alignment) could be met by freezing their promotions until they Google Meet together & hash it out.

I’m reducing into silliness I’m sure, I don’t have any breadth or depth of knowledge here.

Edit: wonder why Apple is ostensibly different. MS seems similar, and don’t know enough about AWS—maybe I’ve seen complaints about them being disjointed, but not as much as Google & Microsoft.

podocarpSep 24, 2026, 7:48 AM
Microsoft is definitely not like that. There are a bazillion versions of copilot. Copilot pc. Copilot vscode. Copilot on Windows. Wtf is going on.

Also famously, the million types of context menus in Windows.

BarbingSep 24, 2026, 3:02 PM
Clarifying ambiguous comment (sorry):

Google and Microsoft are in the disjointed camp. Apple is not. AWS, I do not know.

wamattSep 23, 2026, 9:54 PM
>they are better thought of as a "loose federation of warring tribes".

Googly way to put it! Nice. Might steal that :)

busssardSep 24, 2026, 8:51 AM
most large corporations can be seen like this. Thats why stakeholder management is a full time job in those places.
dekhnSep 24, 2026, 7:48 PM
they even have monkey knife fights (AKA, "resource distribution and trading")
hervalSep 23, 2026, 10:59 PM
united by a Performance Cycle
xatttSep 23, 2026, 8:27 PM
An education account I have lists 3.1 Pro, and 3.6 Thinking and Flash as the available models in the app.

My personal account, 3.6 Flash Lite and 3.5 Thinking.

Meanwhile, I can go hog wild and drain my bank account on GCP. I don’t though, because my family has to eat.

MelatonicSep 23, 2026, 7:58 PM
Seriously. It's very annoying. Add to that the corporate Google Workspace also getting models later
deepfriedbitsSep 23, 2026, 8:27 PM
It is and frankly, it's always been a Google weakness.
dansquizsoftSep 24, 2026, 1:15 AM
I mean, at this point, they are so far behind the frontier, why even bother trying to work with them at all...
simonwSep 23, 2026, 4:19 PM
> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

kingstnapSep 23, 2026, 5:48 PM
Yesterday night I was doing a project with QwenTTS 1.7B.

After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

So yeah the cat is out of the bag for sure.

MacNCheese23Sep 23, 2026, 7:15 PM
Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.

It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.

Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D

NewJazzSep 24, 2026, 5:31 AM
Did you make her breakfast?
Xmd5aSep 25, 2026, 8:02 AM
He cooked eggs for the handsome actor.
yieldcrvSep 23, 2026, 6:15 PM
Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices
pixl97Sep 23, 2026, 7:11 PM
Very few people explore the options they have and tend to stick with the first thing that works.
spuzSep 23, 2026, 9:29 PM
Arguably, no AI voice should sound like a human voice:

https://youtu.be/M-IVVJkZnuo?t=236

yieldcrvSep 24, 2026, 4:47 AM
uninspired

so one of my interests is reducing the payload size for video games

the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism

good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster

although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors

everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences

this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand

peddling-brinkSep 24, 2026, 2:20 PM
> this will vastly supplant the assumed and uninspired “tricking humans” use case

No que por los dos?

Don’t worry, we will be able to destroy careers and scam people at the same time. We don’t have to pick and choose!

yieldcrvSep 24, 2026, 2:39 PM
I can play devil’s advocate too

but let’s not pretend solo developers were ever going to have a voice actor suite and hire that talent en masse

the transactions were never going to happen

and now the outcome will be better than the studios that are making those transactions

peddling-brinkSep 24, 2026, 9:54 PM
It’s genuinely good for some people, I get it. And it’s very cool that some indie games get to be a little better.

But I think you’re dancing on the grave of an entire industry, and every person that’s going to lose their life savings due to this.

yieldcrvSep 25, 2026, 1:54 AM
the current AAA studios will keep that going. they just wont be the future AAA studios

the market doesn’t want 15 year lead times and overly expensive and delayed games that don't experiment on anything

every friction plaguing the industry is solved by distributed indie developers being able to make richer experiences faster and cheaper

rpastuszakSep 23, 2026, 9:06 PM
Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure / approach.
stavrosSep 23, 2026, 9:45 PM
Same, I'd love a link.
kingstnapSep 24, 2026, 5:55 PM
I'll see if I can polish stuff to not be garbage some time. But the general idea was:

1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.

2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).

3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.

4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.

It was honestly pretty vibe coding friendly.

lukanSep 25, 2026, 1:06 PM
(Was dead for some reason unknown to me, vouched for it.)
MulticompSep 23, 2026, 4:21 PM
They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'

and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

gruezSep 23, 2026, 5:14 PM
>and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.

throwa356262Sep 23, 2026, 6:08 PM
This has been possible for quite a long time (probably 1-2 years). There are multiple open models that can do this quite well. Recent example from my YT feed:

https://m.youtube.com/watch?v=WENMgQE9tws

weird-eye-issueSep 24, 2026, 6:49 AM
> I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
schainksSep 24, 2026, 8:24 AM
What's the over/under that Android will roll out a spam filter feature that flags AI voice calls that sound like loved ones?
CROON_tvSep 25, 2026, 11:07 AM
[dead]
miltonlostSep 23, 2026, 4:36 PM
[flagged]
imjonseSep 23, 2026, 4:48 PM
voice cloning is a tool, it is not necessarily evil, even though the scenarios it can be used for nefarious purposes outnumber the legitimate ones.
bakiesSep 23, 2026, 4:59 PM
is the legit ones just like... putting carrie fisher in star wars?
gegtikSep 23, 2026, 5:10 PM
usually people bring up anyone who is losing physical control of their voice faculties, so they can have a synthetic voice that matches their natural voice
jolanSep 23, 2026, 5:15 PM
As an example of this, Sarah Langs uses a synthesized voice due to the progression of her ALS symptoms.

https://en.wikipedia.org/wiki/Sarah_Langs

simonwSep 23, 2026, 5:20 PM
The iPhone has a built in voice cloner hidden in the accessibility settings for exactly this use-case: creating a backup of your voice in case you need it in the future.
lynx97Sep 24, 2026, 6:12 AM
Yeah, but its only available for english IIRC.
gruezSep 23, 2026, 5:16 PM
I mean yeah? If google offers a cloud nmap tool, should everyone get in a tizzy about how google is "evil", even though it saves baddies maybe 5 minutes of work?
kmoserSep 23, 2026, 5:17 PM
Google granted themselves the authority be evil back in 2018. https://en.wikipedia.org/wiki/Don%27t_be_evil
hbnSep 23, 2026, 6:59 PM
It's still in their code of conduct, as the wikipedia article you linked mentioned

https://abc.xyz/investor/board-and-governance/google-code-of...

ctrl/cmd+f "evil"

I don't know why this argument is brought up all the time anyway, it literally means nothing. They can name themselves "Don't Be Evil Inc" and continue to do evil stuff cause evil isn't an objective measure. If squeezing juice out of puppies made money, any business can just say it's "not evil."

And if you think they're evil why would you trust them to follow their own guideline of not doing evil? An evil corp would be more likely to just hide behind that phrase, not quietly remove it as some subtle hint that they want to be openly and proudly evil all of a sudden.

thangalinSep 23, 2026, 4:18 PM
Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:

https://www.youtube.com/watch?v=WAeHgE94rVo

No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.

[1]: https://deepmind.google/models/gemma/gemma-4/

[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design

RGS1811Sep 23, 2026, 9:01 PM
I've been working on a similar project all year and as a tip, you should try Fish Audio or Higgs as a replacement for Qwen3. Both yield much better prosody and are much easier to listen to for long runs.
thangalinSep 23, 2026, 9:58 PM
> Fish Audio or Higgs

I wasn't able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?

MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.

RGS1811Sep 23, 2026, 10:02 PM
For the voice design, these don’t support it, but for the final render, they’re much better. So your pipeline could for example generate voices with one tool and render with another.
loremmSep 23, 2026, 4:25 PM
It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.

I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion

eruSep 23, 2026, 6:54 PM
Awesome! I had been meaning to build something like this for a while now, but never got around to it.

Is it possible to annotate your text with extra 'stage directions' that influence how the book is read out?

thangalinSep 23, 2026, 7:06 PM
> annotate your text with extra 'stage directions'

Good idea, not something I've considered yet. Wouldn't take much to add it since there's already a feature for selecting a quotation and assigning it an intonation. Same infrastructure could be reused to select arbitrary text and assign stage directions.

eruSep 24, 2026, 8:32 AM
Nice! One suggestion: don't just make this apply to text, but also to a specific cursor position (you could say a zero-length text selection where start=end). That's natural if you eg want to say that between one sentence and the next, or one word and the next, you want to insert the sound of a train arriving or a sigh.
MulticompSep 23, 2026, 4:20 PM
<grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>

The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.

Jordan-117Sep 23, 2026, 5:23 PM
Their first sentence literally tells you it's a video of their app? It's not a mystery-meat link.
talon8635Sep 23, 2026, 4:27 PM
[flagged]
idiotsecantSep 23, 2026, 7:09 PM
I guess I should feel bad about wanting to listen to audiobooks of novels that don't have audio while I drive
MulticompSep 23, 2026, 4:24 PM
I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.

Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.

So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.

exhilarationSep 23, 2026, 4:48 PM
You might find this interesting, it seems to be exactly what you need to make an audio drama: https://github.com/Finrandojin/alexandria-audiobook

Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS

I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/

I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!

My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...

throwa356262Sep 23, 2026, 6:14 PM
Then I think what this guy is doing with local Star Trek voice cloning is right up your alley:

https://m.youtube.com/watch?v=jDudeaWppSE

ghostbrainalphaSep 23, 2026, 4:39 PM
Original series or Next Generation?
k12sosseSep 23, 2026, 5:10 PM
Lower decks obviously
hungryhobbitSep 23, 2026, 5:06 PM
Please don't say Nu Trek.
seemazeSep 23, 2026, 4:58 PM
My primary use case for TTS is converting written content (blogs, articles, etc.) in to clips I can listen to on the go.

Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..

exhilarationSep 23, 2026, 7:41 PM
Browser extension? Try these: http://www.paper2audio.com/ https://freevoicereader.com/ https://chromewebstore.google.com/detail/sza-text-to-speech/... https://chromewebstore.google.com/detail/flow-tts/hknidankih...

I found all these on this subreddit: https://www.reddit.com/r/TextToSpeech/

I've also see comments like "Microsoft Edge's read-aloud feature is amazing for TTS" but I haven't tried it myself.

Personally I use Kokoro with a python front-end on my Macbook, I linked to it in this comment: https://news.ycombinator.com/item?id=49818923 it outputs MP3, so I just copy those to my phone and listen as audiobooks.

goldenjmSep 23, 2026, 10:59 PM
I'm the Paper2Audio founder. I hope you enjoy using us for text to speech. Our browser extension adds your open tabs to your listening queue.

Please let me know if you have any questions or feedback.

kyrraSep 23, 2026, 4:59 PM
If you use Chrome on Android, there is a "Listen to this page" option in the menu.

https://support.google.com/chrome/answer/14768725?hl=en

e12eSep 23, 2026, 8:29 PM
Doesn't actually read the page, though. Or not what you see in the browser. For example right now, when I navigate to "reply" it doesn't read the comment I'm replying to - it's reading the login banner you get visiting a reply link without being logged in.
unglaublichSep 23, 2026, 5:46 PM
I vibed a wikipodcast app that has a wikipedia dump, gemma llm for search and summarization and a tts model for audio gen so I could just ask about topics while on the go, and the app would just go off and tell me stuff out of the Wikipedia database.

It was especially nice during a bike trip along the Rhine, I listened to a lot of the history of the industrial area and its cities.

le-markSep 24, 2026, 11:44 AM
That’s really great; an ad hoc tour guide!
janalsncmSep 23, 2026, 5:38 PM
For on the go, I’ve been using ElevenReader. Technically not free but their free tier has been plenty for me.
jonificoSep 24, 2026, 3:26 AM
Edge has a built-in function to read websites.
maelitoSep 23, 2026, 4:40 PM
Related, for embedding small models, this lib is incredible.

Having a voice under 1Mo is crazy, even if it sounds robotic.

https://tts.ampixa.com/sanoTTS/

accountrequiredSep 23, 2026, 4:54 PM
"users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created"

How long is this stored? What could go wrong? :P

kmoserSep 23, 2026, 5:19 PM
Does it detect synthesized consent recordings?
pixl97Sep 23, 2026, 7:14 PM
There is no world in which something like that works over a few months.
nater5000Sep 23, 2026, 4:53 PM
It's giving me an error when I try to generate a voice with Voice Design in AI Studio. It also says voice replication isn't available in my region.

Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?

And there's no pricing listed anywhere.

I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this.

ThaxllSep 23, 2026, 5:52 PM
What is the best open model / tool for text to speech running locally?
rhdunnSep 23, 2026, 6:34 PM
It depends on what you are after (quality, legibility, performance, etc.).

If you're after quality then Qwen3 TTS is a very good model esp. if you take some effort to craft a voice file. It is slow, so isn't practical for real-time voices (like assistants). It can also occasionally switch to a different voice to the one provided, so you may want to break up the text being processed.

I've not yet tried other recent/recentish models.

If you are after performance then two options from older models are:

1. flite with a HTS (Hidden Markov Model) voice like cmu_us_rms (male) or cmu_us_slt (female);

2. espeak/espeak-ng with an MBROLA (an Overlapped Add model) voice (mb-us1, mb-de5-en, etc.).

Alternatively, you could try using Qwen3 TTS or over voice changing model with the CMU Arctic (http://www.festvox.org/cmu_arctic/) voice data which includes audio for the rms and slt voices among others.

If you're feeling adventurous you could also try fine tuning one of the TTS models on that data to create a custom voice, though the data is likely to be in the training data for the voices, so using an audio sample may be sufficient depending on the TTS model.

VariousProgramsSep 23, 2026, 6:20 PM
I've got the best results with BreezeTTS, Higgs Audio v3, and Fish Audio S2 Pro. audio.cpp (https://github.com/0xShug0/audio.cpp) is an easy way to run a lot of different models.

Different models have different strengths. If you throw an entire ebook at a model you're going to get a different result than if you craft a perfect 10 second sentence with a model that supports voice direction and emotion tags, so you should try a bunch depending on your use case.

freedombenSep 23, 2026, 5:55 PM
I've been making audiobooks out of text files that I have lying around, and kokoro TTS has been phenomenal
x3haloedSep 23, 2026, 5:55 PM
I love PocketTTS. It's stupid fast and the voice quality is decent. But it's not the highest quality.
simonwSep 23, 2026, 5:19 PM
I vibe coded a playground UI for trying this out. The conversation mode is neat, and it's very expensive - most of my experiments have cost less than a cent.

https://tools.simonwillison.net/gemini-tts-playground#compos...

simonwSep 23, 2026, 9:12 PM
Example (topical, pelican themed) audio clip here: https://simonwillison.net/2026/Sep/23/gemini-tts-playground/
eisSep 23, 2026, 6:13 PM
Nice but in general I wouldn't put my API key into some third party website, no matter if it claims to not store it. It's not something personal, you are a bit of a celeb here so I wouldn't think you'd save the keys, it's just good general data hygiene. Especially not if a leaked API key can rack up thousands of dollars in fees quickly.
simonwSep 23, 2026, 6:37 PM
Sure. Feel free to copy the code and run it yourself instead.
mohsen1Sep 23, 2026, 6:36 PM
Why tho? AI Studio has this model available and I don't need to paste the API key anywhere

https://aistudio.google.com/generate-speech?model=gemini-3.8...

simonwSep 23, 2026, 6:37 PM
Because building my own is a better way to understand the capabilities of the model and how to use it - and to verify that it can be used via CORS.
CyberDildonicsSep 24, 2026, 11:17 AM
Did you 'build' it or did you vibe code it?
simonwSep 24, 2026, 2:38 PM
Bit of both. I define "vibe coding" as building without looking at the code at all. For this one I was reading the code to understand how it works, and I then followed up several times to adjust how it was working.

Full session here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

iAMkenoughSep 23, 2026, 5:22 PM
Very expensive?
simonlosersonSep 23, 2026, 6:28 PM
he's vibe writing
xnxSep 23, 2026, 4:32 PM
Would be great if this would power the Google Books app feature. The voice system there is pretty out of date.
laweijfmvoSep 23, 2026, 4:39 PM
the ratio of new voice models i see on hackernews to the number actually deployed in any product i use is approximately infinity.
mamudoSep 23, 2026, 4:38 PM
Yes, I am quite disappointed by seeing all this cool AI stuff and yet the same Play Books. Come on, it is the best place to apply AI, in my opinion.
112233Sep 23, 2026, 4:23 PM
"Super tinny monotone robotic voice" does not sound neither tinny nor monotone. Compared to what TTS from 90s sounded like. Or even how actors impersonated robots in movies. Has the model been eating too much hype DJs?
burkamanSep 23, 2026, 4:54 PM
None of these examples are really what the prompt asked for. It's just like image models, once you get over how unbelievable it is that a computer produced this you realize the result isn't actually what you want.
hatingisokSep 23, 2026, 7:32 PM
My ROFLcopter goes: SOISOISOISOISOISOISOISOISOISOISOISOISOISOISOISOI
ramesh31Sep 23, 2026, 11:15 PM
Happier times...
burkamanSep 23, 2026, 4:48 PM
Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good, and usually not particularly close to the prompt. In basically all of these examples some core part of the prompt is completely ignored.
kanbankarenSep 23, 2026, 6:49 PM
Probably.

I listened to some of the voices. The male voices are believable while all the female voices sound the same and artificial. For some reason, it also reminds me of the voices in Toy Story movies.

Bias in the training data?

avazhiSep 23, 2026, 4:58 PM
> Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good,

Um, what?

burkamanSep 23, 2026, 5:07 PM
Technologically incredible as in "I cannot believe it's possible for a computer to do this" and not very good as in "these examples are not what was prompted and I can't think of a use case where these would be acceptable".

Imagine someone showing you that they've trained their dog to hold a paintbrush and paint. There would be no contradiction between "this is incredible" and "these paintings suck".

muzaniSep 24, 2026, 6:59 AM
I've used TTS in videos to replace the speech of the person talking so that we could type up whatever we need them to say in editing.

They sounded exactly like that person... but it's like they're angry at me.

Great toys but not production ready.

sgcSep 23, 2026, 4:29 PM
Sorry if this is in that article, but I am on my phone and can't see it. How much would this cost to batch generate an audiobook? Right now I just listen to things in the 11 labs app which is free, but I would rather just generate audio files.
thevinterSep 23, 2026, 4:32 PM
Roughly 5-10$ for 10h, assuming you few-shot it.

Price per hour:

- 3.8 Flash TTS, standard: $0.81

- 3.8 Flash TTS, batch: $0.41

- 3.8 Flash‑Lite TTS, standard: $0.54

- 3.8 Flash‑Lite TTS, batch: $0.27

sgcSep 23, 2026, 5:35 PM
That is the biggest difference here for me. I have not looked at every solution, but many. They are either much more expensive or garbage quality. For example my next target audiobook is a monster 250k words, 1.5m characters, so about $15 here or $75 using elevenlabs (0.05 per 1k characters). For me that is the difference between I will or I won't use it.
eisSep 23, 2026, 6:21 PM
Until the end of the year, then double that. And that's only the audi output though the text input shouldn't cost much in comparison.

Official pricing can be seen here: https://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-fla...

andrewstuartSep 23, 2026, 4:35 PM
I’ve never found a TYS that does convincing British accents.

They all sound like Americans putting in their best fake British accent.

nitroedgeSep 23, 2026, 5:04 PM
Couldn't see this in the article, does it support API calls for one-shot conversation type responses like ElevenLabs offers?

$0.50 per hour pricing could last a long time with back and forth conversation use.

LarsDu88Sep 23, 2026, 5:30 PM
I've been trying to track the SOTA for years switching between wavenet on Gcloud, to Azure, to ElevenLabs, and now Fish.audio. This is damn good
ttulSep 24, 2026, 4:35 AM
I gave 3.8 a whirl today, replacing Eleven v3 TTS in an internal application that uses TTS to provide a listening function. The Google model produces extremely expressive output. To my ears, it’s on par with Eleven v3, which was already amazing.

Nice work. It’s awesome to have these capabilities so close at hand and so trivially easy to integrate with.

m3kw9Sep 23, 2026, 4:57 PM
Still sounds AI, you can tell they exaggerate all the tone and trailing "high scoring expressive sounds" like your job depends on it.
qlteSep 23, 2026, 8:30 PM
Yes, I find nearly every "SOTA" voice model I try intolerable to listen to because of the fake exaggerated expression/emotion. It's actively distracting because it pulls focus to emphasize randomly. ChatGPT Voice models are so insufferable to put up with for a conversation longer than 45 seconds.

All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.

nmstokerSep 24, 2026, 12:52 AM
Yes, I wish there was more focus on correct pronunciation over emotion. You need a mechanism to control/guide the voice which gets under balance right between not having to specify everything and still letting you fix certain cases (where you know a certain sense of a word is meant)
m3kw9Sep 27, 2026, 1:50 PM
They trained the female voice like it will be used for sex chats. Real life females don't talk like they are flirting with you, at least in my own experience.
breakyerselfSep 24, 2026, 1:53 PM
I remember reading years ago that Google was going to start doing voice transcription on device, but to this day it my Internet connection is cutting out it ruins voice transcription for me.
sharktheoneSep 24, 2026, 12:58 AM
Interesting to see that they are publishing a new tts model. I remember that they refused to release one of them a few years ago because they were too afraid of abusing it. Now they just release it without any much thought lol
yipinwongSep 23, 2026, 7:26 PM
Google is spreading too thin, as gemini isn't really that intelligent.

They are creating gemini SOTA (not really any more), flash versions, text-to-speech, video (omni), etc.

I can see they want to create an ecosystem, but I see no focus in any one area.

MelatonicSep 23, 2026, 8:03 PM
Honestly I think their approach is the long term most useful one. Being SOTA is probably very costly and difficult. Optimising all the smaller stuff that's real world useful for people seems like it would have better long term use.

On top of that I would imagine the research side of Google might disk over things from one modality being useful to another - something like an audio optimising or memory optimising for text to speech could maybe also be useful for translation or world models. Etc etc

yipinwongSep 23, 2026, 9:43 PM
I can see where you are coming from. Vendor lock-in.

Basically spread many AI's everywhere, get people develop/user their ecosystem, and lock them in eventually.

But I find their models' intelligence lacking still.

I don't ever use their built-in gemini features (i have paid gmail) because they don't work well.

e.g. I ask gemini to format my Google docs per Google's material design spec with spacing, etc. It does a real bad job. Many times it does it line by line, and when I finaly get it do it for the whole doc, it does it sloppy, and extremely slow (takes 5 minutes for 10 page doc)

MelatonicSep 23, 2026, 11:24 PM
Vendor lock in sucks - I was talking about the research side.

And yeah integration with Google Workspace is still surprisingly not as good as you would expect.

mewse-hnSep 23, 2026, 8:59 PM
The gemma 4 models are pretty great
yipinwongSep 23, 2026, 9:40 PM
not as good as other chinese openweight models.

At best, good, but not great.

MelatonicSep 23, 2026, 11:24 PM
They've also been out for awhile. And scale super well to tiny sizes
rhdunnSep 24, 2026, 7:00 AM
They have thousands of engineers, so it is very likely that different teams are working on each of those with the relevant domain knowledge.
dangoodmanUTSep 23, 2026, 9:03 PM
In their demo for "Monologue" (2nd video over, after the medeival video game example), it is clearly ignoring the vocal cues like `<chuckles>` and `<laughing>`...
tantalorSep 23, 2026, 8:23 PM
Would pay any amount of money for a zombo.com voice.

The original is the best: https://youtu.be/qxWwEPeUuAg

AyanamiKaineSep 23, 2026, 7:06 PM
Still, all voices sound like they are missing something only real human speech can sound like. But many people will not notice the difference between AI and normal voices.
pixl97Sep 23, 2026, 7:13 PM
So you're saying they are up to support tech afterb5 hours on the same call... the soul has left their body.
perrohunterSep 23, 2026, 4:22 PM
Gemini 3.8 "Flash" says hello
cainxinthSep 23, 2026, 5:17 PM
They keep announcing new 3.8 variants. I wonder why they still haven't updated 3.1 Pro yet.
MelatonicSep 23, 2026, 8:05 PM
Probably $$$ and limit to internal compute. Rumour is they even internal google employees have trouble getting compute for AI stuff. Guessing they're pretty slammed infrastructure wise with adding AI answers to default Google searches
SXXSep 23, 2026, 9:46 PM
IMHO its obviously: they cant compete with SotA models in benchmaxxing. They can and do compete on price though.
dainiusseSep 23, 2026, 5:48 PM
Where is "pro"?
kooldeep7Sep 24, 2026, 4:31 AM
Set up billing before trying out anything? Not for me yet.
OutOfHereSep 23, 2026, 5:18 PM
Pricing isn't noted.
fullstackwifeSep 23, 2026, 4:59 PM
It would be nice to have sound effect generation (use case: games)
GrimblewaldSep 27, 2026, 12:35 AM
I dont know why anyone even pays attention to google anymore given what is on offer from qwen. Much older models are superior in every way to google offerings often months if not years in advance. US AI is dead imo, wouldnt surprise me at all, given the lag in capabilities, if google simply distills qwen models.
melin2024Sep 24, 2026, 7:19 AM
As a pro user, can i use this by api?
drewbittSep 23, 2026, 4:44 PM
Great price at least until December 31 too.
sharktheoneSep 24, 2026, 12:57 AM
Why does "gemini 3.8" make me not be too interested in their tts model?
rryanSep 24, 2026, 1:12 AM
HN is not your therapist, but possibly your decision making is driven by branding instead of judging technology on its merits? Easy mistake to make ;)
talon8635Sep 23, 2026, 4:25 PM
Great. Now in additional to AI email responses I will get AIs impersonating my contacts on the phone too. Lovely.
dadoumSep 23, 2026, 5:24 PM
here is a competitor, if someone wants to compare https://gradium.ai/
newhotelownerSep 23, 2026, 8:22 PM
Where can I get text to speech audio easily?
WarmWashSep 23, 2026, 6:10 PM
I hope this spills over into more voices for android/android auto default assistant voices. The current choices are all so meh
UmYeahNoSep 23, 2026, 7:27 PM
HN: Lament devs losing jobs to AI

Also HN: Fuck those voice actors and their careers.

NeoByteSep 24, 2026, 11:00 AM
[flagged]
xmorseSep 23, 2026, 4:59 PM
[flagged]
droidjjSep 23, 2026, 5:00 PM
It’s just a woman’s voice…
hmokiguessSep 23, 2026, 5:01 PM
That argument says as much about the training as it says about your sexuality, you do realize that right.
xmorseSep 23, 2026, 5:03 PM
you can't use this model if it randomly moans between phrases
konartSep 23, 2026, 5:35 PM
Either you have posted link to a wrong video or we have a very different experience when it comes to moaning.

Onomatopoeia? Sure it is there, and some fillers (or whatever you call those little sounds). But moans?

barrellSep 23, 2026, 5:45 PM
The first <sigh> does sound a lot like a moan. OP linked to the timestamp so I missed it when it first played. I was also confused but on second playback I heard the first <sigh> and also thought wtf.

I would not want that in my product.

codygmanSep 24, 2026, 1:23 AM
Same... That was definitely a moan.