By Jason Snell
OpenAI silences familiar voices

I got a weird email the other day from OpenAI that ended up causing a great deal of work for me:
As part of our continuous upgrade process, we are deprecating the following legacy audio models: tts-1, tts-hd, gpt-4o-mini-tts. Access to these models will be shut off on January 6, 2027.
I use OpenAI’s text-to-speech voice for the Six Colors Audio Newsletter. Ironically, I chose one of OpenAI’s least advanced voices for this brand-new feature of this website because it was one of the only modern text-to-speech models that didn’t occasionally omit, rewrite, or hallucinate text. I need the computer voice to speak all of our words, and our words only.
This is a bummer, and not just because I have only had six months with this feature before having to rebuild it! I was inspired to build it in the first place by the app (and, full disclosure, former sponsor) Listen Later, which is a read-it-later service that turns your saved articles into a podcast read by… you guessed it… a now-deprecated OpenAI TTS voice.
Not too long after I got the email from OpenAI, I heard from Yalım of Listen Later about this. Unsurprisingly, as someone who has built his entire business on this particular set of voices, he was a bit upset! And not just for the technical challenge he now faces.
It’s the human cost on his customers. Because, funny thing, our human brains are great at finding patterns in static and personifying things that aren’t people. Yalım’s customers have spent years listening to their chosen voice. Their brains have made connections to the sound, rhythm, and delivery pattern of those particular voices. Now those voices are being silenced.
(There’s also an accessibility angle here: As you might expect, Listen Later has a lot of visually impaired users. I’m not sure OpenAI has thought through the ramifications of what this change will mean for those users’ lives, but then again, OpenAI strikes me as a company that doesn’t think a lot of things through.)
I can’t speak to Yalım’s users, but obviously I’m also concerned about this for the Six Colors members who subscribe to the Audio Newsletter. (If you’re someone who likes this site and prefers to listen rather than read, I highly recommend becoming a member and subscribing—you’ll get the stories of the day in a podcast, chaptered and fancy like.) Unfortunately, I don’t expect OpenAI to budge on this, so I’ve spent the last week investigating other options, both local and API-based, to replace my six-month-old version.
Much to my surprise, the winner of my bake-off was the brand-new ElevenLabs v4 engine, which is remarkably good. Yes, I’m going to be spending dollars a month instead of pennies, but the results are very good. Even if the voices are not, alas, our old friends. What I have had to do is create a workflow that tests every audio file that comes out of the API to make sure it’s saying the words it’s supposed to be saying, and not hallucinating. Yes, that means articles take a ridiculous round trip from text-to-speech-back-to-text. But it means that I can use more sophisticated models while guarding against AI omissions or hallucinations. So far, so good.
If you’d like to sample the newsletter or just compare the two sets of voices, good news! I’m currently running both models simultaneously as a test. Here’s a recent example of the current OpenAI voice style, and here’s the beta version of the new ElevenLabs alternative.
I hope people like the new voices, and that Listen Later users don’t suffer too much from the discomfort that change can bring.
If you appreciate articles like this one, support us by becoming a Six Colors subscriber. Subscribers get access to an exclusive podcast, members-only stories, and a special community.