This Open-Weights TTS Beat ElevenLabs... Then I Read the License
BBetter Stack
Computing/SoftwareInternet Technology
Transcript
00:00:00An OpenWeight's text-to-speech model just beat 11 labs on artificial analysis.
00:00:05This is Breeze.
00:00:07It's a 3GB download.
00:00:09It runs on your MacBook, and if you actually want to use what it generates for work,
00:00:13you're probably not even allowed to.
00:00:15It's not that the output is that good.
00:00:17Well, it's alright.
00:00:18It's pretty good.
00:00:19But the actual license is the problem.
00:00:21So I want to show you what Breeze can do first, then how we can actually use this.
00:00:30This is Breeze TTS 2.
00:00:33It first dropped on August 25th from a company called Breeze Blue.
00:00:37Now, the benchmark is cool, but how we create a voice is so much better.
00:00:41Most open text-to-speech models nowadays have voice cloning.
00:00:45You give them around 10 seconds of reference audio, usually with an exact transcript,
00:00:50and they try to copy that voice.
00:00:52It works, but we obviously need a voice from the start.
00:00:56Breeze doesn't need one.
00:00:58You just describe the voice you want to actually use.
00:01:01So it could be something along the lines of
00:01:03a warm, thoughtful, wise old man with a calm, reflective delivery.
00:01:08Think Morgan Freeman.
00:01:09And it generates that voice from the description.
00:01:12There's no reference recording.
00:01:14Nowadays, this is called voice design,
00:01:16and we're starting to see more and more open-weight voice models
00:01:19with increasingly good capabilities with voice design.
00:01:23And this fills a real gap in OpenTTS.
00:01:26Kokoru gives you fixed presets.
00:01:29Chatterbox needs referenced audio.
00:01:31Breeze can build the voice from a sentence.
00:01:34So let's actually hear it.
00:01:35If you enjoy coding tools that speed up your workflow, be sure to subscribe.
00:01:39We have videos coming out all the time.
00:01:41Getting this going is mostly done right in my terminal.
00:01:45I just made a virtual environment, and then I pip installed MLX audio.
00:01:49This is running on my Mac, which is why I needed MLX here.
00:01:53If you are on a Linux or have an NVIDIA GPU,
00:01:56then the setup is a bit different, but it's just as straightforward.
00:02:00So I made the same sentence, running it here in my terminal
00:02:03with two different voice instructions.
00:02:05Let's run the first one.
00:02:06Welcome to the BetterStack channel.
00:02:09For all things tech and AI, have you subscribed yet?
00:02:13Now the exact same sentence again with a little different intonation.
00:02:19Welcome to the BetterStack channel.
00:02:21For all things tech and AI, have you subscribed yet?
00:02:27Those were the same words, completely different people, or AI voice.
00:02:33Although this ranks high, I can still tell it's an AI voice, right?
00:02:37So it's good.
00:02:38It's doing rather decent, and the ability to set a motion is actually there too.
00:02:42Now there's one setting here that is key to all this, called CFG scale.
00:02:47That controls how strongly the model follows your voice description.
00:02:51If you lower it, then the model starts drifting back towards a more generic voice.
00:02:55Four is the number in their docs, which I am also using here for a few voices.
00:03:00But let me slide that down to a one real quick, just to get a pure AI voice.
00:03:05Here we go.
00:03:05Welcome to the BetterStack channel.
00:03:08This voice is a clear AI.
00:03:10You see what I mean there, right?
00:03:11That sounded pure AI.
00:03:12So now you can also put vocal events directly inside the text, which is really cool.
00:03:18I can literally type, I can put parentheses.
00:03:20I could put sigh and close those parentheses.
00:03:23At the end, I could put laughing or crying right in the middle of a sentence at the end, at the beginning.
00:03:28Let's hear it.
00:03:30When I try to tell you a joke, I can't hold myself together.
00:03:35That's part of the text, which is a nice feature.
00:03:38But many open-weight voice models are coming out with this now too.
00:03:42So you be the judge of how good Breeze actually sounded.
00:03:46To AI, is it good?
00:03:47Does it deserve the ranking it actually has?
00:03:50Now, Breeze is built from a stack of parts you might actually recognize.
00:03:54A Quen 3 backbone handles language modeling.
00:03:57There's a T5 Gemma 2 text encoder.
00:03:59And MIMI handles the audio codec.
00:04:01The model is around 3 billion parameters, and it outputs 24 kilohertz mono audio.
00:04:07But there's one detail here I want you to remember.
00:04:09The audio tokenizer is derived from Quen 3 TTS, and Quen 3 TTS is Apache 2.0.
00:04:16So part of this model is built on a real open work.
00:04:19Keep that in mind, because in a minute, that becomes rather ironic.
00:04:23First, though, we need to talk about the claim that made Breeze blow up.
00:04:27Breeze beat 11 labs.
00:04:29Yeah, they beat them.
00:04:30That's technically true, but based on the sound output, I'm not sure how, though.
00:04:36This is also where it gets more complicated than just saying that.
00:04:40Artificial Analysis runs a speech arena based on blind, pairwise human votes.
00:04:45They're an independent benchmarking firm, not Breeze Blue.
00:04:49On their provider voice leaderboard, Breeze sits at 1215.
00:04:53That's the ELO score.
00:04:5411 labs conversational is at 1210.
00:04:57So yeah, technically, they beat them by five points.
00:05:00Breeze beats 11 labs.
00:05:01Okay, but now, if I dig into this even more, problems start to pop up.
00:05:06First up, Breeze has the fewest votes near the top.
00:05:09About 1200.
00:05:1111 labs has more than 4,500 votes.
00:05:14So Breeze has the least settled score on that part of the board.
00:05:17And now new models can get a little bit of a welcome bump.
00:05:21Second up here, Breeze isn't actually number one.
00:05:24Cartesia Sonic 3.6 is at 1282.
00:05:27So saying Breeze beats Frontier Proprietary Systems is true in one sense.
00:05:32It's also pretty selective.
00:05:33But the third issue is key to all this.
00:05:36That leaderboard lets each model use its own voices.
00:05:39Artificial Analysis has another leaderboard where the models are given the same exact prompt.
00:05:45The same test.
00:05:46It's basically just apples to apples.
00:05:49And now Breeze looks very different.
00:05:51In that ranking, it scores 1002.
00:05:54Third among openweight models.
00:05:5616th out of 39 overall.
00:05:59And behind Mistral's, Vox Troll.
00:06:01So Breeze is pretty darn good for competing with a lot of these.
00:06:04But it is not number one.
00:06:05And it seems to be especially good when it gets to choose the voices being compared.
00:06:10That changes how well we should actually be looking at this.
00:06:13But it still wasn't the biggest thing that I found.
00:06:16The biggest issue to this whole thing is the license.
00:06:19And this is the part that a lot of people are just skipping over.
00:06:22Yeah, the code is Apache 2.0.
00:06:24There is no problem there.
00:06:26The weights are not.
00:06:27The weights use something called the Breeze Blue Research and Non-Commercial License.
00:06:33And I want to use their exact language here because this distinction actually matters.
00:06:38What this is going to say, what this actually does says is,
00:06:42this agreement permits research and non-commercial use of the model materials free of charge.
00:06:47It does not grant any commercial rights and it is not an open source license.
00:06:52That's their words.
00:06:53Not an open source license.
00:06:55Okay.
00:06:56Now, sometimes you see non-commercial on a model and it basically just means don't resell our weights.
00:07:02Yeah, okay.
00:07:02That makes sense.
00:07:04This is not what this says.
00:07:05Their definition on commercial purpose includes production use,
00:07:09using it in a product or a service, hosting it behind an API.
00:07:13And then there's the part that really just kicks in here.
00:07:16As if that wasn't already enough,
00:07:17it also includes using the outputs for any internal operations beyond limited evaluation.
00:07:23The outputs.
00:07:24Okay, so what does this actually mean?
00:07:26Not just the model itself, the audio it generates.
00:07:29And this whole thing just keeps going.
00:07:31There's no creator exception, no small business exception.
00:07:35There's no revenue threshold.
00:07:37There's nothing here.
00:07:38A lot of non-commercial licenses leave some gray area.
00:07:42This one goes out of its way to basically block a lot of that.
00:07:45So, let's make this more practical now.
00:07:47When could you use this?
00:07:48Put a generated breeze voice inside a product?
00:07:51Nope.
00:07:52That's not going to happen.
00:07:53Use it in a monetized YouTube video?
00:07:56This isn't monetized.
00:07:58Yeah, no.
00:07:58Not going to happen.
00:07:59In some kind of podcast?
00:08:00Again, not going to happen.
00:08:02So, this blocks pretty much everything.
00:08:04The license specifically names sponsorship and subscription revenue.
00:08:09$34 per million characters once you get paying for this.
00:08:12But even there, the license makes something very clear.
00:08:15Paying for the API does not give you commercial rights to self-hosted models
00:08:19or the outputs from their self-hosted model.
00:08:22And once you see that, the whole setup makes a lot of sense.
00:08:25The open weights are the demo.
00:08:27The API is their whole product.
00:08:30So, how usable are those open weights in the first place?
00:08:33You'd be the judge of that.
00:08:34Now, this was very cool.
00:08:36Voice AI is very cool.
00:08:37But I did a breakdown on Vox CPM2 a little while back.
00:08:41Vox CPM2 was honestly insane.
00:08:43It was very, and I mean very, impressive.
00:08:45So, if I stack Breeze against Vox CPM,
00:08:48Breeze would lose this in a heartbeat.
00:08:50And it all has to do with the whole license.
00:08:52Spectacle.
00:08:53So, if you want an honest, open voice clone voice AI,
00:08:56I'm still sticking with Vox CPM2.
00:08:58As it has all the same features as Breeze,
00:09:00and it might be even better.
00:09:02Should you use Breeze?
00:09:03That's the question.
00:09:04I think I answered that, but let's play around here.
00:09:06For playing around with voice AI, yeah, you could.
00:09:09It's cool to see where voice AI is going.
00:09:12Breeze is pretty good, okay?
00:09:14Testing out models is fun,
00:09:16especially as they get better and better.
00:09:18Where are things going?
00:09:19Voice design from a text description
00:09:21is a genuinely really cool capability.
00:09:23But if you're shipping anything commercial,
00:09:25don't use this.
00:09:26I would use something else,
00:09:27regardless of where Breeze sits on the leaderboard.
00:09:30Kokoro is Apache 2.0.
00:09:32Chatterbox is MIT.
00:09:34Vox CPM, we'll look into it,
00:09:35because that might crush them all.
00:09:36Open weights and open source
00:09:38stopped meaning the same thing a while ago.
00:09:40Breeze makes that difference impossible to ignore.
00:09:43The benchmark tells you what the model can do.
00:09:45The license tells you what you can do.
00:09:48I'm Josh from BetterStack.
00:09:49If you enjoy coding tips and tricks like this,
00:09:51be sure to subscribe.
00:09:53We'll see you in another video.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video