This Open-Weights TTS Beat ElevenLabs... Then I Read the License

BBetter Stack
Computing/SoftwareInternet Technology

Transcript

00:00:00An OpenWeight's text-to-speech model just beat 11 labs on artificial analysis.
00:00:05This is Breeze.
00:00:07It's a 3GB download.
00:00:09It runs on your MacBook, and if you actually want to use what it generates for work,
00:00:13you're probably not even allowed to.
00:00:15It's not that the output is that good.
00:00:17Well, it's alright.
00:00:18It's pretty good.
00:00:19But the actual license is the problem.
00:00:21So I want to show you what Breeze can do first, then how we can actually use this.
00:00:30This is Breeze TTS 2.
00:00:33It first dropped on August 25th from a company called Breeze Blue.
00:00:37Now, the benchmark is cool, but how we create a voice is so much better.
00:00:41Most open text-to-speech models nowadays have voice cloning.
00:00:45You give them around 10 seconds of reference audio, usually with an exact transcript,
00:00:50and they try to copy that voice.
00:00:52It works, but we obviously need a voice from the start.
00:00:56Breeze doesn't need one.
00:00:58You just describe the voice you want to actually use.
00:01:01So it could be something along the lines of
00:01:03a warm, thoughtful, wise old man with a calm, reflective delivery.
00:01:08Think Morgan Freeman.
00:01:09And it generates that voice from the description.
00:01:12There's no reference recording.
00:01:14Nowadays, this is called voice design,
00:01:16and we're starting to see more and more open-weight voice models
00:01:19with increasingly good capabilities with voice design.
00:01:23And this fills a real gap in OpenTTS.
00:01:26Kokoru gives you fixed presets.
00:01:29Chatterbox needs referenced audio.
00:01:31Breeze can build the voice from a sentence.
00:01:34So let's actually hear it.
00:01:35If you enjoy coding tools that speed up your workflow, be sure to subscribe.
00:01:39We have videos coming out all the time.
00:01:41Getting this going is mostly done right in my terminal.
00:01:45I just made a virtual environment, and then I pip installed MLX audio.
00:01:49This is running on my Mac, which is why I needed MLX here.
00:01:53If you are on a Linux or have an NVIDIA GPU,
00:01:56then the setup is a bit different, but it's just as straightforward.
00:02:00So I made the same sentence, running it here in my terminal
00:02:03with two different voice instructions.
00:02:05Let's run the first one.
00:02:06Welcome to the BetterStack channel.
00:02:09For all things tech and AI, have you subscribed yet?
00:02:13Now the exact same sentence again with a little different intonation.
00:02:19Welcome to the BetterStack channel.
00:02:21For all things tech and AI, have you subscribed yet?
00:02:27Those were the same words, completely different people, or AI voice.
00:02:33Although this ranks high, I can still tell it's an AI voice, right?
00:02:37So it's good.
00:02:38It's doing rather decent, and the ability to set a motion is actually there too.
00:02:42Now there's one setting here that is key to all this, called CFG scale.
00:02:47That controls how strongly the model follows your voice description.
00:02:51If you lower it, then the model starts drifting back towards a more generic voice.
00:02:55Four is the number in their docs, which I am also using here for a few voices.
00:03:00But let me slide that down to a one real quick, just to get a pure AI voice.
00:03:05Here we go.
00:03:05Welcome to the BetterStack channel.
00:03:08This voice is a clear AI.
00:03:10You see what I mean there, right?
00:03:11That sounded pure AI.
00:03:12So now you can also put vocal events directly inside the text, which is really cool.
00:03:18I can literally type, I can put parentheses.
00:03:20I could put sigh and close those parentheses.
00:03:23At the end, I could put laughing or crying right in the middle of a sentence at the end, at the beginning.
00:03:28Let's hear it.
00:03:30When I try to tell you a joke, I can't hold myself together.
00:03:35That's part of the text, which is a nice feature.
00:03:38But many open-weight voice models are coming out with this now too.
00:03:42So you be the judge of how good Breeze actually sounded.
00:03:46To AI, is it good?
00:03:47Does it deserve the ranking it actually has?
00:03:50Now, Breeze is built from a stack of parts you might actually recognize.
00:03:54A Quen 3 backbone handles language modeling.
00:03:57There's a T5 Gemma 2 text encoder.
00:03:59And MIMI handles the audio codec.
00:04:01The model is around 3 billion parameters, and it outputs 24 kilohertz mono audio.
00:04:07But there's one detail here I want you to remember.
00:04:09The audio tokenizer is derived from Quen 3 TTS, and Quen 3 TTS is Apache 2.0.
00:04:16So part of this model is built on a real open work.
00:04:19Keep that in mind, because in a minute, that becomes rather ironic.
00:04:23First, though, we need to talk about the claim that made Breeze blow up.
00:04:27Breeze beat 11 labs.
00:04:29Yeah, they beat them.
00:04:30That's technically true, but based on the sound output, I'm not sure how, though.
00:04:36This is also where it gets more complicated than just saying that.
00:04:40Artificial Analysis runs a speech arena based on blind, pairwise human votes.
00:04:45They're an independent benchmarking firm, not Breeze Blue.
00:04:49On their provider voice leaderboard, Breeze sits at 1215.
00:04:53That's the ELO score.
00:04:5411 labs conversational is at 1210.
00:04:57So yeah, technically, they beat them by five points.
00:05:00Breeze beats 11 labs.
00:05:01Okay, but now, if I dig into this even more, problems start to pop up.
00:05:06First up, Breeze has the fewest votes near the top.
00:05:09About 1200.
00:05:1111 labs has more than 4,500 votes.
00:05:14So Breeze has the least settled score on that part of the board.
00:05:17And now new models can get a little bit of a welcome bump.
00:05:21Second up here, Breeze isn't actually number one.
00:05:24Cartesia Sonic 3.6 is at 1282.
00:05:27So saying Breeze beats Frontier Proprietary Systems is true in one sense.
00:05:32It's also pretty selective.
00:05:33But the third issue is key to all this.
00:05:36That leaderboard lets each model use its own voices.
00:05:39Artificial Analysis has another leaderboard where the models are given the same exact prompt.
00:05:45The same test.
00:05:46It's basically just apples to apples.
00:05:49And now Breeze looks very different.
00:05:51In that ranking, it scores 1002.
00:05:54Third among openweight models.
00:05:5616th out of 39 overall.
00:05:59And behind Mistral's, Vox Troll.
00:06:01So Breeze is pretty darn good for competing with a lot of these.
00:06:04But it is not number one.
00:06:05And it seems to be especially good when it gets to choose the voices being compared.
00:06:10That changes how well we should actually be looking at this.
00:06:13But it still wasn't the biggest thing that I found.
00:06:16The biggest issue to this whole thing is the license.
00:06:19And this is the part that a lot of people are just skipping over.
00:06:22Yeah, the code is Apache 2.0.
00:06:24There is no problem there.
00:06:26The weights are not.
00:06:27The weights use something called the Breeze Blue Research and Non-Commercial License.
00:06:33And I want to use their exact language here because this distinction actually matters.
00:06:38What this is going to say, what this actually does says is,
00:06:42this agreement permits research and non-commercial use of the model materials free of charge.
00:06:47It does not grant any commercial rights and it is not an open source license.
00:06:52That's their words.
00:06:53Not an open source license.
00:06:55Okay.
00:06:56Now, sometimes you see non-commercial on a model and it basically just means don't resell our weights.
00:07:02Yeah, okay.
00:07:02That makes sense.
00:07:04This is not what this says.
00:07:05Their definition on commercial purpose includes production use,
00:07:09using it in a product or a service, hosting it behind an API.
00:07:13And then there's the part that really just kicks in here.
00:07:16As if that wasn't already enough,
00:07:17it also includes using the outputs for any internal operations beyond limited evaluation.
00:07:23The outputs.
00:07:24Okay, so what does this actually mean?
00:07:26Not just the model itself, the audio it generates.
00:07:29And this whole thing just keeps going.
00:07:31There's no creator exception, no small business exception.
00:07:35There's no revenue threshold.
00:07:37There's nothing here.
00:07:38A lot of non-commercial licenses leave some gray area.
00:07:42This one goes out of its way to basically block a lot of that.
00:07:45So, let's make this more practical now.
00:07:47When could you use this?
00:07:48Put a generated breeze voice inside a product?
00:07:51Nope.
00:07:52That's not going to happen.
00:07:53Use it in a monetized YouTube video?
00:07:56This isn't monetized.
00:07:58Yeah, no.
00:07:58Not going to happen.
00:07:59In some kind of podcast?
00:08:00Again, not going to happen.
00:08:02So, this blocks pretty much everything.
00:08:04The license specifically names sponsorship and subscription revenue.
00:08:09$34 per million characters once you get paying for this.
00:08:12But even there, the license makes something very clear.
00:08:15Paying for the API does not give you commercial rights to self-hosted models
00:08:19or the outputs from their self-hosted model.
00:08:22And once you see that, the whole setup makes a lot of sense.
00:08:25The open weights are the demo.
00:08:27The API is their whole product.
00:08:30So, how usable are those open weights in the first place?
00:08:33You'd be the judge of that.
00:08:34Now, this was very cool.
00:08:36Voice AI is very cool.
00:08:37But I did a breakdown on Vox CPM2 a little while back.
00:08:41Vox CPM2 was honestly insane.
00:08:43It was very, and I mean very, impressive.
00:08:45So, if I stack Breeze against Vox CPM,
00:08:48Breeze would lose this in a heartbeat.
00:08:50And it all has to do with the whole license.
00:08:52Spectacle.
00:08:53So, if you want an honest, open voice clone voice AI,
00:08:56I'm still sticking with Vox CPM2.
00:08:58As it has all the same features as Breeze,
00:09:00and it might be even better.
00:09:02Should you use Breeze?
00:09:03That's the question.
00:09:04I think I answered that, but let's play around here.
00:09:06For playing around with voice AI, yeah, you could.
00:09:09It's cool to see where voice AI is going.
00:09:12Breeze is pretty good, okay?
00:09:14Testing out models is fun,
00:09:16especially as they get better and better.
00:09:18Where are things going?
00:09:19Voice design from a text description
00:09:21is a genuinely really cool capability.
00:09:23But if you're shipping anything commercial,
00:09:25don't use this.
00:09:26I would use something else,
00:09:27regardless of where Breeze sits on the leaderboard.
00:09:30Kokoro is Apache 2.0.
00:09:32Chatterbox is MIT.
00:09:34Vox CPM, we'll look into it,
00:09:35because that might crush them all.
00:09:36Open weights and open source
00:09:38stopped meaning the same thing a while ago.
00:09:40Breeze makes that difference impossible to ignore.
00:09:43The benchmark tells you what the model can do.
00:09:45The license tells you what you can do.
00:09:48I'm Josh from BetterStack.
00:09:49If you enjoy coding tips and tricks like this,
00:09:51be sure to subscribe.
00:09:53We'll see you in another video.

Key Takeaway

Breeze TTS 2 matches proprietary models on benchmarks through custom voice selection, but its non-commercial weight license strictly blocks all production and monetized use.

Highlights

  • Breeze TTS 2 requires a 3GB download and runs locally on a MacBook using MLX audio.

  • Artificial Analysis ranks Breeze at 1215 ELO, narrowly beating 11Labs conversational model scored at 1210.

  • Breeze utilizes a Qwen 3 backbone, a T5 Gemma 2 text encoder, and a MIMI audio codec for 24 kilohertz mono audio generation.

  • The model weights use the Breeze Blue Research and Non-Commercial License, which explicitly prohibits production use, APIs, and generated audio in monetized content.

  • Apples-to-apples benchmarking places Breeze at 1002, ranking third among open-weight models and 16th overall out of 39 systems.

Timeline

Voice Design Capabilities and Local Setup

  • Breeze generates custom voices directly from descriptive text prompts instead of requiring reference audio.
  • Local installation runs through the terminal by creating a virtual environment and installing MLX audio.
  • The CFG scale controls how strictly the model adheres to the provided voice description.
  • Model architecture combines a Qwen 3 backbone, T5 Gemma 2 text encoder, and MIMI audio codec.

Text-to-speech workflows typically require ten seconds of reference audio for voice cloning, but voice design allows direct generation from text descriptions like a calm, reflective delivery. Users can adjust the CFG scale to control adherence to the prompt, dropping it to lower values to produce a generic AI voice. The system also supports inserting vocal events like sighs or laughter directly into the text syntax. The underlying 3 billion parameter model outputs 24 kilohertz mono audio and integrates Apache 2.0 components from Qwen 3 TTS.

Benchmark Analysis and Leaderboard Realities

  • Artificial Analysis ranks Breeze at 1215 ELO based on blind pairwise human votes.
  • Breeze holds the fewest votes near the top of the provider voice leaderboard with roughly 1200 votes.
  • Apples-to-apples testing with identical prompts drops the Breeze score to 1002.
  • Cartesia Sonic 3.6 leads the leaderboard with an ELO score of 1282.

Independent benchmarking firm Artificial Analysis places Breeze slightly above 11Labs conversational models on the provider voice leaderboard. However, this ranking allows each model to select its own optimal voices. When evaluated on identical test prompts, the score drops significantly, placing Breeze behind alternative open-weight models like Vox Troll. Selective voice choices inflate leaderboard placement, altering the perceived competitiveness against frontier proprietary systems.

Licensing Restrictions and Commercial Limitations

  • The code uses an Apache 2.0 license, while the model weights operate under a non-commercial research license.
  • The license explicitly bans production use, hosting behind an API, and internal operational use of generated audio.
  • Paid API tiers do not grant commercial rights to self-hosted models or their outputs.
  • Alternative models like Kokoro and Chatterbox offer permissible Apache 2.0 and MIT licenses for commercial deployment.

While the repository code is open, the model weights restrict usage exclusively to research and non-commercial purposes. The agreement defines commercial purpose broadly to include product integration, monetization, sponsorship revenue, and internal operations. Creators cannot use generated audio in monetized YouTube videos or podcasts. Because the open weights serve primarily as a demo for the API, projects requiring commercial viability must rely on alternatives like Vox CPM, Kokoro, or Chatterbox.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video