Transcript
00:00:00A brand new AI stealth model has dropped and it is posting benchmarks that compete with the best
00:00:05models OpenAI and Anthropic have to offer. It's called Union Alpha and we can currently try it out
00:00:10for free through providers like OpenRouter. Now we still have no idea which AI lab is responsible
00:00:16for Union Alpha but with benchmarks like these it's worth our time to take a closer look which
00:00:21is exactly what we're going to do today. So I'm going to show you how to get Union Alpha on your
00:00:26machine. I'm going to break down what its stats are and then we're going to put it head-to-head
00:00:31against models like Fable, Astra, and Opus. So before we jump into the head-to-head testing let's just
00:00:37talk a little bit more about this model and how you can get access to it. So Union Alpha no idea who's
00:00:43responsible for this thing in terms of its context window. We're looking at 262,000 tokens and right
00:00:50now it is completely free. Now when it comes to performance we have some unofficial reports. Right
00:00:55here is the DeepSuite score. This graph came from the CEO of OpenRouter so maybe some conflict of
00:01:00interest. Either way if this is true this is showing Union Alpha over here with the star performing at the
00:01:05same level as Astra and Fable as Opus 5 at a fraction of the cost which is wild. Although I will say
00:01:13someone who is a big fan of DeepSuite I don't know how great it is anymore when we have stuff like
00:01:19Gemini 3.8 flash scoring so high. Like you see stuff like that and you're just like is this the
00:01:25end-all be-all for benchmarks? No and no benchmark really is. But if this is true that's still a great
00:01:31sign. Alex then posted more stats. Again this is the CEO and co-founder of OpenRouter saying that the
00:01:37outperforms 5.6 Sol on Terminal Bench 2.1 and the Swee Bench Verified and also outperforms Opus 5 on
00:01:43Terminal Bench 2.1 as well. Now we've gone well beyond Terminal Bench 2.1 so again take all of these
00:01:50BenchBrush with a grain of salt. But if this is true it is a highly performing model. So how do you
00:01:57get access to this thing? I think the easiest way to do it is through OpenRouter. So if you go to
00:02:01OpenRouter.ai you just set up an account and this again is totally free so it's not going to cost you
00:02:06anything. Now once you have an OpenRouter account set up then we need some sort of harness to use
00:02:11Union Alpha. For this demo I am using OpenCode. So really easy to connect OpenRouter to OpenCode.
00:02:18Once you do that then you simply go on here down here and you choose your model. In this case I am
00:02:23using Union Alpha. Now let's talk about some of the benchmarks we're going to run. I did two tests
00:02:27related to front-end design. I had it do some game design. I had it do motion design. And then lastly
00:02:33I did essentially a modified eight needle test to see how well it does on larger context windows and
00:02:39being able to grab the correct information. But before we hop into the head-to-head tests
00:02:42a quick word from today's sponsor me. Now inside of Chase AI Plus I have just released an updated
00:02:48Claude Code Masterclass and inside of here you also get a Codex Masterclass and this is the perfect place
00:02:53to go from zero to AI Dev no matter your technical background. We focus on real use cases. I assume
00:02:59you know nothing going in here. So if you're someone who wants to get more serious with this stuff but has
00:03:04no idea where to start this is the place for you. You can find a link to it in the pinned comment.
00:03:09Now for test number one I wanted to take a look at how it does with front-end design. And so I gave it a
00:03:14very simple non-prescriptive prompt that I wanted it to execute without pulling in any
00:03:19outside skills. The prompt was very simple. Create me a landing page for a fake boutique hotel in London
00:03:25don't use any skills. And what it created in three minutes and 37 seconds was this. So again no skills
00:03:32were used. Extremely simple prompt. Hero section very simplistic. We have sort of this rolling banner
00:03:38right here. As I scroll down I don't hate essentially what it's created. I do like the fact that it doesn't
00:03:46necessarily scream AI in terms of like some of the things you normally see which is like
00:03:51rounded corners like those very standardized cards that you see everywhere. So it's a little bit off
00:03:58the beaten path in that regard. And in general in terms of like the color scheme I'm not super
00:04:02against it. Yeah. Is it very very simple and like not gonna blow you away? Sure of course but that's a
00:04:08function of the prompt. And it actually knocked this out relatively quickly right. Three and a half minutes
00:04:12is not bad. So I would say this output is decent all things considered. Now I went ahead and gave Opus
00:04:185 the exact same prompt on a high effort setting. There's no effort settings when it comes to Union
00:04:23Alpha. And this is what it created. So I do like the fact that right off the bat we get this sort of
00:04:29ability to check the availability. Now if we compare that here it's just like a check availability button
00:04:34here that takes you all the way to the bottom and it isn't as intuitive as this one right. This makes a lot
00:04:39more sense. It went a little more in depth in terms of creating this background with what I assume our
00:04:44SVGs. And it has a little bit of like motion as I scroll down the page you might have seen that.
00:04:53So in general I kind of like what Opus created a little bit more. I think it is slightly better. Again
00:05:00very simplistic but that is a function of the prompt at large. So in general I would give the edge to Opus.
00:05:09Now and this goes for all the benchmarks. If Opus does better than this thing that doesn't
00:05:14necessarily mean that Union Alpha is quote unquote worse. Until we know the actual cost of this in
00:05:19reality it's kind of hard to make a judgment call there. If the costs kind of match up with what we
00:05:25saw in that deep sweep sort of benchmark that OpenRouter put out then this is a great output.
00:05:30But that being said I would put Opus slightly ahead. Now for our second test I had it again create a
00:05:35landing page but this time the prompt is much more detailed. Now I still give it room to kind of use
00:05:40its judgment in terms of layout implementation. You can see that on the last line but I gave it some
00:05:45guardrails. So I said build a polished responsive homepage for Dune House a fictional boutique desert
00:05:51hotel in Joshua Tree California. And this is what it created. So here's what it created for us and right
00:05:56away you can tell this is a huge leap forward from that first test. This is just a very simple
00:06:01prompt and here's something with just slightly more guidance. I love the image it used in the hero
00:06:06section and as I scroll down you can tell this is definitely much more polished. I will say I wish
00:06:11there was something at the top in terms of finding availability like we saw in the Opus section but
00:06:17just the use of images alone totally elevates this web page it built. And here's a look at what Astor
00:06:23created using that same prompt. Now they both went for this like hero section with the image in the
00:06:27background which I really like. I will say I do like what Astor did here similar to what we saw with Opus
00:06:33where it has you know a very clear section to like check availability. But as we scroll down you see a
00:06:40lot of similarities in terms of how it uses these images and the difference between what we got with
00:06:45Astor and what we got here with Alpha Union isn't that big. Now I would probably give the edge just a
00:06:51little bit to Astor but it's not a huge gap. I mean we're a couple problems away from pretty much making them
00:06:56the same thing. So you know with that in mind huge thumbs up for Alpha Union. Now for the third test
00:07:03I wanted to lean into some game design. This is something we had done in a previous video
00:07:07where we looked at Fable versus Astor and had them essentially create a Fortnite clone.
00:07:11I gave Union Alpha the exact same prompt included the same reference images and it just straight
00:07:16up struggled to do that. It aired out every single time. I don't know if that was an issue
00:07:20just with open router in general or it's literally just a capability problem with the model and I just gave
00:07:25it too much. Hard to say. So instead what we did is I had it build a 3D arcade racing game for the
00:07:32browser. I said I wanted to be on a stylized coastal highway and then I gave it some guard rails in terms
00:07:37of what it should be trying to build. Now when I had it build this game it ran into some errors and I had
00:07:41to restart the session. I'm not sure if it was an issue with the model itself or it was just some sort of
00:07:46routing problem. And this is the game it created so let's check it out.
00:07:57So this is all using 3JS. Controls seem to work pretty fine. It's relatively smooth. Now this is
00:08:04nothing complicated. It's just a single circular track. If I hit things there is some physics involved
00:08:11but overall it handles pretty nicely and in terms of the graphics it's not bad. And here's what Astra
00:08:17created. Right away you see a lot of similarities with the loading screen. So let's start this one up.
00:08:26So they ended up using you can't really hear the audio and I'm going to spare you from it because
00:08:30it's kind of annoying but they use the exact same audio. And in regards to the controls and how it feels,
00:08:37again pretty smooth. We have physics. I can bump into guardrails. It actually gives me some hints of
00:08:43how to actually drive better but overall solid and I would put it exactly the same tier as what we got
00:08:49with Union Alpha. Not a huge difference in either one and they pretty much gave us the same result.
00:08:54And frankly with any of these benchmarks if Union Alpha is holding strong with these top models I
00:09:00consider that a win for Union Alpha. For the next test I wanted to see how it would do with motion graphics.
00:09:05So I told it I wanted to use the Higgs field motion skill and I wanted it to create a 2d motion
00:09:10explainer about how your message travels across the internet. And this is the video it created.
00:09:30So that was really solid. Let's take a look at what Astra created.
00:09:49So both of those really strong. Now that test is really built on the back of that motion skill and
00:09:55it's really just calling on an MCP server to create these using SeedDance. But it's good to see that if
00:10:00we throw some skills at something like Union Alpha it's going to be able to do just fine and operate
00:10:05in the same manner again as these Frontier models. Now for Union Alpha's final test we ran the eight
00:10:10needle test. Now we did it on the official 128k version as well as a modified 200k version so it could
00:10:17fit its context window. Basically what this test is looking at is it wants to see how well can this model
00:10:22answer questions when we flood its context window. So for the 200k test basically what we did is we filled
00:10:29up Union Alpha's context window with 200,000 tokens worth of text. Inside of that text were eight similar
00:10:37responses. So I want you to imagine that inside that text we had eight different things that said
00:10:42this is a poem about flamingos. So it had a poem about flamingos number one and a poem about flamingos
00:10:47number two and number three and number four but all those poems are slightly different and at the end
00:10:52what we ask is we say hey bring me back exactly the second poem about flamingos. So first of all it needs
00:11:00to be able to identify what is a poem about flamingos and then it needs to be able to separate all eight
00:11:06of them and then give me exactly the second one. Now when I ran this test on the 128k version it scored a
00:11:13100 percent when we ran it on the modified 200k it got a 99.6 percent. So it was off by five apostrophes.
00:11:21So that tells us when we wanted to answer questions accurately when we pretty much filled up its
00:11:25context window is it going to do so. The answer is yes it's going to be pretty accurate which is great.
00:11:30So all in all when we look at all those benchmarks this looks like a really solid model pretty much across
00:11:35the board. It did just as well as Opus. It did just as well as Fable. It did just as well as Astra. So
00:11:40hopefully when this thing comes out it's pretty competitive price wise because if that holds up
00:11:45across a number of different domains this could be something we could all kind of implement into our
00:11:50stack. So definitely keep an eye out for this guy and as always make sure to check out Chase AF+ if
00:11:57you want to get your hands on my Claude Code and Codex Masterclass but besides that I'll see you around.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video