I Tested GPT 6 Astra vs Fable 5.1 (No Hype Assessment)

CChase AI
Computing/SoftwareVideo & Computer GamesInternet Technology

Transcript

00:00:00So this week was the Battle of the AI Titans. We got the release of Claude Fable 5.1 and GPT-6
00:00:06Astra. Now, not everyone has access to GPT-6 Astra. They're slowly phasing it in, but I was
00:00:11lucky enough to get early access. And today, I'm going to show you some head-to-head tests
00:00:15with GPT-6 versus Claude Fable 5.1. So you can kind of figure out what makes the most sense for
00:00:23you. Yes, we'll take a look at the benchmarks at the beginning, just to kind of give you an idea
00:00:27of the numbers, but I figured it was more important to see some real use cases. We're
00:00:32going to take a look and see how well they do head-to-head in things like front-end design,
00:00:36motion graphics, web applications, and even some longer-running agentic tasks related to
00:00:41game design. The goal here is to give you a snapshot of where these two models stand when
00:00:46compared head-to-head so you can make a more informed decision about where you should invest
00:00:50your time. So let's hop in. So first, let's very quickly go over the numbers here, the benchmarks
00:00:55that they reported. I know we kind of all take these with a grain of salt, but there's still
00:00:59some information we can pull from here. Now, kind of a bummer that Anthropic didn't give
00:01:03us more data. You'll notice there is a lot of benchmarks that they just didn't report compared
00:01:08to OpenAI. OpenAI kind of gave us everything you could possibly want. And based on the GPT-6
00:01:13numbers, it pretty much beats Fable 5.1 in almost everything. Well, based on these numbers, literally
00:01:20everything, which is very impressive. Now, like just because on DeepSuite version 1.1 and it's a 74.1,
00:01:28does that mean then it's, you know, it's actually seven and a half percent better than Fable 5.1? And
00:01:34in fact, Gemini 3.8 Flash is better than everything else except GPT-6. No, you really can't look at these
00:01:39numbers in the vacuum. But when we take a look at it holistically, what do we see? We see a huge jump
00:01:45forward, especially when you compare it to something like 5.6. So GPT-6 definitely is a step change from
00:01:50previous GPT models. And Fable 5.1 across the board, what do we see? Well, we just see a very strong
00:01:56model that by the numbers is better than Fable 5. And if you're someone who's used Fable 5, you know
00:02:00for a fact that this is a solid model. Like it's good. So by the numbers, what can we genuinely say?
00:02:06Well, we can say GPT-6 is a big leap forward on the OpenAI side. And Fable 5.1 is just giving us
00:02:11more of what we already like. So good start. Now, when it comes to the numbers, the actual performance,
00:02:17the percentage of scoring on these benchmarks are just one part of the equation. The second half of
00:02:21the equation is the cost. How many tokens is it taking each of these models to achieve those numbers?
00:02:27Because historically, the GPT models have actually been a lot better in this regard. And when we look
00:02:31at the stats, that kind of plays out here as well. Now, I'm looking at OpenAI's press release for this
00:02:37for full disclosure. And if we look at something like Terminal Bench 4.0, and I look at Astra,
00:02:41which is here with the stars compared to Fable 5.1, which is here in the orange, we can see that at
00:02:47their best, they're pretty close. So at max, we're getting 55.8% accuracy with Fable 5.1. And with Astra,
00:02:55it's 56.7. So, you know, we're basically the same. But where they aren't the same is the cost. At max,
00:03:02I'm at $10.35 for Astra, and I'm at $19.50 with Fable 5.1. And that's even more of a significant
00:03:11gap when we look at stuff like extra high here versus extra high there. And in general, you can
00:03:17look at this whether we look at really like any of these different benchmarks. GPT 6 tends to be a lot
00:03:23cheaper with performance that matches 5.1. Now it's time to actually put them to the test
00:03:28on some real use cases. But before we do that, a quick word from today's sponsor, me. So inside
00:03:33of Chase AI+, I have not only just released a Cloud Code Masterclass, I have also released
00:03:38a Codex Masterclass. So if you're someone who's trying to figure out how can I master any of the
00:03:43tools you see here in today's video, Chase AI+, is the place for you. I assume you have no knowledge
00:03:48going into any of this, you don't need to be technical. And we focus on real use cases,
00:03:52real projects, so you can actually apply this to whatever your pet project is. So if that sounds
00:03:58like something you're interested in, definitely check us out. There will be a link in the pinned comment.
00:04:02So for our first test, I'm going to see how well Fable and GPT 6 can recreate Fortnite. I was
00:04:09inspired by this video that I saw on YouTube. It's called Claude Fable 5.1 is insane by this guy
00:04:14named Cole. He only has like 7k subs, so definitely check it out. But he was showing how he did this
00:04:18with Fable. And I thought, hey, let's try it out. So I gave it the prompt you see here. I gave it
00:04:24several reference images. And I was like, pretty much, let's create Fortnite that I can play in the
00:04:29browser with 99 bots and different weapons and all those things. And this is what they were able to
00:04:34create. So first up, we have GPT 6. You can see this is sort of the selection screen, I can choose my
00:04:39different game modes. We're just going to do solo, but I can add some bot teammates if I want.
00:04:44Move back over here. It has sort of like how to play, it gives me the different buttons,
00:04:49I can do my settings, but let's just jump in and see what happens. So I'll hit ready. And here we are
00:04:55waiting for the battle bus. Now we are inside of the bus. This is also supposed to have an entire like
00:05:02storm element. You can see here. As I jump, I can do a glider, which is pretty good. And also
00:05:09remember, this is a one shot as well. I just gave it the prompt I showed you guys earlier. And
00:05:14that was pretty much it. So let's see how this plays out. You can see the bots kind of floating
00:05:20around. We have chests up top here. Let's see how that works. I can grab guns. I can give myself
00:05:32potions. So you can see my shields going up there on the left. Looks like I got a shotgun here. The
00:05:39camera's a bit shaky. But there's a bot. He's dead. The bots are not very good. I purposely told
00:05:51it to make the bots kind of bad. But you can see I can zoom in. Like I can reload. I can pick up
00:06:01different guns from people. Right. So I can switch it in my inventory. And so I can tab. I can see the
00:06:09whole map. And honestly, besides the bots being terrible, like not a bad job. See what happens
00:06:17if I try to like harvest something like I get brick. It doesn't really show anything coming out. There's
00:06:22like small little like specks of dusts. But honestly, overall, pretty solid. This took Codex about 45
00:06:30minutes to do this. Now I don't have an exact token amount, unfortunately, because of how the early
00:06:37access worked. It wasn't like tied to usage. And so I can see that. But for a one shot here, pretty
00:06:42solid. And you probably can't hear the actual audio here at all. But it's kind of just like audio it
00:06:48created on its own. So it's nothing too impressive. Oh, yeah, I can also build stuff. Pretty cool.
00:06:55Let's see what else I can build like. See if it actually works. So yeah, I can like jump
00:07:01up it. So overall, not bad for a one shot. So that was Codex. Now let's take a look at what Fable
00:07:08created. So here's Fable, kind of like relatively similar. I'll move this over here so you can see a
00:07:15more. When I look at Codex is I kind of like it has the map thing over here. I don't see that on
00:07:20Fortnites. So a little less clean, I would say. But when I look at the Fortnight lobby, it has more
00:07:28things for me to click on. But these don't actually do anything. And I click on them. Now let's see.
00:07:33So let's say we're ready to go. It's filling up the lobby. And also for reference, this took Claude
00:07:39about like an hour and a half to create this and about 750,000 tokens. Here's the battle bus area.
00:07:48The camera controls are a little more janky, but I'm not going to hit on that too much space. I will say
00:07:55this has a super annoying noise as I'm jumping out. And in general, the camera again is kind of janky.
00:08:03Like I can't really, it's really hard to kind of move it. Graphics overall, I would say are a little
00:08:11bit worse than what we had with what Codex created. But not like huge difference, to be honest. So here
00:08:20I'm in front of the tree trying to mine it. I will say the animations are a bit wonky. And that was
00:08:25very strange with how the tree disappeared. I can still build things. Interesting kind of way to do
00:08:33it. When I try to move into, you know, what I've actually just built kind of janky just slides me
00:08:38through it. I finally found a gun. I will say in terms of like gunplay with this one, definitely
00:08:42leave something to be desired. Like where I shoot versus where it goes is like this crazy bloom going
00:08:48on. But hey, I didn't eliminate somebody. And so overall, I would say just feels a little less
00:08:54polished. Again, this is one shot. And it took like an hour and a half. So it is interesting to kind of
00:08:59see what we can do on a single pass when we give it a very long prompt like that with a bunch of
00:09:03different references. So overall, it's kind of impressive. It was able to do this definitely
00:09:08didn't get to what I saw in the reference video that I showed you earlier. I don't really know if
00:09:13you one shot it was a bunch of different passes like sometimes you just never know. But for just a
00:09:17straight one shot, I will say there kind of is no competition here. I feel like codex kind
00:09:22of nailed it. Now for test number two, I wanted to take a look at front end design. This was the
00:09:26prompt I gave it. I said, I want you to create a landing page for an AI travel website. They put in
00:09:32there where they're going and you create an itinerary or something relatively simple. What I was really
00:09:37focused on is what is it actually going to look like? I said they can do whatever research they want.
00:09:41They can bring in whatever tools or skills they deem appropriate. And this is what we got on the first pass.
00:09:45Now this is what GPT-6 created. It generated or found this image. I didn't actually check which
00:09:51remember that GPT-6 has access to its own internal image model, which is really nice for this sort of
00:09:56stuff. So the hero section looks pretty clean. They have like this little thing where I can put where
00:10:02I'm departing from and where I'm going. I told it I didn't care about the functionality. So I wasn't
00:10:06actually going to test like, does this actually work under the hood? I was just focused kind of on
00:10:10aesthetics. You go down here, it kind of shows you what's happening. Hey, five days, a thousand
00:10:15little stories. Again, we have some imagery. And I will say in general, as I look at this,
00:10:20this isn't screaming AI to me. This isn't screaming AI sloth. This is actually pretty clean. If I was
00:10:27someone like, hey, what is this website about? What do they actually do? I think it does a pretty good job
00:10:32of explaining that. So it goes down here. Hey, we bring the curiosity, blah, blah, blah. Lots of images,
00:10:38right? It definitely like bringing in a ton of images, a little FAQ section, and then the footer. So overall,
00:10:45I think this is solid. This isn't blowing anyone away. This isn't an awards level website. But for
00:10:52just like, clean, simple, figure out what you want to do and do it. I didn't give a ton of creative
00:10:56direction. I'm like, actually pretty impressed. On the other hand, with the exact same prompt,
00:11:02this is what Fable 5.1 gave me, which is kind of like, eh. Right? Like, this comes across very,
00:11:10very basic. You know, like, the coloring isn't very inspired. Like, this is a very generic kind of hero
00:11:19section where we have text on the left, some sort of image on the right. You know, it's not like GPT-6,
00:11:25where it's able to create its own images. It didn't decide to find an image. Instead,
00:11:29it just like created this using its own internal graphics system. It has this little border here,
00:11:35which looks very AI. We have sort of our typical cards, you know, rounded corners. There's not even
00:11:42really any motion there whatsoever. And in general, it just feels very like it's light blue into dark blue.
00:11:51And honestly, I was kind of disappointed with what Fable 5.1 gave me because in general,
00:11:57I feel like the anthropic models have been better when it comes to front end design. So I'm not sure
00:12:03if it just fit off on like, maybe the wrong front end design skill or just like research something
00:12:08weird. But overall, I felt when you compare these two things, kind of night and day, kind of night and
00:12:14day. Now, when it comes to front end design, in general, I will say there is a huge gap between
00:12:19those who are good at using AI and those who are bad. And so the better you are, the less the model
00:12:24almost matters because you understand how to give it reference images, how to go like what components
00:12:29should look like where to go outside, you know, outside of these models on the internet and found like
00:12:34fine component libraries. So there's a ton of skill involved here. And like, we can have this argument
00:12:38about how, how much does it matter if a model is baseline really good at front end design, if I can
00:12:43sort of coach it. In reality, that's not most people, most people just like build me this and
00:12:48we see what we get. And I think like, if this is the median output, we are judging here. If you're
00:12:55not good at prompting, kind of no question. Again, I would give Astra the W here. So for the next test,
00:13:02I want to look at motion graphics. So I hooked up both models to the Higgs field MCP and had them call
00:13:07the same motion graphics skill I've created. And what I wanted them to do was create a 15 second
00:13:132d explainer on how internet messaging works, right? Or if you're on your phone and you text someone,
00:13:19how the heck does that work? So this is what Codex came up with.
00:13:37Honestly, pretty good. So here's what Fable gave us one tap, half a planet away chopped into packets.
00:13:50A dozen hops, milliseconds, every message, every time.
00:13:57So overall, I thought both these models did really well, I would call it a tie. They're both really
00:14:01good at calling these outside tools like the Higgs field MCP routing their prompt to something like
00:14:06Seed Dance 2.5 and giving you solid motion graphics. Now for the final test, I had both models execute
00:14:12this prompt. This is also somewhat front and design related, but I really wanted to see how they could
00:14:17push the bounds in terms of creativity versus our landing page design, where one of the parameters was
00:14:22like, this needs to be just functional. Like this is, we're trying to sell a product. Someone should be
00:14:26able to very clearly understand who we are, what we are about and how to use it. And this one,
00:14:31right, it had some room for the visual spectacle. So we wanted to create this 3D globe dashboard kind of
00:14:38web app. And this is Orbit. And this is what GPT-6 created. So I put where I'm departing from, let's
00:14:45say we're coming from New York, I go wherever I'm going to go, let's say we're going to Cape Town. And
00:14:52you can see it shows it on the map. I can kind of click wherever I want, which is kind of neat. If I
00:14:57move this over here, you can see there's information on the bottom. So we're going to go to Cape Town,
00:15:02going from SFO to CPT, one stop, 22 hours, it's 19 degrees out there. And for the sample round trip,
00:15:08we're looking at $864 per person. If I then click on explore Cape Town, I have this information over
00:15:14here on the left, right? Gives me a cool little picture. It gives me some information about the
00:15:18place. It gives me some potential things to do, right? Hike Lion's Head. And I can even save the
00:15:24journey. It also had this thing called meet Cape Town at golden hour. Like one of the things it added was
00:15:28this chase the sun functionality, where it shows me on the globe, like where golden hour is at any one
00:15:35time, which is kind of cool, right? I gave it some room for creativity. And that's what it came up with.
00:15:39And not bad. Again, what I really like about GPT-6, it's just clean, very, very clean design. So I
00:15:45like this. Then we have Arclight, which is what Fable 5.1 created. And right off the bat, it definitely
00:15:53goes for the visual spectacle, not just with that loading screen. But this is just a lot,
00:15:58very, very bright, tons of things going on. And like we have like, the text is like following
00:16:04the route itself. I have a bunch of different lines. There's like this craziness going over
00:16:09here with like the fares and their live fares. I can see what the percentage has increased as of late.
00:16:14So honestly, kind of crazy off the bat, like this looks really cool. But when we compare this right
00:16:21away to the Astra design, like not as clean, it's almost too much. Although I think it's cool. And this
00:16:28could definitely be cleaned up again. This was all one shot. Now, similar to what we saw with Astra,
00:16:32I sort of have like the solar clock thing going on. And as I move it and change the time, the fares
00:16:39update with it, I can see certain ones that are trending. And if I click on any of these,
00:16:44like let's say I click on Helsinki, we get this sort of cool animation going on where it brings us to
00:16:51Helsinki over here and then gives us some information about it, the local time, the weather, the winds,
00:16:54visibility, all this stuff, you know, I can hold the fair right and like really cool animations.
00:17:02Although it's like I can kind of see it. It's actually kind of hard to see. It's sort of blurry
00:17:06because it's so bright in the background. So it almost feels like visual spectacle became the number
00:17:12one priority with fable 5.1 versus functionality. And like, that does look sweet. Like this whole
00:17:19animation. But when I see this, I'm almost like, Hmm, what if we could take some of this, you know,
00:17:23craziness and like sort of temper it with what Astra was able to generate. So in general,
00:17:30I would call this another win for Astra, to be totally honest, especially when I can kind of do
00:17:35like this whole like explore Cape down thing, like you see over here on the left, and that really goes
00:17:38for any of these places. Just like I like it looks very professional. And I don't feel like I would have
00:17:45as much work to do versus here. Like with fable 5.1, I feel like I'm many prompts away before this is like
00:17:52usable, in my opinion. So where does that ultimately leave us? Well, I mean,
00:17:57based on the tests I showed you, we did kind of four tests. Astra won three of them. Fable 5.1,
00:18:03I would say was even with it on the motion graphics. And when we look at the raw benchmarks,
00:18:07we see that GPG six by the numbers also kind of edges ahead of the competition, plus the fact it's
00:18:13actually cheaper by tokens. But these were just for tests, they were just one shots. How important is
00:18:22it to you to have a model that has built an image generation? How comfortable are you in
00:18:27something like the codex desktop platform versus the anthropic platform? Are you someone who even
00:18:31needs to choose between the two? Are you someone who actually uses both? You know, we watch a video
00:18:36like this, and I hope you were able to come away with something. And really what you should come away
00:18:41with is that both of these models are just really good. You know, as someone who's been using Fable 5.1 a
00:18:45ton over the last three days, I almost feel like these tests, I put it through, put it through kind of
00:18:50downplay what it's able to do, because I really have yet to come across someone who's like,
00:18:54Fable 5.1 sucks. It's just isn't it. At the same time, Astra feels great. It really does. All the
00:19:02generations it had were very smooth. I'm not going to say AGI, because the definition of AGI changes
00:19:09day by day. But I think more importantly, there is a question now of Astra versus 5.1. There was
00:19:14certainly a contingent of people out there over the last month or so who were like, "Hey, 5.6 actually does
00:19:19compete with Fable 5." But I think that was a very small community of people. And when we take into
00:19:25account like how Anthropic has been treating sort of like its user base, you know, take the great
00:19:31examples how like they lowered limits and then pretended they were actually bumping it up when
00:19:35reality had been less over the last month. Like it just kind of leaves a bad taste in your mouth.
00:19:40Fable 5.1, we still only get like 50% of our weekly usage. And then 20x isn't really 20x.
00:19:46And then on the other side of the equation, we have OpenAI, which is just like giving us a reset
00:19:50every three days, although they did take away some of the bank's resets. But in general, it feels like
00:19:54they're giving you more than Anthropic. So I don't think it's such a big difference between
00:20:02Astra and Fable 5.1 that if you like Anthropic, you should just like throw it away and cancel it.
00:20:06But there's a real question about it now. And if you're someone who's been on the 20x Fable plan,
00:20:12and guess what? 20x doesn't actually mean 20x. I think we're at a place where you should really
00:20:16consider like maybe I do the 5x plan with OpenAI and the 5x plan with Anthropic for a while. Still
00:20:24paying $200 at the end of the month. But now I have options. And now I can really test it in my
00:20:28day to day because no video and no test and no benchmark is really going to be the same as when
00:20:32you get in there and use it yourself for your projects. So that's where I'm going to leave you.
00:20:39As always, if you want to learn more about how to master cloud code and codex,
00:20:43I have masterclasses for both of those inside of Chase AI+. But besides that, let me know what you
00:20:47think. Would love to see what you are able to do with Astra because they are starting to roll it out
00:20:52for everyone starting today. I've seen some posts about that. But besides that, I'll see you around.

Key Takeaway

GPT-6 Astra surpasses Claude Fable 5.1 in benchmark scores, cost efficiency, and one-shot generation speed across front-end design and gaming applications.

Highlights

  • GPT-6 Astra outperforms Claude Fable 5.1 across benchmark scores while consuming fewer tokens on comparable tasks.

  • GPT-6 Astra costs $10.35 at maximum usage on Terminal Bench 4.0 compared to $19.50 for Claude Fable 5.1.

  • GPT-6 Astra takes 45 minutes to generate a functional browser-based Fortnite clone in a single shot, while Claude Fable 5.1 requires one and a half hours and 750,000 tokens.

  • GPT-6 Astra generates native images directly via its internal image model during front-end landing page creation.

  • GPT-6 Astra wins three out of four head-to-head tests against Claude Fable 5.1, with a tie in motion graphics.

Timeline

Model Benchmarks and Cost Efficiency

  • GPT-6 Astra represents a significant step change from previous OpenAI models.
  • OpenAI provides comprehensive benchmark reporting, whereas Anthropic omits data for multiple metrics.
  • GPT-6 Astra delivers performance matching Claude Fable 5.1 at a lower token cost.

Initial evaluations compare reported benchmark scores and token costs between GPT-6 Astra and Claude Fable 5.1. OpenAI discloses broad benchmark metrics showing across-the-board advantages, while Anthropic supplies limited data. Cost analysis on benchmarks like Terminal Bench 4.0 reveals that maximum expenditure reaches $10.35 for Astra and $19.50 for Fable 5.1, establishing Astra as the more cost-efficient model.

Browser-Based Game Recreation Test

  • GPT-6 Astra builds a functional browser-based Fortnite clone in 45 minutes on a single pass.
  • Claude Fable 5.1 requires one and a half hours and 750,000 tokens to complete the same game replication task.
  • GPT-6 Astra delivers cleaner map layouts and more responsive mechanics than Claude Fable 5.1.

Both models receive identical prompts and reference images to recreate a playable browser-based Fortnite game with bots and weapons. GPT-6 Astra completes the task in 45 minutes with polished UI elements and working mechanics. Claude Fable 5.1 takes ninety minutes, consumes 750,000 tokens, and produces jankier camera controls and lower visual fidelity.

Front-End Design and Travel Landing Page Test

  • GPT-6 Astra utilizes its internal image model to generate relevant visuals for an AI travel website landing page.
  • Claude Fable 5.1 produces a basic interface with standard cards and minimal aesthetic styling.
  • GPT-6 Astra requires significantly less manual refinement to achieve a professional-grade appearance.

Testing evaluates aesthetic generation for a travel website landing page without functional constraints. GPT-6 Astra generates clean hero sections, integrates appropriate imagery autonomously via internal capabilities, and organizes content cleanly. Claude Fable 5.1 relies on basic blue gradients and generic layouts that require substantial subsequent prompting to become usable.

Motion Graphics and 3D Globe Web App Test

  • Both models achieve identical performance in motion graphics when connected to the Higgs field MCP.
  • GPT-6 Astra delivers a cleaner and more professional 3D globe travel dashboard than Claude Fable 5.1.
  • Claude Fable 5.1 prioritizes visual spectacle and bright animations over clean user interface design.

Final tests connect both models to external tools for motion graphics generation and require a 3D globe travel dashboard application. Both models tie in the motion graphics explainer task. In the 3D globe application test, GPT-6 Astra produces a clean, professional itinerary interface with time-zone features, while Claude Fable 5.1 creates an overly bright, blurry visual spectacle that obscures usability.

Community Posts

View all posts