GPT 6 Astra Is Here (And It's Better Than Fable 5.1?)

CChase AI
Computing/SoftwareInternet Technology

Transcript

00:00:00So right after Anthropic released Fable 5.1, OpenAI strikes back and releases GPT-6 Astra.
00:00:07Now let's kick this off by looking at some of the benchmarks for GPT-6 Astra. And shout out to
00:00:10OpenAI for giving us a lot more benchmarks to look at than what Anthropic showed us with Fable
00:00:145.1. When it comes to computer use, pretty much runs the table. We don't even have a lot of data
00:00:19here when it comes to Fable 5.1. When we look at the professional benchmarks, things like BenchCAD
00:00:24and browse comp. Again, GPT-6 does amazing. The only place where it doesn't win is in the
00:00:31artificial analysis intelligence index, where it scores a 61.2 compared to Fable 5.1, 65.7.
00:00:38Next, we have coding. And again, kind of the same story. That's either beating Fable 5.1 or it's
00:00:43neck and neck. If we look specifically at TerminalBench 4.0, we see a huge leap from GPT 5.6 solo
00:00:49from 37.3 over to 57.7. And if we look at DeepSuite, which is probably my favorite coding
00:00:55benchmark, it really focuses on long-running agentic tasks. It scores a 74.1%, which again,
00:01:02is better than everything else out there. GPT-6 also crushes the academic benchmarks as well as the
00:01:08science and health ones. And when we take a look at cybersecurity benchmarks, again, really I'm paying
00:01:13attention to 5.6 solo here. Huge leaps, 78.5 to 100%. 55.9 all the way to 88. 5.5% on the exploit
00:01:22bench all the way up to 39%. So this is a step change from the five series models. And in terms
00:01:27of long context, GPT-6 also kills it. 100% score when doing the eight needle test from 256 to 512k
00:01:34token. So essentially a quarter to a completely full context window. And then when that window has 512,000
00:01:41tokens to a million, it still scores a 96.3%. So if you're someone who's always throwing a ton of
00:01:46context at your AI model of choice, GPT-6 is going to do really well in those scenarios. But the raw
00:01:51benchmark scores aren't enough. We need to know how much it costs to actually achieve these scores
00:01:55because Astra may have the best score in the world, but if it costs 10x everything else, what's the
00:01:59point? Now, luckily for the GPT models, historically, they've been very token efficient. So when we look
00:02:04here at the terminal bench 4.0 score, we see that kind of plays out. So GPT-6 Astra here with the stars,
00:02:11first thing I want to note is that like many of these models, we see sort of a drop in efficiency
00:02:16as we go higher up the effort level chain. So when I go from low to medium to high, I continue to get
00:02:23better scores at somewhat a linear cost increase. But as I move from high to extra high, and then max,
00:02:30in fact, I'm actually getting a worse score at a higher cost. So at high for $7.21, I'm getting a
00:02:3957.9% accuracy. When I compare this to Fable 5.1 right here, and I'm on high, well, I'm getting a 49.4%
00:02:49accuracy at $10.50. So more expensive, and it doesn't score as well. Now, Fable does score pretty high
00:02:57when I put it all the way to max effort 55.8, but it's costing me $19.50. So for a same score,
00:03:06again, about 55%, it's almost $20 of Fable 5. And for GPT-6 Astra, it's $7.20. So infinitely cheaper
00:03:16for essentially the same accuracy. And when we look at Frontier Code 1.1, we see a similar pattern,
00:03:22where at the high ends, we are getting similar scores. When we look at GPT-6 versus Fable 5.1
00:03:29or Fable 5. In this case, Fable 5 actually scored higher here. But it's just way more expensive to
00:03:34run these cloud models due to the token efficiency with GPT-6 Astra. Now, there are a few functionalities
00:03:39beyond the benchmarks that OpenAI calls out when it comes to GPT-6. The first is that they're calling it
00:03:43the world's best computer use model. Now, there are a few sort of benchmarks that test this sort of
00:03:48thing, things like the agent's last exam. And like we've seen with other benchmarks, what does it show
00:03:52us? It shows Astra kind of crushing it, both in terms of its accuracy and the API costs, which can
00:03:57never be discounted. But arguably even more important than the accuracy is the speed. Codex actually is
00:04:01really good computer use, especially if you use it in conjunction with the voice mode. And Astra is
00:04:06supposed to be 1.9 times faster than GPT-5.6, which is great if you're someone who's been using that a
00:04:11whole lot over the last month or so. The other thing they call out is adherence to templates. So if you
00:04:15give it some sort of template of, say, a slide deck, like you see here on the left, and you ask it to
00:04:20create another slide deck in that same fashion about a different topic, well, you get something
00:04:25like we see on the right, something that adheres to the style very, very well. So if you have some
00:04:31sort of template that you use over and over, whether it's for slide decks or websites, think of it as a
00:04:35design system for really any sort of output, GPT-6 is great for that. You can see that again here with
00:04:41Excel documents and document styling in general. Here was essentially the reference image it was given.
00:04:46Here's what we would get before, and here's what we get on the right after with it looking at that
00:04:51reference image. Another interesting thing is this line here where they talk about GPT-6 Asher
00:04:55bringing stronger visual judgment to websites, gains, applications, and rendering it built, aka they're
00:05:00saying GPT-6 has better taste. And you've probably heard it over and over again that AI has no taste.
00:05:05Well, they're saying, well, GPT-6 does. Now, let's be honest, there's always going to be issues when it
00:05:10comes to AI in taste, because if you give it a bad prompt, it's going to have some sort of regression
00:05:15to the mean, no matter what. And whatever that mean is, people are going to associate that with AI
00:05:20slop, no matter how good it really is. But they show a few examples here of essentially like an Unreal
00:05:25Engine walkthrough they did, Blender modeling, as well as some still images here. Now, a really cool thing
00:05:31they're adding with Astra is changes to how it does auto compaction when the context window fills. Now,
00:05:37usually you fill up your context window, it's going to auto compact, it's going to create like a summary
00:05:41of everything you've been talking about in that last session, and then start a new session. Well,
00:05:46that creates issues, sometimes it leaves out details. But now, Astra keeps notes across context
00:05:52windows, instead of just having a single summary. So instead of one document, it essentially has a bunch
00:05:57of different sticky notes about what it thinks is important. And it can reference those at any time
00:06:01versus it kind of just being a one shot of like, here's a summary, hope this is everything we need.
00:06:06And in fact, earlier context windows remain searchable. So again, it isn't like this is the
00:06:12only document we have in terms of a summary. And if it's not in here, we're screwed. Just not the case
00:06:16anymore. Now, this is an experimental feature. So you're going to have to enable it in the codex
00:06:20config, but it will become the default in a few weeks. Now, when it comes to cybersecurity,
00:06:24Astra is treating this similar to how Anthropic treats Fable, where they're like, this model is so
00:06:28powerful, it can create a bunch of exploits. So if it thinks you're trying to create some sort of
00:06:33exploit, it's just not going to let you do it. It's not going to adhere, it's not going to answer
00:06:37your problem. Another great improvement in GPT-6 compared to the 5 series is a lower hallucination
00:06:42rate. GPT-5.6 Sol and like the 5.6 models in general actually hallucinated quite a bit compared
00:06:48to the other frontier models. Yet head to head, we see we've gone from something like a 9.4%
00:06:54hallucination rate all the way down to a 2% hallucination rate. Now is Astra available? Well,
00:07:01it's open to a limited set of organizations. And in the next few days, it will be open to essentially
00:07:06everybody. As for the API, it is available for developers. And in regards to price, again,
00:07:11matches Fable pricing. We're looking at $10 per million input tokens and $50 per million output tokens.
00:07:17So based on the numbers, OpenAI is giving us a very solid model that can totally compete
00:07:22with the best of what Anthropic has. And they're doing it at a lower price point. So
00:07:26I'm super excited to try this out. I think having competition amongst the biggest players here
00:07:30is ultimately great for us, the end user.

Key Takeaway

GPT-6 Astra achieves higher benchmark accuracy and lower hallucination rates than competing models while maintaining competitive API pricing of $10 per million input tokens.

Highlights

  • GPT-6 Astra scores 74.1% on the DeepSuite long-running agentic coding benchmark.

  • TerminalBench 4.0 accuracy for GPT-6 Astra reaches 57.7% compared to 37.3% for the previous solo model.

  • Cybersecurity benchmark scores on exploit bench increase from 5.5% in the 5 series to 39% in GPT-6 Astra.

  • Hallucination rates drop from 9.4% in the previous 5 series to 2% in GPT-6 Astra.

  • API pricing for GPT-6 Astra matches Fable pricing at $10 per million input tokens and $50 per million output tokens.

Timeline

Benchmark Performance and Capabilities

  • GPT-6 Astra dominates professional, coding, academic, and science benchmarks.
  • DeepSuite scores reach 74.1% for long-running agentic tasks.
  • Cybersecurity exploit bench accuracy rises to 39%.

GPT-6 Astra surpasses previous models across multiple evaluation domains. TerminalBench 4.0 jumps from 37.3% to 57.7%. Long context tests maintain a 100% score on the eight needle test up to 512k tokens and 96.3% up to one million tokens.

Cost Efficiency and Effort Levels

  • Higher effort levels beyond high yield worse scores at increased costs.
  • GPT-6 Astra provides 57.9% accuracy at $7.21 on TerminalBench 4.0.
  • Competing models cost nearly $20 for equivalent accuracy levels.

Token efficiency remains high for GPT-6 Astra, though pushing effort levels past high introduces cost inefficiencies and diminished accuracy returns. Compared to Fable 5.1 at max effort, GPT-6 Astra delivers similar accuracy at a fraction of the cost.

Computer Use, Styling, and Context Management

  • GPT-6 Astra operates as a computer use model with speeds 1.9 times faster than GPT-5.6.
  • Context window updates utilize sticky notes across sessions instead of single summaries.
  • API access is available immediately for developers with broad rollout arriving shortly.

New features include strong template adherence for slide decks and documents, along with improved visual judgment for rendered applications. Experimental auto compaction retains searchable notes across context windows to prevent detail loss.

Community Posts

View all posts