스크립트
00:00:00So right after Anthropic released Fable 5.1, OpenAI strikes back and releases GPT-6 Astra.
00:00:07Now let's kick this off by looking at some of the benchmarks for GPT-6 Astra. And shout out to
00:00:10OpenAI for giving us a lot more benchmarks to look at than what Anthropic showed us with Fable
00:00:145.1. When it comes to computer use, pretty much runs the table. We don't even have a lot of data
00:00:19here when it comes to Fable 5.1. When we look at the professional benchmarks, things like BenchCAD
00:00:24and browse comp. Again, GPT-6 does amazing. The only place where it doesn't win is in the
00:00:31artificial analysis intelligence index, where it scores a 61.2 compared to Fable 5.1, 65.7.
00:00:38Next, we have coding. And again, kind of the same story. That's either beating Fable 5.1 or it's
00:00:43neck and neck. If we look specifically at TerminalBench 4.0, we see a huge leap from GPT 5.6 solo
00:00:49from 37.3 over to 57.7. And if we look at DeepSuite, which is probably my favorite coding
00:00:55benchmark, it really focuses on long-running agentic tasks. It scores a 74.1%, which again,
00:01:02is better than everything else out there. GPT-6 also crushes the academic benchmarks as well as the
00:01:08science and health ones. And when we take a look at cybersecurity benchmarks, again, really I'm paying
00:01:13attention to 5.6 solo here. Huge leaps, 78.5 to 100%. 55.9 all the way to 88. 5.5% on the exploit
00:01:22bench all the way up to 39%. So this is a step change from the five series models. And in terms
00:01:27of long context, GPT-6 also kills it. 100% score when doing the eight needle test from 256 to 512k
00:01:34token. So essentially a quarter to a completely full context window. And then when that window has 512,000
00:01:41tokens to a million, it still scores a 96.3%. So if you're someone who's always throwing a ton of
00:01:46context at your AI model of choice, GPT-6 is going to do really well in those scenarios. But the raw
00:01:51benchmark scores aren't enough. We need to know how much it costs to actually achieve these scores
00:01:55because Astra may have the best score in the world, but if it costs 10x everything else, what's the
00:01:59point? Now, luckily for the GPT models, historically, they've been very token efficient. So when we look
00:02:04here at the terminal bench 4.0 score, we see that kind of plays out. So GPT-6 Astra here with the stars,
00:02:11first thing I want to note is that like many of these models, we see sort of a drop in efficiency
00:02:16as we go higher up the effort level chain. So when I go from low to medium to high, I continue to get
00:02:23better scores at somewhat a linear cost increase. But as I move from high to extra high, and then max,
00:02:30in fact, I'm actually getting a worse score at a higher cost. So at high for $7.21, I'm getting a
00:02:3957.9% accuracy. When I compare this to Fable 5.1 right here, and I'm on high, well, I'm getting a 49.4%
00:02:49accuracy at $10.50. So more expensive, and it doesn't score as well. Now, Fable does score pretty high
00:02:57when I put it all the way to max effort 55.8, but it's costing me $19.50. So for a same score,
00:03:06again, about 55%, it's almost $20 of Fable 5. And for GPT-6 Astra, it's $7.20. So infinitely cheaper
00:03:16for essentially the same accuracy. And when we look at Frontier Code 1.1, we see a similar pattern,
00:03:22where at the high ends, we are getting similar scores. When we look at GPT-6 versus Fable 5.1
00:03:29or Fable 5. In this case, Fable 5 actually scored higher here. But it's just way more expensive to
00:03:34run these cloud models due to the token efficiency with GPT-6 Astra. Now, there are a few functionalities
00:03:39beyond the benchmarks that OpenAI calls out when it comes to GPT-6. The first is that they're calling it
00:03:43the world's best computer use model. Now, there are a few sort of benchmarks that test this sort of
00:03:48thing, things like the agent's last exam. And like we've seen with other benchmarks, what does it show
00:03:52us? It shows Astra kind of crushing it, both in terms of its accuracy and the API costs, which can
00:03:57never be discounted. But arguably even more important than the accuracy is the speed. Codex actually is
00:04:01really good computer use, especially if you use it in conjunction with the voice mode. And Astra is
00:04:06supposed to be 1.9 times faster than GPT-5.6, which is great if you're someone who's been using that a
00:04:11whole lot over the last month or so. The other thing they call out is adherence to templates. So if you
00:04:15give it some sort of template of, say, a slide deck, like you see here on the left, and you ask it to
00:04:20create another slide deck in that same fashion about a different topic, well, you get something
00:04:25like we see on the right, something that adheres to the style very, very well. So if you have some
00:04:31sort of template that you use over and over, whether it's for slide decks or websites, think of it as a
00:04:35design system for really any sort of output, GPT-6 is great for that. You can see that again here with
00:04:41Excel documents and document styling in general. Here was essentially the reference image it was given.
00:04:46Here's what we would get before, and here's what we get on the right after with it looking at that
00:04:51reference image. Another interesting thing is this line here where they talk about GPT-6 Asher
00:04:55bringing stronger visual judgment to websites, gains, applications, and rendering it built, aka they're
00:05:00saying GPT-6 has better taste. And you've probably heard it over and over again that AI has no taste.
00:05:05Well, they're saying, well, GPT-6 does. Now, let's be honest, there's always going to be issues when it
00:05:10comes to AI in taste, because if you give it a bad prompt, it's going to have some sort of regression
00:05:15to the mean, no matter what. And whatever that mean is, people are going to associate that with AI
00:05:20slop, no matter how good it really is. But they show a few examples here of essentially like an Unreal
00:05:25Engine walkthrough they did, Blender modeling, as well as some still images here. Now, a really cool thing
00:05:31they're adding with Astra is changes to how it does auto compaction when the context window fills. Now,
00:05:37usually you fill up your context window, it's going to auto compact, it's going to create like a summary
00:05:41of everything you've been talking about in that last session, and then start a new session. Well,
00:05:46that creates issues, sometimes it leaves out details. But now, Astra keeps notes across context
00:05:52windows, instead of just having a single summary. So instead of one document, it essentially has a bunch
00:05:57of different sticky notes about what it thinks is important. And it can reference those at any time
00:06:01versus it kind of just being a one shot of like, here's a summary, hope this is everything we need.
00:06:06And in fact, earlier context windows remain searchable. So again, it isn't like this is the
00:06:12only document we have in terms of a summary. And if it's not in here, we're screwed. Just not the case
00:06:16anymore. Now, this is an experimental feature. So you're going to have to enable it in the codex
00:06:20config, but it will become the default in a few weeks. Now, when it comes to cybersecurity,
00:06:24Astra is treating this similar to how Anthropic treats Fable, where they're like, this model is so
00:06:28powerful, it can create a bunch of exploits. So if it thinks you're trying to create some sort of
00:06:33exploit, it's just not going to let you do it. It's not going to adhere, it's not going to answer
00:06:37your problem. Another great improvement in GPT-6 compared to the 5 series is a lower hallucination
00:06:42rate. GPT-5.6 Sol and like the 5.6 models in general actually hallucinated quite a bit compared
00:06:48to the other frontier models. Yet head to head, we see we've gone from something like a 9.4%
00:06:54hallucination rate all the way down to a 2% hallucination rate. Now is Astra available? Well,
00:07:01it's open to a limited set of organizations. And in the next few days, it will be open to essentially
00:07:06everybody. As for the API, it is available for developers. And in regards to price, again,
00:07:11matches Fable pricing. We're looking at $10 per million input tokens and $50 per million output tokens.
00:07:17So based on the numbers, OpenAI is giving us a very solid model that can totally compete
00:07:22with the best of what Anthropic has. And they're doing it at a lower price point. So
00:07:26I'm super excited to try this out. I think having competition amongst the biggest players here
00:07:30is ultimately great for us, the end user.