Opus 5 Is Here And Its BETTER THAN FABLE?!
CChase AI
Computing/SoftwareBusiness NewsInternet Technology
Transcript
00:00:00So Claude Opus 5 just dropped in, and Anthropic is claiming we now have a model that is getting
00:00:04close to Fable 5 outputs at half the cost. So let's take a look at what they're giving us.
00:00:09So let's take a look at some of the benchmarks, and right away, kind of wild numbers. We see that
00:00:14it beats Fable 5 and blows Opus 4.8 out of the water on a number of these tests, specifically
00:00:20agentic terminal coding, knowledge work, agentic search. Also, it's beating things like GPT's 5.6
00:00:27in these categories as well. The only areas that we see Opus 5 doing worse than Fable 5 is
00:00:33multidisciplinary reasoning, which isn't by much. In fact, it beats it when it uses tools, as well as
00:00:39legal and health. But every other spot, it's actually beating their top model. And if those numbers are to
00:00:46be believed, that's crazy. Right here, it's saying on cursor bench 3.2 at max effort, the model performed
00:00:51within 0.5% of Fable 5's peak score, but at half the cost per task. Nuts. And you can see that right
00:00:57here. We have Opus 5 in the red, and then Fable 5 in the orange, and down here in the blue is Opus 4.8.
00:01:03So yes, at max, max effort, Fable 5 does pull ahead, but most people aren't actually operating
00:01:08at max. They're either operating at high or extra high. And when we look at extra high and high
00:01:13compared to Fable, we're pretty much getting the same outputs with Opus 5, but again, at like
00:01:17half the cost. And in some of these other benchmarks, like artificial analysis coding agent index,
00:01:23Opus 5 is just straight up beating Fable 5. And same thing with the frontier bench. Like
00:01:27it's, it's actually whooping Fable and it's cheaper. Now let's take a look at some knowledge
00:01:32work and problem solving tasks. And again, I think what's really crazy isn't just this whole
00:01:37like Fable 5, Opus 5 dynamic. It's the jump from Opus 4.8 to Opus 5. Like all of these
00:01:43across the board, just a massive, massive leap forward. And again, if this was like a Sonnet
00:01:495 thing where we saw it doing better than the previous model, yet it was just like prohibitively
00:01:53expensive because Sonnet 5 is somehow the most expensive model out there when we break it down
00:01:57by task. That's not the case here. Opus looks to be extremely token efficient compared to these
00:02:03two other models. And again, like the fact it's actually beating Fable 5, I really can't get
00:02:07over that. Another interesting improvement is visual output. So here's a look at Opus 5 doing
00:02:13some wind tunnel work. And then over here, we have it building a 3D interactive animal cell
00:02:19artifact. Now, one of the downsides with Cloud is it like can't create its own images and stuff. So
00:02:24it being able to do better with these sort of like 3D created SVG, almost like HTML graphics is huge.
00:02:31Another big improvement is that Anthropic is saying it's much stronger in verifying its work
00:02:35and iterating carefully until it succeeds. What does this mean? This means when we use Opus 5
00:02:40on long horizon agentic tasks, things that just aren't one shot. These are like loops, right? This
00:02:45is like, you know, the whole graph engineering thing people are talking about. Anything that requires
00:02:49Claude to essentially check its work against some sort of goal we give it and keep doing that over and
00:02:53over and over again. Well, Opus 5 is much better than Opus 4.8 in that regard. And one interesting
00:02:58example here is Opus 5 was given a drawing of a machine part and told to write code to rebuild it as a
00:03:033D free CAD model. But the model wasn't given any way to actually view the drawing. So, you know,
00:03:09it was told to reach this goal. It wasn't told specifically how to do it. So it went ahead and
00:03:14wrote its own computer vision pipeline to pull the geometry from raw pixels and then reconstructed the
00:03:19whole machine part. Again, anything that is long horizon, given some sort of North Star to reach for,
00:03:25Opus 5 is much better than 4.8 in actually completing that. In terms of misaligned behavior,
00:03:29it's also a step above all these previous models. Lower is better here. And we see a huge increase in its capability
00:03:35from Opus 4.8 to find vulnerabilities and exploits in open source code. Now the other cool thing is the
00:03:40safeguards. This isn't a Fable scenario where if you mention anything about cybersecurity biology,
00:03:45it's going to get shut down. Those guardrails are much less strict with Opus 5 versus Fable 5.
00:03:50Now, while there are still some classifiers in there, they say it's 85% less likely that they actually
00:03:56trigger versus Fable. Now, lastly, the big thing is cost. Like I said, this is half the price of Fable.
00:04:01We're looking at $5 per million input and 25 per million output. And like we saw from the benchmarks
00:04:07in the charts, this is a relatively token efficient model. So I'm really excited to try this one out
00:04:12because if the numbers are true, we kind of almost have a better Fable for half the cost.