Transcript
00:00:00So we got a brand new Opus model.
00:00:02We've gone all the way from Claude 5 and jumped to 5.5.
00:00:06But what sort of changes do you need to know about?
00:00:08Well, the highlight is that Anthropic is saying Claude Opus 5.5
00:00:12essentially operates at the same level as Fable 5.1, but costs 40% less.
00:00:20So we're basically getting a cheaper Fable 5.1.
00:00:24Now to start, let's take a look at the benchmarks.
00:00:26We have Opus 5.5 compared to Fable 5.1, Opus 5, as well as GPT-6 Astra and 5.6 Sol.
00:00:33Across the board, Opus 5.5 is showing numbers that beat pretty much everybody.
00:00:38The only two places it loses are to GPT-6 Astra,
00:00:42and that's on business workflows and agentic scientific research.
00:00:45On the business workflows benchmark, it loses by 1.4%.
00:00:50And on the scientific research benchmark, it loses by about 6%.
00:00:54But when we compare Opus 5.5 to Fable 5.1, and again, it is 40% cheaper,
00:01:00it beats it pretty much everywhere.
00:01:01And on top of that, we see huge leaps when we compare it to the Opus 5 model
00:01:04and pretty much everything related to coding.
00:01:07But when we look at benchmarks, it doesn't really do us much good
00:01:09just to see the performance levels.
00:01:11We need to see what it looks like at different effort levels
00:01:14and what it looks like in terms of cost per task.
00:01:16Because in theory, we can say, okay, per million tokens, Opus costs this,
00:01:21but that doesn't matter if it uses way more tokens to complete a task.
00:01:24So what we have right here is the Terminal Bench 4.0,
00:01:26and we have Opus 5.5 in the red.
00:01:29We have Fable 5.1 in the green, and Astra over here in the gray.
00:01:32Now, purely in terms of performance, we see that Opus 5.5 is beating pretty much everybody,
00:01:37especially when we look at high, extra high, and max.
00:01:40And in terms of price, it also looks like it's fairly efficient.
00:01:44When we look at it at high, we're scoring 64.2% at $3.88.
00:01:50For Astra on high, it's $7.21.
00:01:53So, you know, almost twice as expensive.
00:01:56And then when we compare it to Fable, 5.1 on high, not only does it perform way worse,
00:02:0149.4% versus 64.2, it's also way more expensive.
00:02:06At $10.50, again, compared to $3.88.
00:02:10So, in general, Opus 5.5 looks like it performs better and is more efficient.
00:02:17And when we look at the Frontier Code and Cursor Bench benchmarks,
00:02:20it kind of plays out the same way.
00:02:22It's performing better at virtually every single effort level,
00:02:25and at each effort level, it also does it cheaper.
00:02:28Now, when we look at these benchmarks, what we do kind of see,
00:02:31and we've noticed this across a whole lot of different models,
00:02:34is that we tend to get diminishing returns as we go from high all the way to extra high and max.
00:02:39When we compare highs, output 64.2 versus 64.8, pretty much the same,
00:02:45yet the cost is almost four times as much.
00:02:49We see the same thing with Cursor Bench.
00:02:51Again, high gives us pretty much max outputs at a way cheaper cost.
00:02:54And it's even more exaggerated on the Frontier Code set,
00:02:57where medium effort level actually gives us the best performance at 54.6 compared to max at 54.4,
00:03:05yet it is six, seven times cheaper, which is wild.
00:03:09I mean, this one's really crazy because Opus 5.5 on medium outperforms every single model out there
00:03:14while being way cheaper than every other model's low setting, which is crazy.
00:03:18And when we take a look at the knowledge work benchmarks, again, we see pretty much the same thing play out,
00:03:23the one exception being the automation bench, where Astra does pull ahead.
00:03:27Now, in terms of its straight up token cost per million, Opus 5.5 is $4 per input and $20 per output,
00:03:35which is cheaper than Opus 5.
00:03:37On top of that, we have cheaper cash reads and writes, which is huge.
00:03:40For reference, Astra and Fable 5.1, their input is $10 and their output is $50.
00:03:46So if these benchmarks are to be taken at face value, then we have an extremely efficient model on our hands.
00:03:51Now, one of the big issues with Opus 5 was how it communicated.
00:03:55We pretty much had the Opus slot mean running around, where it was extremely jargon heavy,
00:04:00and it was difficult to even understand what it was saying.
00:04:03This is something Anthropic has heard, and they're telling us that they have solved this problem with Opus 5.5,
00:04:09making it a much better collaborator, because we didn't actually understand what it's saying.
00:04:13Specifically, they say it is less likely to use jargon or idiosyncratic phrases,
00:04:17and it follows the writing rules you give it.
00:04:20And they have some comparisons here, where it shows what Opus 5 would tell us,
00:04:23versus what Opus 5.5 would tell us, whether that's a bug, or summarizing a thread,
00:04:27or explaining a design change.
00:04:29And again, if you've used Opus 5 at all, you understand.
00:04:32It is very annoying to actually talk to it.
00:04:34And lastly, we have safeguards.
00:04:36Anthropic is telling us that Opus 5.5 pretty much falls in line
00:04:38with the Fable series in terms of like cybersecurity and biology type questions.
00:04:43So if it deems that you're asking questions about cybersecurity that you're going to use
00:04:47in a negative fashion, like you're trying to use it to like ask somebody,
00:04:52it's going to reroute your question.
00:04:53And when it reroutes it, it's going to reroute it to Opus 4.8.
00:04:57This also applies to biology.
00:04:59So in general, if you can't ask Fable 5.1 this question,
00:05:02it reroutes you with Fable.
00:05:03It's going to reroute you when you ask the same question to Opus 5.5.
00:05:07Also similar to Fable is how Opus 5.5 handles distillation.
00:05:11It stopped API users from editing Claude's prior context in an attempt to extract Claude's reasoning.
00:05:16So same exact sort of method to stop these attacks.
00:05:21And lastly, in terms of availability, it is available everywhere today.
00:05:25So on paper, what's not to love here?
00:05:27We have a model that is just as good as Fable and is cheaper.
00:05:31Now, will it hold up in reality?
00:05:33Well, it's just going to take time.
00:05:34We're going to have to play around to see if that's the case.
00:05:36But if you've been on Twitter at all lately,
00:05:39you have seen some of the outputs people have been getting with Opus
00:05:42when it's been silently routed to Opus 5.5 over the last week or so.
00:05:46And the results have been impressive.
00:05:48So I'm excited for this model because honestly,
00:05:51it is a pain in the butt to use the Anthropic plans,
00:05:54especially when they have reduced our usage of late.
00:05:56And we can only use half of our weekly allowance on Fable.
00:05:59And so your second half almost has felt wasted with Opus 5.
00:06:02But if I can pretty much get Fable 5.1 and then Opus 5.5,
00:06:06which is basically Fable for the other 50%,
00:06:08well, then I'm not going to complain.
00:06:10And that's big because for a lot of people, it's been like,
00:06:12well, why don't we just go over to OpenAI and chat GPT and start using Astra?
00:06:15Because they don't limit us in that sense.
00:06:17Well, if you're able to get a $200 plan, that is.
00:06:20So really excited for this one.
00:06:22Definitely check it out and let me know what you think.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video