Sonnet 5 is LIVE And It Competes With Opus
CChase AI
Computing/SoftwareSmall Business/StartupsManagement
Transcript
00:00:00Anthropic just released Sonnet 5, so let's run through everything you need to know about it.
00:00:04So let's begin with the benchmarks. And remember, Sonnet 5 is a less powerful,
00:00:08less expensive model than things like Opus and definitely Fable. We haven't had an upgrade here
00:00:13since 4.6. Now, comparing 5 to 4.6, we see this is a huge leap forward in terms of agentic coding,
00:00:21multidisciplinary reasoning, computer use, and knowledge work. So straight upgrade across the
00:00:26board. Now, when we compare it to Opus 4.8, the question isn't, is it close? It's how much of a
00:00:32falloff do we have? And honestly, it's not that much of a falloff. In fact, it boasts better numbers
00:00:39when we see knowledge work, which is kind of crazy. The computer use is only 2% behind, and it's only
00:00:442% behind agentic coding, and only a handful of percentage points behind it on multidisciplinary
00:00:49reasoning. The biggest gap is over here with Sweebench Pro when we look at 63 versus 69%.
00:00:56But hey, Terminal Bench 2.1, 80 versus 82. So this isn't much of a gap. And this discussion
00:01:03about performance has to be done in the context of the pricing. Remember, Fable 5, Mythos 5,
00:01:08we're looking at $10 per input tokens, which is double Opus. Opus is $5 per million input tokens.
00:01:16And for output, it's at $25. So $5, $25. For Sonnet 5, we're looking at $2. So only 40% of the cost of
00:01:26Opus in input tokens. And for output, it's $10. So significantly cheaper, less than half the cost of
00:01:34Opus. Yet when we look at the numbers, it's pretty dang close. So this is huge news if you're somebody who
00:01:40isn't doing necessary bleeding edge tasks with AI, you're doing it for more routine tasks, and you
00:01:46still want power, well, Sonnet 5 is now plugging that gap. Now let's look at these charts that go
00:01:51beyond just the benchmarks. Here we're looking at agentic search performance by effort level, comparing
00:01:55Opus 4.8 to Sonnet 5 and Sonnet 4.6. So right away, I think you'll notice here, Sonnet 5, there's a huge
00:02:04gap in the pass rates as we move between the effort levels. So on low, it does 55%. But if we look at
00:02:13low on Sonnet 4.6, low on Sonnet 4.6 actually performs better. Yet in terms of cost, it's way less.
00:02:19We move over to medium, we're essentially getting the same performance that we saw with Sonnet 4.6,
00:02:25but at a significant discount. And it's only until we move to high effort levels, do we begin to surpass
00:02:34Sonnet 4.6. But at that point, we're getting better performance than Sonnet 4.6, but at a cost that
00:02:41would have been on the low effort level for Sonnet 4.6. So this sort of effort level sort of distinction,
00:02:49I think is very important because it's not just, hey, Sonnet 5 is better across the board, where low
00:02:53on 5 is necessarily better than low on 4.6. Low on Sonnet 5 really just means this is going to be super
00:02:59cheap. But in terms of getting a higher level performance, what we've seen in the past,
00:03:03probably need to do something that's in the high range. That being said, when we're on the high
00:03:08range with Sonnet 5, we're pretty much paying the same amount we would be with Opus. Opus on medium
00:03:14and high costs the same as Sonnet on high, but we're getting a better pass rate. So this sort of brings
00:03:21a lot of nuance into this discussion. Do we need to use Sonnet 5 versus Opus 4.5 and things like
00:03:27Egentic Search? I think it really depends. If it's a more complex problem, in the end, Opus 4.8 may be
00:03:33cheaper just because it is so token efficient and Sonnet just isn't the right thing for the job.
00:03:38However, when we're doing tasks that just aren't as challenging for these AI models, then perhaps
00:03:44it doesn't make sense to use Opus 4.8 at all. That's when we bring in Sonnet 5. So I think this
00:03:49is an interesting place we're in now where it's kind of a case-by-case basis as for which model
00:03:54you should use. Now, here's another interesting one when we see Egentic computer use. Now, across the
00:03:58board, for the most part, Sonnet 5 is beating Sonnet 4.6 and it's doing it at a pretty low cost. So
00:04:05Egentic Search, that definitely wasn't the case, but Egentic computer use, all of a sudden,
00:04:09we would always use Sonnet 5 over 4.6. And again, this goes back to the idea that it's now a case-by-case
00:04:14basis as for which model you should use. Interesting enough, though, look at Opus. Opus High performs
00:04:20better than Mac's Sonnet 5 and it's cheaper. So, I mean, you kind of look at these and you're kind of like,
00:04:26wow, like, Opus is actually pretty efficient for what it does. Not to mention it kind of hits these higher
00:04:30levels that Sonnet just can't reach. So things to think about, honestly, just because the cost per million
00:04:36tokens is cheaper with Opus, there's going to be a lot of cases where Opus is still the best. Now, I want to be
00:04:41an anthropic model release without talking about misaligned behavior. We can see it's an improvement
00:04:45with Sonnet 5, but still below what we would expect with Opus 4.8 and Mythos Preview. And in terms of
00:04:50exploits and this whole cybersecurity question that got Mythos pretty much locked up, that's not a
00:04:55problem or something you need to worry about with the Sonnet class. So overall, I wouldn't pay too much
00:04:59attention to these benchmarks, to be honest. These sort of benchmarks, these sort of charts give me a lot
00:05:05more pause because the question I don't think is Sonnet 5 versus 4.6. We're pretty much always going to
00:05:11use Sonnet 5. It's Sonnet 5 versus Opus 4.8. And the question is, is Opus actually cheaper? And
00:05:18I would love to see more tests and more stuff by anthropic that kind of spells out here's sort of
00:05:24that like line of demarcation between, okay, Sonnet 5 can do just fine on this and it's cheaper than
00:05:29Opus versus, hey, this is now a more complicated task. Opus is actually going to be more efficient than
00:05:34Sonnet at completing this. So things to think about, definitely a nice to have. I think overall lower
00:05:40level tasks on it is going to be a boon because for those of us who do use API tokens for, you know,
00:05:46our applications or things that aren't in the Claude Max Plan ecosystem, it's kind of tough to use
00:05:51Anthropic. It's just so damn expensive. And I've been telling people for a long time, like, just go look
00:05:55at OpenAI, you know, look at some of their mini models because that's often more than enough for a lot
00:05:59of people's jobs. Well, we now kind of have a middle tier option with Anthropic, which is a nice
00:06:05to have. So that's where I'm going to leave you guys today. This is available everywhere. I think
00:06:09it's going to be a lot of trial and error figuring out where this makes the most sense. So let me know
00:06:13what you thought. Make sure to check out Chase AI Plus if you want to get your hands on my Claude
00:06:17Code Masterclass and I'll see you around.