Transcript
00:00:00Anthropic just released Claude Opus 5, and it might actually be better than Fable while being
00:00:04half the price. If that is true, this is my new favourite model, so let's find out.
00:00:13So this is how they're describing Opus 5. Opus 5 is a thoughtful and proactive model that comes
00:00:18close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work
00:00:23evaluations like Frontier Bench and GDP Val, Opus 5 is the new state-of-the-art that remains
00:00:28behind Mythos 5 on cybersecurity tasks. Opus 5 is designed to be used every day, it works more
00:00:33efficiently than other models, and it's the new default model on Claude Max and the strongest model
00:00:38on Claude Pro plans. After this, they then go on to show the benchmarks, and it is winning on a lot of
00:00:43these, despite the fact that Fable 5 is here, and it's not even winning by small amounts. According
00:00:47to them, it's nearly 10% better on Frontier Bench than Fable 5, but one of the craziest ones to me
00:00:52is its score on Arc AGI. Opus 4.8 scored 1.5% here, whereas Sol scored 7.8%, which was the previous high
00:01:00score, but Opus 5 scores 30.2%. The team behind this benchmark actually showed that Opus had a new
00:01:06capability that they hadn't seen before in any other models, where it would turn these tests into
00:01:10algebraic notation to perform advanced logical reasoning. This is seriously impressive stuff.
00:01:15If you're wondering, by the way, why Fable isn't on here, that's because of Fable's data retention
00:01:19policy. Arc AGI don't want to risk leaking any details of that benchmark to Anthropic, but the
00:01:24good news is, Opus 5 does not have data retention requirements for general access. That is going to
00:01:29make a lot of people very happy. Moving on though, let's take a look at the third-party benchmarks by
00:01:33Artificial Analysis. On the coding index, it comes in second, losing only by 0.3 points to GPT 5.6 Sol
00:01:39Extra High, but it does beat Fable by 1.5 points, and it's a very nice jump above Opus 4.8.
00:01:45Shout out to Kimi K3 here, by the way, for breaking up the big guys. I did a video on that
00:01:49when it released, so you should definitely subscribe to stay up to date with the AI and
00:01:52developer news. If we then take a look at the agentic index, Opus 5 actually takes the lead
00:01:57here, being a nice small jump above Fable 5 and beating OpenAI's sole model. It's also the same
00:02:02story on the intelligence index, where Opus 5 is ahead there too. So on the performance side of
00:02:06things, it looks like Opus 5 is a seriously good model, and honestly, I'm pretty surprised. I kind of
00:02:11thought Fable was where they were going to focus on now, and that Opus would become the sort of new
00:02:15Sonic model, but now it seems that Fable and Opus are almost neck and neck, until we factor
00:02:19in pricing. If we take a look at the cost to complete a task on artificial analysis, Opus
00:02:235 actually comes in 72 cents cheaper than Fable 5, and it's only a little bit more costly than
00:02:28Opus 4.8. Now I would say if you were going to go for price to performance, GPT 5.6 Sol
00:02:33is still going to be the winner here, costing nearly half the price of Opus 5. You can see
00:02:37on this chart where green is the most attractive quadrant, the GPT models sit just about in there,
00:02:42along with Kimi K3, and Anthropic models kind of have their own section in the higher price
00:02:46side of things. Cursor Bench also mirrors all of these benchmarks that we've looked at. On
00:02:50here, Opus 5 Max has nearly the exact same score as Fable 5, but Fable costs $17 for that
00:02:56task, and Opus costs $8. The price itself, by the way, is the same as Opus 4.8, so it's
00:03:01$5 for a million input tokens, and $25 for a million output tokens, so if you're a fan of
00:03:06Opus 4.8, I don't think there's any reason why you wouldn't use this model, and as a fan
00:03:10of Fable 5 myself, I'm definitely going to be using Opus 5 a lot to save my subscription
00:03:14and my credits so they don't run out as often. Honestly, my best guess on Fable versus Opus
00:03:18is that Fable is still going to be the top model if you want really long horizon hard problems,
00:03:23but for 95% of the work, Opus 5 can now take its place. But now let's go ahead and actually
00:03:28see it in action with some local tests. The first test I did was to ask the models to create
00:03:32a complete Formula 1 racing game in a single HTML file using 3.js. Now, I'll talk about how long
00:03:37each model took as well as how much that would cost in the API at the end, but first, let's
00:03:40just see the results. This is the result that Opus 5 gave me, and I've got to say I am seriously
00:03:45impressed with the models here. It uses no external assets or anything like that itself. This is all
00:03:50done by Opus 5, and everything looks really good. Now, the track layout was a bit backwards there,
00:03:55but the handling is all good. The lap times work, the positioning works, the order works,
00:03:59all of that works well. Yes, we are driving under grass here, but you can see collision works,
00:04:03and there is actually a track underneath this grass, so I can very easily prompt that out.
00:04:07And also, the gravel traps seem to work really nice as well. Honestly, I am very impressed with
00:04:12the result that we got from Opus 5 here. These F1 cars look the best out of all of them, in my opinion.
00:04:16I'll let you decide, though. This is the result that Fable gave me. I think this one is also really
00:04:20nice with the track layout, but the cars are a little more basic than what we got out of Opus 5.
00:04:25I do actually really like the camera wobble that it does when we turn, and again, all of the
00:04:28other features work, like track positioning, track timing, and all of that. So Fable 5 has
00:04:33also done a really good job, but honestly, I do think I prefer the Opus one. This is the result
00:04:37that I got from GPT 5.6 Soul, also on the max effort. I'd say this one has done a nice job with
00:04:43the modeling and sort of the camera look, but you can see there there is no track, and there's some
00:04:47assets that are backwards, and also about halfway around this track, the track does break. So it's
00:04:51done a worse job, in my opinion, than Fable and Opus. I definitely think that Opus is still
00:04:56winning here. And the last model I tested is Kimi K3, and this is the result they gave me,
00:05:01and honestly, super impressive for an open weight model. It's done a pretty good job,
00:05:05pretty similar to GPT 5.6 Soul. The barriers work, the track works, everything like that,
00:05:09positioning, everything is working. It's built me a pretty nice game in a single shot. If we take a
00:05:14look at the cost and timings of those runs, Opus 5 actually spent 46 minutes and 51 seconds
00:05:19to complete the F1 game, and would have cost around $12.99 if we used the API. So for me,
00:05:25Opus 5 was actually more expensive than Fable here, and looking into it, it was because it used
00:05:30five times more tokens. I'm really hoping this is an outlier, as the other benchmarks don't seem to
00:05:35report the same thing. I will add that I'm calculating the cost of those clawed models by using the clawed
00:05:39code usage tool, so there might be some accuracy issues there, as you don't get the exact price when
00:05:44you use the subscription. Also, as I said earlier, these models are still a premium, so if you come down
00:05:49to something like GPT 5.6, you are going to get a quicker and cheaper model, but the results were,
00:05:54in my opinion, worse. Let's try another test though. This time, I'm asking them to create a personal
00:05:58finance management dashboard, so it's going to be a full stack application with a front end and a back
00:06:02end, and I also list a few features that they can add here. Now, this is the result that I got back
00:06:06with Opus 5, and honestly, I've got to say I am pretty impressed. I really like this UI, it's just a
00:06:11style that I personally like. Let me know in the comments if you do as well, and everything is
00:06:15fully functional. I have tested all of these work, and we do have separate pages for each of the
00:06:20features that I asked for here, and again, all of the features are working between the front end and
00:06:24the back end, and I just really like the UI style that it's gone with. I think it's the best one out
00:06:28of the lot. As for the code itself, it used React and React Router on the front end, which I'm a big
00:06:33fan of, and in the back end, it also gets plus points because it did use a real database. It used
00:06:37Node SQLite here for the database, and for the actual server it used Express. This is the result
00:06:42that Fable 5 gave me, and I'm a bigger fan of the Opus 5 UI, but this one also isn't too bad,
00:06:48but these charts are a bit useless, to be honest. This one down here is fairly nice, and again,
00:06:53all of these features do work. I just think Opus 5 put a bit more energy into the UI, and it's
00:06:57definitely better at that side of things. In the code, it did something similar to Opus 5, where it used
00:07:01React, but I will say it didn't pick a routing library. It actually just built a small routing
00:07:05library itself. I probably would prefer that it sticks to something like React Router,
00:07:09and then for the database, it did the same thing where it's using Node SQLite here, and this one
00:07:13is also built on Express, so a pretty similar back end there, but I'd say Opus actually chose the
00:07:18better stack for the front end. This is what GPT 5.6 Sol gave me, and I've got to say this one is up
00:07:22there for UI design. It's definitely matching Opus. I think it's just a different style, so it's
00:07:26probably going to come down to personal preference, but I do really like this one. It's done an
00:07:31incredible job of mapping out all of the UI, and again, all of the features do work, but this one
00:07:35didn't actually use different pages for these individual features. It just goes to a different
00:07:38part of the page. It also chose the most unique stack out of all of them. It went with NextJS with
00:07:43a combination of Drizzle, which I do really like for the database, and is a pretty valid approach for
00:07:47something like this. But then it also used Cloudflare and Vinext to actually host it, which I do think is a
00:07:52bit of an odd decision, because as you can see, Vinext is very early, and I don't think it's been
00:07:56tested yet, and they're definitely putting something in the prompt to actually force it into using Cloudflare
00:08:00and Vinext. And I said this in my Kimi K3 video, it's probably because that is how their own hosted
00:08:04ChatGPT sites work. Finally, the last model that I tested is Kimi K3, and this is the result that we
00:08:09got back, and I'd say it's done a very good job. This UI is sort of something I would have gotten out
00:08:14of GPT 5.4 or 5.5, so I think it's a little bit behind in UI design, but again, all of the features
00:08:19are working, and I can't really complain too much about this UI. It does exactly what I asked it to do.
00:08:24The code does get some minus points from me though. It did choose React for its frontend,
00:08:28and again, it built its own routing library like Fable did, but then if we go to the server,
00:08:32it actually only has an express server, and it didn't build out any database using Node SQLite
00:08:36or anything like that. It actually just built out its own database using JavaScript, so it saves
00:08:40everything to a JSON file. That is definitely not the approach that I would want it to take.
00:08:44As for the timing and pricing of the models on that finance test, Fable was the most expensive,
00:08:48costing $30.79, and also taking 56 minutes and 55 seconds, but again, with Opus, it seems to be
00:08:55pretty close in price to Fable, as it cost me $29.82, and it actually took longer at 59 minutes and 15
00:09:02seconds. The GPT models are just really well priced in comparison, it's about three and a half times
00:09:07cheaper, and I'd really hope this competition will start to bring those Anthropic prices down.
00:09:11So it's a bit of a mixed bag for me. Obviously, I only ran each test once, and I only did two tests,
00:09:16so I'm hoping that the artificial analysis benchmarks are more accurate, where Opus is
00:09:20actually a third of the price of Fable, but I'll definitely have to use it a bit more to work that
00:09:24out. Maybe it's just the tool that I'm using to actually measure the costs. Another point for Opus
00:09:28though is that it's slightly less strict on cybersecurity work than Fable 5. The safeguards are pretty
00:09:33similar to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber
00:09:37tasks. Apparently, the safeguards will allow Opus 5 to find vulnerabilities in source code,
00:09:42but it will block binary-based vulnerability scanning, as well as penetration testing and
00:09:46exploit generation. Their testing found that their classifiers intervene around 85% less often
00:09:51than they do for Fable 5. So there we go, the model competition is genuinely quite impressive
00:09:55right now. We have GPT 5.6 showing that prices can be bought down, we have open models like Kimmy
00:10:00matching the big guys in intelligence, and we have Anthropic who are constantly pushing that
00:10:04intelligence further and further. What is your favourite model, and do you like the sound of Opus?
00:10:09Let me know in the comments down below, while you're there subscribe, and as always, see you in the next one.