스크립트
00:00:00Over the last 36 hours, a lot of new AI models dropped, and you'll see plenty of videos about them, I'm sure.
00:00:06In this video, I'll simply tell you which one was the best one, what's the disappointing one,
00:00:11the meh one, the overlooked one, which we all missed, or most of us missed,
00:00:15and the not-so-hidden stars, so that you know which model to use for what.
00:00:19I'll not even make it too interesting, I'll tell the best one right away.
00:00:25Probably unsurprising, that would be Opus 5.5.
00:00:29Claude released or announced Opus 5.5 a couple of hours ago, you can already use it in Claude code in your subscription.
00:00:37And of course, like always, we have amazing benchmark numbers, it's essentially better than the other top models in most of these areas.
00:00:45And of course, we have to take this with a grain of salt, you all know how those benchmarks work.
00:00:51It's easy and it's done, that you optimize models to do well in benchmarks, we had models that had great benchmark
00:00:59numbers but did not perform that well, but still of course, we look at those numbers and yeah, Opus 5.5 is quite impressive there.
00:01:06Now, they themselves then shared more details, for example, and that is important, that it is cheaper.
00:01:13And it's not just cheaper compared to Opus 5, regarding the input and output token cost, it uses less compute to serve and it seems to be more efficient.
00:01:24So they say that it should cost around 40% less than Opus 5, which I also understand is, it'll go farther in your subscription usage, if you have the Claude subscription, you should be able to get more Opus usage out of it.
00:01:40It's also interesting that if we look at comparisons like this, where we basically see the cost per attempt, so lower is better, and the score higher is better, we see that Opus, of course, vastly outperforms, according to this benchmark, not just GPT-6 Astra, but also Fable, which is the green one.
00:02:02Which, of course, kind of means there shouldn't really be any cases where you need to reach for Fable anymore, if you trust those benchmarks, because Opus is just better.
00:02:14And especially, like, the high usage here seems to be a real sweet spot, according to this one benchmark, of course, there are more, because you get a very high score and a cost here that is much lower than what Fable would be giving you.
00:02:29Keep in mind that this is log scale.
00:02:31We can also see this in this chart, which includes more models, where I just highlighted the latest Opus model, Fable and Astra, that the latest Opus 5.5 graph basically gives us higher intelligence across all reasoning efforts with lower cost, typically, especially also compared to Fable, where the max Fable reasoning effort gives us a lower intelligence and costs more than the max Opus 5.5 effort level.
00:03:00Of course, benchmarks are just that, benchmarks as mentioned, but from the receptions I saw on X, people also seem to be very happy about Opus 5.5.
00:03:10Again, we have to be careful here, we all know that as a new model is released, we got the typical influencers coming out that finally they can talk about the model and that it's really amazing and best and blah, blah, blah, you know it all.
00:03:23But I've seen more and more posts by people who don't seem to be those influencers, who seem genuinely impressed by the model and in my first steps with it, it also seemed to be intelligent, stick to instructions, work its way through complex problems and also not to be underestimated.
00:03:43It's something which Anthropic also announced in their post here on X, that it is a better communicator.
00:03:49And they showed this comparison here, that it no longer spits out those long jargon-heavy explanations and essays, but that it instead is more concise, gives you the answer you're looking for in normal English.
00:04:05And that is something I also can confirm, it's much better there.
00:04:10So that seems to be a really amazing model.
00:04:13And if you got an Anthropic subscription, a Claude subscription, I think it's a no-brainer to use that.
00:04:19I currently don't really see why you would still use Fable 5.1 in certain situations, but I have to do more testing on that.
00:04:28It's definitely the default over Opus 5, of course.
00:04:31And if you don't have a Claude subscription yet, this may be a good reason to get one.
00:04:35Now, before I dive into all these other models, this is probably a good place for some advertisement for myself.
00:04:42I'm launching a program, it's not live yet, but you can sign up to be notified when enrollments do open up, where I share how I build applications, web applications with AI.
00:04:53This is a two-week program, I'll walk you through the entire process of researching, planning, setting up the project, setting up guardrails, discussing constraints, doing the systems design parts, using my favorite tech stack, in fact, Alchemy Cloudflare.
00:05:10To show you and take you with me on how I build with AI, without vibe coding, without getting lost, with great results.
00:05:19By the end of the two-week program, we'll have a finished application.
00:05:22And therefore, yeah, you'll find a link below, of course.
00:05:25And if that sounds interesting, register to know when signups open up.
00:05:30The disappointing one is not GPT-6 Sol, we'll talk about this soon.
00:05:35It's not the winner, but it's also not the disappointing one.
00:05:38The disappointing one instead is Grok 4.7.
00:05:42I think launched around 30 hours ago or so.
00:05:48Also, of course, impressive benchmark numbers, though, if we're honest, it's not the clear winner in all these categories.
00:05:56It's indeed only the winner in two of the categories here, and none of them is about coding in this benchmark published here.
00:06:03But it's a step up in software engineering compared to 4.6, for example, and it has the same pricing.
00:06:10So all good.
00:06:11Nope.
00:06:12Grok 4.7 earns that place here because in my testing and also what you can read from other people,
00:06:23it's just not very efficient.
00:06:26It's slow and it takes very long to complete tasks.
00:06:31I found multiple posts like this where people essentially mentioned that it just burns through more usage
00:06:38and doesn't seem smarter than 4.6.
00:06:42My personal experience has been basically the same.
00:06:46I used it for various tasks and it wasn't bad.
00:06:50It's not a bad model, obviously.
00:06:52It's just not a huge step up from 4.6, as it seems, and it definitely was much slower and burned more tokens.
00:07:01So for Grok, you should be using Grok 4.6 instead of 4.7.
00:07:09I don't think it's worth the upgrade, at least right now.
00:07:12If they make it better still, if they tweak the pricing or anything like that, things may change.
00:07:17But right now, it's the disappointing one.
00:07:20Now, what's the meh one?
00:07:22Well, maybe you already guessed it.
00:07:24That would be GPT-6 Sol.
00:07:27We did not just get Opus 5.5.
00:07:30We also got GPT-6 Sol by OpenAI.
00:07:33And that is the meh one.
00:07:37Astra GPT-6 is a pretty good model.
00:07:40It has its quirks.
00:07:41It can be a bit hard to wield and to get the best results out of it.
00:07:47But it's a good model in my experience.
00:07:50GPT-6 Sol, and that's important, is also a good model.
00:07:54Well, I guess if we wouldn't have gotten Opus 5.5 yesterday, receptions would have been better.
00:08:02But it just feels like a model that doesn't fit in that well.
00:08:09If we take a look at the official benchmark here, the automation bench numbers published by OpenAI,
00:08:15we can see that GPT-6 Sol is smarter and more cost-efficient than 5.6 Sol.
00:08:23And indeed, that's the way to think about it and the way they want you to think about it.
00:08:28It's a better version of 5.6 Sol.
00:08:31It is worse, intelligence-wise, than Astra, though.
00:08:36It is more cost-efficient, though.
00:08:38And therefore, that's why it's meh.
00:08:40It's not disappointing.
00:08:42It's kind of what it is intended to be, a better 5.6 Sol.
00:08:48It's just not on the same level as Astra, intelligence-wise.
00:08:54There is a big gap here in this benchmark alone.
00:08:58It is more cost-efficient, though.
00:09:00So, the idea essentially is that in your ChatGPT, in your OpenAI subscription,
00:09:06you get more usage out of that subscription,
00:09:09and you use GPT-6 Sol for all the tasks
00:09:12where you don't need the highest level of intelligence.
00:09:16The problem, of course, just is that with Opus 5.5,
00:09:21that's not a trade-off you have to make.
00:09:24Of course, we'll soon get probably Fable 5.5 or whatever,
00:09:29and then things will shift again.
00:09:30But with the launch yesterday,
00:09:32we essentially have an Opus model,
00:09:35which is cheaper than Fable,
00:09:37which is better than Fable.
00:09:40Or at least on the same level, depending on how you look at it.
00:09:43For GPT-6 Sol, that's not the case.
00:09:46It is definitely worse than Astra.
00:09:50Now, I still have to do more testing myself,
00:09:53how I feel about it in my projects.
00:09:55And I definitely, at least right now,
00:09:58do see myself using it for a lot of day-to-day work
00:10:02because of the increased cost efficiency.
00:10:06But when working on complex projects,
00:10:09you really want to go for the highest possible intelligence.
00:10:13That is the reason why before,
00:10:15you probably should have used Fable
00:10:17and not Settle for Opus 5 for everything
00:10:21or for GPT-5.6 Sol for everything.
00:10:25That's the reason why you want to use Astra
00:10:27and keep it under control
00:10:29and maybe not GPT-6 Sol.
00:10:33So it's meh.
00:10:35It has its place.
00:10:36It has its use.
00:10:39It's just a bit more limited
00:10:41and you need to think about it more
00:10:43than you do have for Opus 5.5.
00:10:46You don't have to think a lot about that.
00:10:48You can see it in this intelligence index
00:10:50versus cost per intelligence graph here
00:10:54on artificial analysis.
00:10:57There we have Opus 5.5,
00:10:59which is that top line here.
00:11:02I showed you that before.
00:11:03And GPT-6 Sol,
00:11:05which is this black line here.
00:11:07This other black line here is Astra.
00:11:09That, as you can see,
00:11:10is a bit worse than Opus 5.5,
00:11:12but still good,
00:11:13but costly.
00:11:15GPT-6 Sol is dumber,
00:11:18if you want to put it like this,
00:11:19but cheaper.
00:11:21So, yeah, it's kind of meh.
00:11:24If you have a ChatGPT subscription
00:11:26and that's your only subscription,
00:11:28you're using it for coding,
00:11:30my recommendation would be
00:11:31to stick to Astra
00:11:33and only switch to GPT-6 Sol
00:11:36if you're really struggling
00:11:37with your usage
00:11:38and or if you're tackling
00:11:41smaller, less complex tasks,
00:11:45I guess.
00:11:46But then again,
00:11:47that may be the wrong way
00:11:48of doing agentic engineering,
00:11:50of programming with AI.
00:11:51But still,
00:11:52that's how I would use it
00:11:53if I had to use it.
00:11:54If you have the choice,
00:11:56right now,
00:11:56Opus 5.5 is probably the best bet.
00:11:59Now, before I go to the overlooked one,
00:12:01let's talk about
00:12:01the not-so-hidden star
00:12:03because that is a GPT-6 model,
00:12:06but it's Luna.
00:12:08I will say right away,
00:12:09this is not a model
00:12:11that qualifies for all the use cases.
00:12:14Definitely not the model
00:12:15you want to use for complex
00:12:18or even regular coding work
00:12:22or tasks with AI.
00:12:25But GPT-5.6 Luna
00:12:26was already pretty interesting
00:12:28for certain cases.
00:12:29And for GPT-6 Luna,
00:12:31that's true too.
00:12:32And again,
00:12:33I'll go back
00:12:33to this official benchmark here.
00:12:36What's important here
00:12:38is that the x-axis
00:12:40is log scale.
00:12:41It's not linear.
00:12:43And that actually does
00:12:45GPT-6 Luna,
00:12:47which is this graph here,
00:12:48a disservice.
00:12:50It makes it look worse
00:12:51than it is
00:12:51because this model
00:12:53is dirt cheap.
00:12:55The max effort version
00:12:59of that model
00:13:00comes out at below 5 cents
00:13:04per task
00:13:05in this benchmark.
00:13:07The intelligence
00:13:08is on the same level
00:13:11as with GPT-5.6 Sol
00:13:15in the medium effort setting
00:13:18or the GPT-6 Sol
00:13:22model
00:13:23at the low effort setting.
00:13:26So again,
00:13:27that is of course
00:13:28not necessarily
00:13:28what you want to use
00:13:29for frontier coding work.
00:13:32But for trivial tasks
00:13:34and for non-coding tasks,
00:13:37that is definitely
00:13:38a model to look at
00:13:39because not everybody
00:13:39is doing coding
00:13:40all day long.
00:13:41And this model
00:13:42is essentially for free
00:13:44in your subscription
00:13:45and also if you use it
00:13:47through the API.
00:13:49I mean,
00:13:49if you want to get
00:13:50the best version of Astra,
00:13:52the most powerful version
00:13:53of Astra,
00:13:54we're looking at
00:13:56just below
00:13:57two dollars
00:13:58per task here.
00:14:00The best version of Luna,
00:14:01which of course
00:14:02is far dumber,
00:14:03but again,
00:14:04for certain use cases,
00:14:05great,
00:14:06is a fraction of that.
00:14:09And that's why
00:14:10I wanted to give it
00:14:10a shout out,
00:14:11so to say.
00:14:12Why I want you
00:14:13to be aware of it.
00:14:14GPT-6 Luna
00:14:15is actually
00:14:17a great model
00:14:19whenever you don't
00:14:21need the top-notch
00:14:23intelligence.
00:14:24So,
00:14:25if you have
00:14:26some more trivial tasks,
00:14:29some tasks
00:14:30where maybe
00:14:31a bunch of data
00:14:33needs to be
00:14:34transformed
00:14:35from format A
00:14:36to format B
00:14:37or anything like that,
00:14:38anything that does
00:14:39not require
00:14:40a lot of intelligence,
00:14:41but where you need
00:14:42a workhorse,
00:14:45look at Luna.
00:14:46It may be
00:14:46a good idea.
00:14:48And just to be
00:14:48very clear,
00:14:49this is far
00:14:51less intelligence
00:14:52than Astra.
00:14:53It will easily
00:14:54beat all the models
00:14:55we had a year ago
00:14:57or so
00:14:57on the top
00:14:58reasoning effort.
00:14:59It's kind of
00:15:00expectation management
00:15:01you have to do here,
00:15:02but this is really
00:15:04something you should
00:15:05not overlook.
00:15:06What is the
00:15:07overlooked one,
00:15:08though for me,
00:15:09and I'm only
00:15:10talking about
00:15:11recent model launches,
00:15:12that would be
00:15:12MIMO 2.6.
00:15:14And maybe you
00:15:15never heard about
00:15:15that before.
00:15:16I'll admit,
00:15:17I never heard
00:15:17about the MIMO
00:15:18models before.
00:15:20It's a model
00:15:21by Xiaomi.
00:15:22And they launched
00:15:23two variants,
00:15:24Pro and Flash,
00:15:25and I only used
00:15:26them through
00:15:27Open Router,
00:15:28so I paid
00:15:29API prices
00:15:30in my testing.
00:15:31I don't have
00:15:32any subscription
00:15:33there,
00:15:33and I think
00:15:34you can only
00:15:34get a subscription
00:15:35with a Chinese
00:15:37VPN or something
00:15:38like this,
00:15:39but this is a
00:15:40good model.
00:15:40They published
00:15:41some benchmark
00:15:42numbers.
00:15:44It's worth noting
00:15:45that they did not
00:15:46include Astra here,
00:15:47not Fable 5.1,
00:15:48obviously not
00:15:49Opus 5.5,
00:15:51but their model,
00:15:532.6 Pro,
00:15:55beats Fable 5,
00:15:56for example,
00:15:57in the deep
00:15:58software engineering
00:15:59benchmark here.
00:16:01And it's really good
00:16:02in those benchmarks,
00:16:03and I can confirm,
00:16:04I used it a lot
00:16:05over the last day.
00:16:07It's a good model.
00:16:08Now, I will say,
00:16:09it's thinking
00:16:10is from a
00:16:11different world.
00:16:13I made a post
00:16:14about this yesterday.
00:16:15It will think
00:16:16for 10 minutes,
00:16:1815 minutes.
00:16:18I had it think
00:16:19forever before
00:16:20it actually starts
00:16:21doing something.
00:16:22Now, obviously,
00:16:23you can control
00:16:24effort level,
00:16:25and I'm talking
00:16:26about the Pro
00:16:27model here,
00:16:29the Flash one,
00:16:30obviously, as the name
00:16:31suggests, is quicker,
00:16:32but also dumber.
00:16:33But yeah,
00:16:34it takes long,
00:16:35but it's so, so cheap.
00:16:37I spent a few bucks,
00:16:39really just a few bucks,
00:16:40on having it work
00:16:41for hours
00:16:43through not super
00:16:45complex problems,
00:16:46but through
00:16:47complex problems,
00:16:49I would say.
00:16:50And I was really
00:16:51impressed by the results.
00:16:52Now,
00:16:53I guess it's not a model
00:16:55that will find
00:16:56widespread use
00:16:57because, again,
00:16:58it's not in an
00:16:59easily accessible
00:17:00subscription or
00:17:01anything like that.
00:17:02But if you're
00:17:03building applications
00:17:05that need AI
00:17:06inside of them,
00:17:07that need a brain,
00:17:08if you're building
00:17:09your own coding agent
00:17:10where you can't
00:17:11or don't want
00:17:12to use
00:17:12a subscription,
00:17:15anything like that,
00:17:16that's also a model
00:17:17to look at.
00:17:18I mean, in general,
00:17:19and that's no news.
00:17:20All these Chinese models
00:17:21are pretty amazing,
00:17:23and I'm super excited
00:17:24to see what
00:17:25Quen 4 will give us
00:17:27and so on.
00:17:28But yeah,
00:17:28this one launched
00:17:30over the last
00:17:31couple of hours
00:17:32and I guess
00:17:33most people missed it
00:17:34and I think
00:17:34you shouldn't miss it.
00:17:35It has its place
00:17:36and it's a good model,
00:17:39especially if we
00:17:40look at something
00:17:41like Grok 4.7,
00:17:42which was rather
00:17:43disappointing
00:17:43and is slow
00:17:45and expensive.
00:17:46This one is slow
00:17:47in the Pro version,
00:17:48but it's not expensive
00:17:49and it's pretty good.
00:17:50So that's my overview
00:17:52and my opinions
00:17:54on these models
00:17:55and what I would
00:17:56recommend for using
00:17:57where.
커뮤니티 글
아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!
이 영상에 대해 글쓰기