This Open Source Model Just Beat Claude Fable 5
CChase AI
Computing/SoftwareBusiness NewsInternet Technology
Transcript
00:00:00Did a Chinese openweight model just overthrow GPT 5.6 and Fable 5?
00:00:05Well, if you've been anywhere near YouTube or Twitter in the last 24 hours,
00:00:08and you have been bombarded with these charts showing KimiK3 performing just as well,
00:00:13if not better, than the best Anthropic and OpenAI has to offer,
00:00:17all while doing it at a fraction of the price.
00:00:20But should you believe the hype?
00:00:22Well, that's exactly what we're going to answer in today's video
00:00:24as we take a deeper look at what these benchmarks actually tell us,
00:00:28and we do a head-to-head test between GPT 5.6, Fable 5, and K3
00:00:34in the one area that KimiK3 is posting better numbers than anybody, front-end design.
00:00:40So by the end of this video, we'll have a much better idea if KimiK3 is the new king.
00:00:46So let's begin with the numbers and benchmarks when it comes to KimiK3
00:00:49because there's a lot we can actually pull out of these
00:00:51and there's some nuance involved that goes beyond these charts
00:00:53that you've seen floating around all over the place, these two being the most popular.
00:00:57Now, when we talk about KimiK3, it is a 2.8 trillion open-weight model.
00:01:02When we talk about open-weight, we talk about open-source,
00:01:05this does not mean you can run this on your computer.
00:01:07We're talking millions of dollars of hardware required to run one of these things.
00:01:11But it's still great that it's open-weight,
00:01:12that we can actually see how all the parameters are tuned
00:01:15versus something like Fable 5 or GPT 5.6, where that is invisible to us.
00:01:20Now, how does it hold up on the numbers?
00:01:23This comes to us courtesy of Kimi and their moonshot lab,
00:01:27and it shows it competing, again, with Fable 5 and GPT 5.6
00:01:31on all these major programming benchmarks.
00:01:34The other chart you've seen floating around is this one.
00:01:36This is coming from arena.ai, and it's measuring their front-end code.
00:01:40Now, how this benchmark works is if you go to arena.ai and you ask it,
00:01:44hey, create me a landing page,
00:01:45it will give you two responses from two different models,
00:01:49and you don't know which one is which.
00:01:50It's like the Pepsi challenge, and you choose whichever one you want.
00:01:53Over time, we see which ones have been voted against its competitor.
00:01:57And right now, Kimi K3 has done better than every other model in this space.
00:02:01But understand this is very, very subjective.
00:02:03And so when we look at these two charts,
00:02:05and we compare that to the pricing between these models,
00:02:08we see that Kimi K3 is also extremely cheap.
00:02:13Its input is $3 per million tokens compared to Fable, which is 10,
00:02:17compared to 5.6, which is 5.
00:02:20So it's 60% of the cost of 5.6,
00:02:23and it's 30% of the cost of Fable 5 on the input.
00:02:26And output-wise, it's half the cost of 5.6,
00:02:29and it is 30% of the cost of Fable 5.
00:02:32So way cheaper, and it's basically just as good if not better.
00:02:35So what's not to love?
00:02:36But this is where we need to start diving a little bit deeper
00:02:38in terms of cost and speed.
00:02:40Because it's one thing to see, oh, per million tokens,
00:02:43it's $3 and that one's $10.
00:02:45Well, how many tokens does it actually use?
00:02:46Not all models are treated equal in terms of token efficiency.
00:02:49In fact, these open-source Chinese models tend to be token hogs,
00:02:53which means while on paper, it's very cheap.
00:02:56What does that mean in reality?
00:02:57And how does it compare to these guys
00:02:58who are actually very, very token-efficient, like GPT 5.6?
00:03:01Well, if we look at this from artificial analysis,
00:03:03which is an independent benchmarker,
00:03:05they have a cost per intelligence index task.
00:03:09So how much does it cost to complete these tasks?
00:03:11So Kimi K3, right here, looking at $0.95.
00:03:15Versus GPT 5.6, sole, is $1.04.
00:03:20So slightly more expensive.
00:03:23If we look at 5.5 extra high, you know, it's over here.
00:03:26If we look at Terra, which is 5.6 on max,
00:03:29it's half the price of Kimi K3.
00:03:30Now, versus Claude, right, pretty expensive.
00:03:35Fable 5, $2.75.
00:03:38So this cost savings definitely applies
00:03:40when we compare Fable to Kimi K3.
00:03:42But when we compare it to 5.6, not a huge difference.
00:03:46We're talking like 10% or less.
00:03:49And in fact, there's certain instances where Terra,
00:03:51some of these smaller models in the 5.6 realm,
00:03:54are actually half the cost of Kimi K3.
00:03:56Now, the other thing we want to take note of is time,
00:03:58because time is certainly a currency
00:03:59we want to pay attention to.
00:04:01Historically, these open-source models from China are slow.
00:04:04And you can see that here.
00:04:05The slowest model of the bunch is the old Kimi K2.6.
00:04:09Now, Kimi K3 has gotten significantly better.
00:04:12In this chart, it's tracking at six minutes.
00:04:14If we compare that to Fable 5,
00:04:17Fable 5 is at five minutes and 5.6 is at 4.7.
00:04:21So, you know, it's less token efficient and it's slower.
00:04:27But overall, much cheaper than Fable 5
00:04:32and about on par with 5.6.
00:04:34Now, the last benchmark I want to talk about
00:04:35is the omniscience index.
00:04:37This is measuring hallucinations.
00:04:39They give the model a series of very difficult questions
00:04:42that it may or may not know the answer to.
00:04:44And it's graded in terms of, did it get it right?
00:04:46Did it get it wrong?
00:04:47Or does it say, I don't know?
00:04:49You get a point if you get it right.
00:04:50You lose a point if you get it wrong.
00:04:52And you get a zero if you just say, I don't know.
00:04:54And we can see here, and this is out of 100,
00:04:57Fable 5 scores the best at 40 points.
00:05:005.6 is way behind at 22.
00:05:03And then Kimi K3 is at 18.
00:05:06And I think this benchmark is really important
00:05:08if you're using these agents in a context
00:05:10that just requires a lot of nuance
00:05:11and there's a lot of gray area.
00:05:13So when we take all these benchmarks together
00:05:14in the aggregate, I think what we should walk away from
00:05:16is that Kimi K3 on paper is extremely powerful.
00:05:19However, when we talk about how cheap it is,
00:05:21that's rather overblown,
00:05:23especially when we compare it to the GPT models
00:05:25that are hyper token efficient.
00:05:27On top of that, Kimi K3 is just a tad bit slower as well.
00:05:30So then it just becomes a question of
00:05:32what are those three things is most important to you
00:05:34in regards to power, cost, and speed.
00:05:38So now let's go to the front
00:05:39and test in between these three models.
00:05:40So right here on the left,
00:05:42I have Fable 5 running inside of Claude Code.
00:05:45And on the right, I have Kimi K3
00:05:46also running inside of Claude Code.
00:05:49And back here, I have Codex.
00:05:51So this is the prompt I'm going to give it.
00:05:54And I'm telling it I wanted to build
00:05:55a 3D globe travel dashboard
00:05:57that feels like a scene from my sci-fi film,
00:05:59not just a website.
00:06:01I'm telling it I wanted to essentially
00:06:02make it a visual spectacle.
00:06:04And I want it to get rather creative.
00:06:06In fact, I say surprise me
00:06:08with at least one idea I haven't seen before.
00:06:10I try not to get too specific with this prompt
00:06:13because I kind of want to see
00:06:14how all these models diverge
00:06:15and what they actually build me.
00:06:16So let's put them to the test.
00:06:20All right, so all three are done
00:06:22building their website.
00:06:23So we're going to go through all of them.
00:06:24And then at the end,
00:06:25we'll compare the actual cost,
00:06:28the amount of tokens used, and the time.
00:06:30So let's start with Kimi 3.
00:06:34So here we go.
00:06:35We have the 3D Earth.
00:06:37I can zoom in and out.
00:06:38I can move it around.
00:06:41Not super high fidelity,
00:06:43but that's all right.
00:06:43If I click on something like Mexico City,
00:06:46nothing really happens.
00:06:51What else?
00:06:51It has some other points on here,
00:06:53like Cape Town.
00:06:55Over here on the left,
00:06:56it shows some active routes.
00:06:58So if I click on an active route like Tokyo,
00:07:01it actually brings me to it.
00:07:02And then over here on the right-hand side,
00:07:04I'll move over for a little bit.
00:07:05We can see, you know,
00:07:07it's a long time,
00:07:08weather,
00:07:09the best window to travel there,
00:07:10entry,
00:07:11and then sort of like
00:07:12how much it costs right now.
00:07:16And as I click through stuff over here on the left,
00:07:19you know,
00:07:20it brings me to these different towns
00:07:21as well as all their stats.
00:07:23So overall,
00:07:24pretty cool.
00:07:24And you can see how it also has the Earth
00:07:26kind of like divided into,
00:07:29you know,
00:07:29daytime versus nighttime and all that.
00:07:31So overall,
00:07:31it looks pretty sweet.
00:07:32Now let's take a look at Fable 5.
00:07:34So right off the bat,
00:07:36I would say these graphics
00:07:38look a little bit cleaner,
00:07:39right?
00:07:39If we compare these two,
00:07:40this just looks a little more crisp.
00:07:43We also have some lighting,
00:07:44which is cool.
00:07:45We have that same like
00:07:46daytime,
00:07:46nighttime breakdown,
00:07:47but I can actually,
00:07:49I have a scroll wheel,
00:07:51a little thing right here
00:07:52where I can switch between day and night,
00:07:55which is pretty awesome.
00:07:57I can't zoom in as far,
00:08:00but I can see more of these sort of cities
00:08:02and it looks a little bit more high fidelity.
00:08:04So if I click on one of these,
00:08:06right,
00:08:06same thing as before,
00:08:08brings up some stats.
00:08:09Lat long shows the conditions
00:08:11in terms of weather and time
00:08:12and then sort of what the fair is.
00:08:15Shows return fair,
00:08:16although I'm not totally sure
00:08:17like where this is actually coming from.
00:08:19If I click here over on the left,
00:08:21same sort of thing,
00:08:22brings me to that city.
00:08:24I can see some of the stats
00:08:25as well as what it would cost.
00:08:27I think overall,
00:08:29in terms of sort of the UI polish
00:08:33between these two,
00:08:34I think I prefer
00:08:36what Fable 5 put out,
00:08:38especially this sort of like,
00:08:40you know,
00:08:41day versus night thing
00:08:42it has happening.
00:08:44And interesting enough,
00:08:45it looks like
00:08:46when I change like where it is
00:08:48during the day,
00:08:48you can see it updates
00:08:50the cost
00:08:51on the right-hand side
00:08:52depending on when you would
00:08:53actually be landing there.
00:08:55So pretty good.
00:08:56And now let's look at GPT 5.6.
00:08:59So initial thoughts,
00:09:01feels like a step down
00:09:02from what we saw
00:09:03with Kimmy K3
00:09:04and Fable
00:09:04in terms of the globe itself.
00:09:06Like these globes
00:09:07definitely have a lot more going on
00:09:09and are much more
00:09:09of a visible spectacle.
00:09:11This,
00:09:11I think they almost went more
00:09:12of like a sci-fi look.
00:09:15That kind of has like random,
00:09:17it's like,
00:09:17are these supposed to be?
00:09:19Okay.
00:09:20So same sort of thing
00:09:21if I click on certain cities here
00:09:23like Singapore,
00:09:24this pops up,
00:09:24but it kind of overlaps
00:09:26the globe itself
00:09:27and everything is so small text-wise,
00:09:29it's kind of hard to even read.
00:09:31I do have here at the bottom
00:09:33sort of like a day versus night
00:09:35kind of sweep
00:09:36like we had over there
00:09:37with Fable 5,
00:09:39but it's not as great.
00:09:42And when it's nighttime,
00:09:43you like can't even see
00:09:44the continent itself.
00:09:45I guess those lights
00:09:46are supposed to be like cities.
00:09:47If I click over here on the left,
00:09:49yeah,
00:09:50same as before
00:09:51brings us to those actual cities.
00:09:53But again,
00:09:53definitely feels like a step down
00:09:54for the last two we saw.
00:09:56So overall,
00:09:56when we compare these three,
00:09:57I would put Fable 5 in the lead.
00:10:00Not far behind,
00:10:02I would have Kimi K3
00:10:04and then in third place,
00:10:05I would put GPT 5.6.
00:10:07So now let's compare tokens,
00:10:09time and cost
00:10:10because we saw the end result,
00:10:11but what did we actually
00:10:13have to pay for it?
00:10:15So we had Kimi,
00:10:16we had Fable
00:10:17and then we had Codex,
00:10:19which I'm just going to put as X.
00:10:21So in terms of tokens,
00:10:24Kimi was by far the heaviest.
00:10:26It took it 21.5 million tokens
00:10:30to come up with that.
00:10:31And in terms of time,
00:10:33it took one hour
00:10:35and 33 minutes.
00:10:37So over 90 minutes
00:10:39to create that
00:10:39in 21.5 million tokens.
00:10:42Cost,
00:10:42and this is all coming
00:10:43from Open Router,
00:10:44was $8.66.
00:10:47Now let's talk about Fable.
00:10:49In terms of tokens,
00:10:50it used 3.5 million.
00:10:53In terms of time,
00:10:55it took just 17 minutes
00:10:56and total cost
00:10:58was 11.64.
00:11:00So again,
00:11:02what did I talk about
00:11:02at the beginning?
00:11:03We talked about speed
00:11:04and token efficiency.
00:11:06Kimi K3 is clearly
00:11:07way behind Fable
00:11:09in that regard,
00:11:09especially when it comes
00:11:11to time.
00:11:11And so when we looked
00:11:12at that token-to-token cost
00:11:14per million,
00:11:15it was like,
00:11:16oh,
00:11:16it's a third of the cost
00:11:17of Fable.
00:11:18Well,
00:11:19in reality,
00:11:20not so much.
00:11:20Still cheaper than Fable.
00:11:22Don't get me wrong,
00:11:23but not nearly
00:11:25as efficient.
00:11:26Then we had Codex,
00:11:27which I was kind of
00:11:28disappointed with its output
00:11:28to be totally honest.
00:11:29It took 5.6 million tokens.
00:11:32For time,
00:11:33it took 25 minutes
00:11:35and its total cost
00:11:36was 5.66.
00:11:38If you're wondering
00:11:39like tokens
00:11:40and cost
00:11:42and why doesn't
00:11:42that exactly line up
00:11:43for like inputs
00:11:44versus outputs
00:11:44versus what's advertised,
00:11:46understand there's a lot
00:11:47of caching going on
00:11:48for all three of these.
00:11:49So what does this
00:11:51kind of mean?
00:11:52Well,
00:11:52big thing is Kimi
00:11:53right?
00:11:55Very heavy
00:11:56on the tokens
00:11:57and it took forever.
00:11:59It took an hour
00:12:00and a half.
00:12:00To be honest,
00:12:01when I was making this video,
00:12:01I was like,
00:12:02oh,
00:12:02maybe I'll do like,
00:12:03you know,
00:12:03three,
00:12:04maybe even four examples.
00:12:05No,
00:12:06I'm not sitting here
00:12:07for 20 hours to do this.
00:12:09And,
00:12:09you know,
00:12:10so way behind
00:12:11in those two places
00:12:12against the Frontier models,
00:12:13but overall,
00:12:14its output was solid.
00:12:15You know,
00:12:15I thought it did
00:12:16a really good job,
00:12:17all things considered.
00:12:18And the most expensive
00:12:19of the bunch,
00:12:20obviously,
00:12:20was Fable
00:12:21with Codex being
00:12:22the cheapest.
00:12:24So what can we really pull
00:12:26from all this?
00:12:27Well,
00:12:27I think what we can pull
00:12:29at least here
00:12:29on the front end side,
00:12:31and it probably applies
00:12:32the same to the coding
00:12:32if we want to take
00:12:33the benchmarks
00:12:33at face value,
00:12:34and that's Kimi
00:12:36can compete.
00:12:37Kimi K3
00:12:38is an open source model
00:12:39that can compete
00:12:40with Fable in 5.6.
00:12:42However,
00:12:43it's not as cheap
00:12:45as you would expect.
00:12:46In certain cases,
00:12:47it's more expensive.
00:12:49And it's just slow.
00:12:51It is so slow.
00:12:54Which is kind of
00:12:55a deal breaker
00:12:55for a lot of people
00:12:56depending on what you're doing.
00:12:58But,
00:12:58if that's okay with you,
00:13:01you know,
00:13:02then maybe it isn't
00:13:03so much of a downside.
00:13:03But I think
00:13:04the big thing
00:13:05that kind of gets lost
00:13:06in the benchmarks
00:13:06and the numbers
00:13:07is really the cost.
00:13:08Right?
00:13:09Because it's not
00:13:09as cheap
00:13:10as you would expect
00:13:11because of this
00:13:12inefficiency
00:13:12when it comes to
00:13:14actually using tokens.
00:13:15So that's where
00:13:16I'm going to leave you guys.
00:13:16Hope I was able
00:13:17to shed some more light
00:13:18on Kimi K3.
00:13:19I think it's awesome.
00:13:20We have an open source model
00:13:21that can compete.
00:13:22But let's not get ahead
00:13:23of ourselves
00:13:23in terms of
00:13:24its raw power
00:13:25and specifically
00:13:26its cost.
00:13:27Both in terms of money
00:13:28and time.
00:13:29So,
00:13:30let me know in the comments
00:13:31what you thought.
00:13:31Make sure to check out
00:13:32Chase AIA Plus
00:13:32if you want to get your hands
00:13:33on my Cloud Code Masterclass
00:13:34that also includes
00:13:35a Codex Masterclass
00:13:36these days.
00:13:37Besides that,
00:13:39I'll see you around.