This Open Source Model Just Beat Claude Fable 5

CChase AI
Computing/SoftwareBusiness NewsInternet Technology

Transcript

00:00:00Did a Chinese openweight model just overthrow GPT 5.6 and Fable 5?
00:00:05Well, if you've been anywhere near YouTube or Twitter in the last 24 hours,
00:00:08and you have been bombarded with these charts showing KimiK3 performing just as well,
00:00:13if not better, than the best Anthropic and OpenAI has to offer,
00:00:17all while doing it at a fraction of the price.
00:00:20But should you believe the hype?
00:00:22Well, that's exactly what we're going to answer in today's video
00:00:24as we take a deeper look at what these benchmarks actually tell us,
00:00:28and we do a head-to-head test between GPT 5.6, Fable 5, and K3
00:00:34in the one area that KimiK3 is posting better numbers than anybody, front-end design.
00:00:40So by the end of this video, we'll have a much better idea if KimiK3 is the new king.
00:00:46So let's begin with the numbers and benchmarks when it comes to KimiK3
00:00:49because there's a lot we can actually pull out of these
00:00:51and there's some nuance involved that goes beyond these charts
00:00:53that you've seen floating around all over the place, these two being the most popular.
00:00:57Now, when we talk about KimiK3, it is a 2.8 trillion open-weight model.
00:01:02When we talk about open-weight, we talk about open-source,
00:01:05this does not mean you can run this on your computer.
00:01:07We're talking millions of dollars of hardware required to run one of these things.
00:01:11But it's still great that it's open-weight,
00:01:12that we can actually see how all the parameters are tuned
00:01:15versus something like Fable 5 or GPT 5.6, where that is invisible to us.
00:01:20Now, how does it hold up on the numbers?
00:01:23This comes to us courtesy of Kimi and their moonshot lab,
00:01:27and it shows it competing, again, with Fable 5 and GPT 5.6
00:01:31on all these major programming benchmarks.
00:01:34The other chart you've seen floating around is this one.
00:01:36This is coming from arena.ai, and it's measuring their front-end code.
00:01:40Now, how this benchmark works is if you go to arena.ai and you ask it,
00:01:44hey, create me a landing page,
00:01:45it will give you two responses from two different models,
00:01:49and you don't know which one is which.
00:01:50It's like the Pepsi challenge, and you choose whichever one you want.
00:01:53Over time, we see which ones have been voted against its competitor.
00:01:57And right now, Kimi K3 has done better than every other model in this space.
00:02:01But understand this is very, very subjective.
00:02:03And so when we look at these two charts,
00:02:05and we compare that to the pricing between these models,
00:02:08we see that Kimi K3 is also extremely cheap.
00:02:13Its input is $3 per million tokens compared to Fable, which is 10,
00:02:17compared to 5.6, which is 5.
00:02:20So it's 60% of the cost of 5.6,
00:02:23and it's 30% of the cost of Fable 5 on the input.
00:02:26And output-wise, it's half the cost of 5.6,
00:02:29and it is 30% of the cost of Fable 5.
00:02:32So way cheaper, and it's basically just as good if not better.
00:02:35So what's not to love?
00:02:36But this is where we need to start diving a little bit deeper
00:02:38in terms of cost and speed.
00:02:40Because it's one thing to see, oh, per million tokens,
00:02:43it's $3 and that one's $10.
00:02:45Well, how many tokens does it actually use?
00:02:46Not all models are treated equal in terms of token efficiency.
00:02:49In fact, these open-source Chinese models tend to be token hogs,
00:02:53which means while on paper, it's very cheap.
00:02:56What does that mean in reality?
00:02:57And how does it compare to these guys
00:02:58who are actually very, very token-efficient, like GPT 5.6?
00:03:01Well, if we look at this from artificial analysis,
00:03:03which is an independent benchmarker,
00:03:05they have a cost per intelligence index task.
00:03:09So how much does it cost to complete these tasks?
00:03:11So Kimi K3, right here, looking at $0.95.
00:03:15Versus GPT 5.6, sole, is $1.04.
00:03:20So slightly more expensive.
00:03:23If we look at 5.5 extra high, you know, it's over here.
00:03:26If we look at Terra, which is 5.6 on max,
00:03:29it's half the price of Kimi K3.
00:03:30Now, versus Claude, right, pretty expensive.
00:03:35Fable 5, $2.75.
00:03:38So this cost savings definitely applies
00:03:40when we compare Fable to Kimi K3.
00:03:42But when we compare it to 5.6, not a huge difference.
00:03:46We're talking like 10% or less.
00:03:49And in fact, there's certain instances where Terra,
00:03:51some of these smaller models in the 5.6 realm,
00:03:54are actually half the cost of Kimi K3.
00:03:56Now, the other thing we want to take note of is time,
00:03:58because time is certainly a currency
00:03:59we want to pay attention to.
00:04:01Historically, these open-source models from China are slow.
00:04:04And you can see that here.
00:04:05The slowest model of the bunch is the old Kimi K2.6.
00:04:09Now, Kimi K3 has gotten significantly better.
00:04:12In this chart, it's tracking at six minutes.
00:04:14If we compare that to Fable 5,
00:04:17Fable 5 is at five minutes and 5.6 is at 4.7.
00:04:21So, you know, it's less token efficient and it's slower.
00:04:27But overall, much cheaper than Fable 5
00:04:32and about on par with 5.6.
00:04:34Now, the last benchmark I want to talk about
00:04:35is the omniscience index.
00:04:37This is measuring hallucinations.
00:04:39They give the model a series of very difficult questions
00:04:42that it may or may not know the answer to.
00:04:44And it's graded in terms of, did it get it right?
00:04:46Did it get it wrong?
00:04:47Or does it say, I don't know?
00:04:49You get a point if you get it right.
00:04:50You lose a point if you get it wrong.
00:04:52And you get a zero if you just say, I don't know.
00:04:54And we can see here, and this is out of 100,
00:04:57Fable 5 scores the best at 40 points.
00:05:005.6 is way behind at 22.
00:05:03And then Kimi K3 is at 18.
00:05:06And I think this benchmark is really important
00:05:08if you're using these agents in a context
00:05:10that just requires a lot of nuance
00:05:11and there's a lot of gray area.
00:05:13So when we take all these benchmarks together
00:05:14in the aggregate, I think what we should walk away from
00:05:16is that Kimi K3 on paper is extremely powerful.
00:05:19However, when we talk about how cheap it is,
00:05:21that's rather overblown,
00:05:23especially when we compare it to the GPT models
00:05:25that are hyper token efficient.
00:05:27On top of that, Kimi K3 is just a tad bit slower as well.
00:05:30So then it just becomes a question of
00:05:32what are those three things is most important to you
00:05:34in regards to power, cost, and speed.
00:05:38So now let's go to the front
00:05:39and test in between these three models.
00:05:40So right here on the left,
00:05:42I have Fable 5 running inside of Claude Code.
00:05:45And on the right, I have Kimi K3
00:05:46also running inside of Claude Code.
00:05:49And back here, I have Codex.
00:05:51So this is the prompt I'm going to give it.
00:05:54And I'm telling it I wanted to build
00:05:55a 3D globe travel dashboard
00:05:57that feels like a scene from my sci-fi film,
00:05:59not just a website.
00:06:01I'm telling it I wanted to essentially
00:06:02make it a visual spectacle.
00:06:04And I want it to get rather creative.
00:06:06In fact, I say surprise me
00:06:08with at least one idea I haven't seen before.
00:06:10I try not to get too specific with this prompt
00:06:13because I kind of want to see
00:06:14how all these models diverge
00:06:15and what they actually build me.
00:06:16So let's put them to the test.
00:06:20All right, so all three are done
00:06:22building their website.
00:06:23So we're going to go through all of them.
00:06:24And then at the end,
00:06:25we'll compare the actual cost,
00:06:28the amount of tokens used, and the time.
00:06:30So let's start with Kimi 3.
00:06:34So here we go.
00:06:35We have the 3D Earth.
00:06:37I can zoom in and out.
00:06:38I can move it around.
00:06:41Not super high fidelity,
00:06:43but that's all right.
00:06:43If I click on something like Mexico City,
00:06:46nothing really happens.
00:06:51What else?
00:06:51It has some other points on here,
00:06:53like Cape Town.
00:06:55Over here on the left,
00:06:56it shows some active routes.
00:06:58So if I click on an active route like Tokyo,
00:07:01it actually brings me to it.
00:07:02And then over here on the right-hand side,
00:07:04I'll move over for a little bit.
00:07:05We can see, you know,
00:07:07it's a long time,
00:07:08weather,
00:07:09the best window to travel there,
00:07:10entry,
00:07:11and then sort of like
00:07:12how much it costs right now.
00:07:16And as I click through stuff over here on the left,
00:07:19you know,
00:07:20it brings me to these different towns
00:07:21as well as all their stats.
00:07:23So overall,
00:07:24pretty cool.
00:07:24And you can see how it also has the Earth
00:07:26kind of like divided into,
00:07:29you know,
00:07:29daytime versus nighttime and all that.
00:07:31So overall,
00:07:31it looks pretty sweet.
00:07:32Now let's take a look at Fable 5.
00:07:34So right off the bat,
00:07:36I would say these graphics
00:07:38look a little bit cleaner,
00:07:39right?
00:07:39If we compare these two,
00:07:40this just looks a little more crisp.
00:07:43We also have some lighting,
00:07:44which is cool.
00:07:45We have that same like
00:07:46daytime,
00:07:46nighttime breakdown,
00:07:47but I can actually,
00:07:49I have a scroll wheel,
00:07:51a little thing right here
00:07:52where I can switch between day and night,
00:07:55which is pretty awesome.
00:07:57I can't zoom in as far,
00:08:00but I can see more of these sort of cities
00:08:02and it looks a little bit more high fidelity.
00:08:04So if I click on one of these,
00:08:06right,
00:08:06same thing as before,
00:08:08brings up some stats.
00:08:09Lat long shows the conditions
00:08:11in terms of weather and time
00:08:12and then sort of what the fair is.
00:08:15Shows return fair,
00:08:16although I'm not totally sure
00:08:17like where this is actually coming from.
00:08:19If I click here over on the left,
00:08:21same sort of thing,
00:08:22brings me to that city.
00:08:24I can see some of the stats
00:08:25as well as what it would cost.
00:08:27I think overall,
00:08:29in terms of sort of the UI polish
00:08:33between these two,
00:08:34I think I prefer
00:08:36what Fable 5 put out,
00:08:38especially this sort of like,
00:08:40you know,
00:08:41day versus night thing
00:08:42it has happening.
00:08:44And interesting enough,
00:08:45it looks like
00:08:46when I change like where it is
00:08:48during the day,
00:08:48you can see it updates
00:08:50the cost
00:08:51on the right-hand side
00:08:52depending on when you would
00:08:53actually be landing there.
00:08:55So pretty good.
00:08:56And now let's look at GPT 5.6.
00:08:59So initial thoughts,
00:09:01feels like a step down
00:09:02from what we saw
00:09:03with Kimmy K3
00:09:04and Fable
00:09:04in terms of the globe itself.
00:09:06Like these globes
00:09:07definitely have a lot more going on
00:09:09and are much more
00:09:09of a visible spectacle.
00:09:11This,
00:09:11I think they almost went more
00:09:12of like a sci-fi look.
00:09:15That kind of has like random,
00:09:17it's like,
00:09:17are these supposed to be?
00:09:19Okay.
00:09:20So same sort of thing
00:09:21if I click on certain cities here
00:09:23like Singapore,
00:09:24this pops up,
00:09:24but it kind of overlaps
00:09:26the globe itself
00:09:27and everything is so small text-wise,
00:09:29it's kind of hard to even read.
00:09:31I do have here at the bottom
00:09:33sort of like a day versus night
00:09:35kind of sweep
00:09:36like we had over there
00:09:37with Fable 5,
00:09:39but it's not as great.
00:09:42And when it's nighttime,
00:09:43you like can't even see
00:09:44the continent itself.
00:09:45I guess those lights
00:09:46are supposed to be like cities.
00:09:47If I click over here on the left,
00:09:49yeah,
00:09:50same as before
00:09:51brings us to those actual cities.
00:09:53But again,
00:09:53definitely feels like a step down
00:09:54for the last two we saw.
00:09:56So overall,
00:09:56when we compare these three,
00:09:57I would put Fable 5 in the lead.
00:10:00Not far behind,
00:10:02I would have Kimi K3
00:10:04and then in third place,
00:10:05I would put GPT 5.6.
00:10:07So now let's compare tokens,
00:10:09time and cost
00:10:10because we saw the end result,
00:10:11but what did we actually
00:10:13have to pay for it?
00:10:15So we had Kimi,
00:10:16we had Fable
00:10:17and then we had Codex,
00:10:19which I'm just going to put as X.
00:10:21So in terms of tokens,
00:10:24Kimi was by far the heaviest.
00:10:26It took it 21.5 million tokens
00:10:30to come up with that.
00:10:31And in terms of time,
00:10:33it took one hour
00:10:35and 33 minutes.
00:10:37So over 90 minutes
00:10:39to create that
00:10:39in 21.5 million tokens.
00:10:42Cost,
00:10:42and this is all coming
00:10:43from Open Router,
00:10:44was $8.66.
00:10:47Now let's talk about Fable.
00:10:49In terms of tokens,
00:10:50it used 3.5 million.
00:10:53In terms of time,
00:10:55it took just 17 minutes
00:10:56and total cost
00:10:58was 11.64.
00:11:00So again,
00:11:02what did I talk about
00:11:02at the beginning?
00:11:03We talked about speed
00:11:04and token efficiency.
00:11:06Kimi K3 is clearly
00:11:07way behind Fable
00:11:09in that regard,
00:11:09especially when it comes
00:11:11to time.
00:11:11And so when we looked
00:11:12at that token-to-token cost
00:11:14per million,
00:11:15it was like,
00:11:16oh,
00:11:16it's a third of the cost
00:11:17of Fable.
00:11:18Well,
00:11:19in reality,
00:11:20not so much.
00:11:20Still cheaper than Fable.
00:11:22Don't get me wrong,
00:11:23but not nearly
00:11:25as efficient.
00:11:26Then we had Codex,
00:11:27which I was kind of
00:11:28disappointed with its output
00:11:28to be totally honest.
00:11:29It took 5.6 million tokens.
00:11:32For time,
00:11:33it took 25 minutes
00:11:35and its total cost
00:11:36was 5.66.
00:11:38If you're wondering
00:11:39like tokens
00:11:40and cost
00:11:42and why doesn't
00:11:42that exactly line up
00:11:43for like inputs
00:11:44versus outputs
00:11:44versus what's advertised,
00:11:46understand there's a lot
00:11:47of caching going on
00:11:48for all three of these.
00:11:49So what does this
00:11:51kind of mean?
00:11:52Well,
00:11:52big thing is Kimi
00:11:53right?
00:11:55Very heavy
00:11:56on the tokens
00:11:57and it took forever.
00:11:59It took an hour
00:12:00and a half.
00:12:00To be honest,
00:12:01when I was making this video,
00:12:01I was like,
00:12:02oh,
00:12:02maybe I'll do like,
00:12:03you know,
00:12:03three,
00:12:04maybe even four examples.
00:12:05No,
00:12:06I'm not sitting here
00:12:07for 20 hours to do this.
00:12:09And,
00:12:09you know,
00:12:10so way behind
00:12:11in those two places
00:12:12against the Frontier models,
00:12:13but overall,
00:12:14its output was solid.
00:12:15You know,
00:12:15I thought it did
00:12:16a really good job,
00:12:17all things considered.
00:12:18And the most expensive
00:12:19of the bunch,
00:12:20obviously,
00:12:20was Fable
00:12:21with Codex being
00:12:22the cheapest.
00:12:24So what can we really pull
00:12:26from all this?
00:12:27Well,
00:12:27I think what we can pull
00:12:29at least here
00:12:29on the front end side,
00:12:31and it probably applies
00:12:32the same to the coding
00:12:32if we want to take
00:12:33the benchmarks
00:12:33at face value,
00:12:34and that's Kimi
00:12:36can compete.
00:12:37Kimi K3
00:12:38is an open source model
00:12:39that can compete
00:12:40with Fable in 5.6.
00:12:42However,
00:12:43it's not as cheap
00:12:45as you would expect.
00:12:46In certain cases,
00:12:47it's more expensive.
00:12:49And it's just slow.
00:12:51It is so slow.
00:12:54Which is kind of
00:12:55a deal breaker
00:12:55for a lot of people
00:12:56depending on what you're doing.
00:12:58But,
00:12:58if that's okay with you,
00:13:01you know,
00:13:02then maybe it isn't
00:13:03so much of a downside.
00:13:03But I think
00:13:04the big thing
00:13:05that kind of gets lost
00:13:06in the benchmarks
00:13:06and the numbers
00:13:07is really the cost.
00:13:08Right?
00:13:09Because it's not
00:13:09as cheap
00:13:10as you would expect
00:13:11because of this
00:13:12inefficiency
00:13:12when it comes to
00:13:14actually using tokens.
00:13:15So that's where
00:13:16I'm going to leave you guys.
00:13:16Hope I was able
00:13:17to shed some more light
00:13:18on Kimi K3.
00:13:19I think it's awesome.
00:13:20We have an open source model
00:13:21that can compete.
00:13:22But let's not get ahead
00:13:23of ourselves
00:13:23in terms of
00:13:24its raw power
00:13:25and specifically
00:13:26its cost.
00:13:27Both in terms of money
00:13:28and time.
00:13:29So,
00:13:30let me know in the comments
00:13:31what you thought.
00:13:31Make sure to check out
00:13:32Chase AIA Plus
00:13:32if you want to get your hands
00:13:33on my Cloud Code Masterclass
00:13:34that also includes
00:13:35a Codex Masterclass
00:13:36these days.
00:13:37Besides that,
00:13:39I'll see you around.

Key Takeaway

While Kimi K3 proves that large open-weight models can match the output quality of proprietary systems like Fable 5 and GPT 5.6, its significantly lower token efficiency and slower processing times prevent it from being the clear cost-saving leader.

Highlights

  • Kimi K3 is a 2.8 trillion parameter open-weight model that competes with Fable 5 and GPT 5.6 on programming benchmarks.

  • Front-end coding tests demonstrate that Kimi K3 is slower and less token-efficient than industry-leading frontier models.

  • Kimi K3 required 21.5 million tokens and 93 minutes to build a 3D globe travel dashboard, while Fable 5 finished the same task in 3.5 million tokens and 17 minutes.

  • The total cost to generate the dashboard using Kimi K3 was $8.66, compared to $11.64 for Fable 5 and $5.66 for Codex.

  • Hallucination benchmarks rate Fable 5 at 40 points, followed by GPT 5.6 at 22 points and Kimi K3 at 18 points.

Timeline

Model Performance and Benchmarking

  • Kimi K3 is a 2.8 trillion parameter model that is open-weight but requires massive infrastructure to run.
  • Benchmarks indicate Kimi K3 is highly competitive with Fable 5 and GPT 5.6 in programming and front-end design.
  • Token inefficiency in Chinese open-source models often offsets their lower per-token pricing compared to models like GPT 5.6.

Public charts suggest Kimi K3 outperforms current industry standards in front-end design, yet these benchmarks are subjective and prone to nuances. Despite a low base input cost of $3 per million tokens, independent analyses show that the actual cost to complete tasks is often comparable to or only slightly lower than GPT 5.6. Furthermore, Kimi K3 lags behind in hallucination tests, scoring 18 out of 100 on the omniscience index compared to 40 for Fable 5.

Head-to-Head Front-End Development Test

  • Three models were tasked with building a 3D globe travel dashboard that emphasizes visual spectacle.
  • Fable 5 produced the most polished UI, including an interactive day-to-night light transition.
  • Kimi K3 delivered a solid result but felt slightly less refined than Fable 5.
  • GPT 5.6 performed the worst, with poor text readability and problematic UI overlaps.

Each model attempted to create a sci-fi inspired 3D globe. Kimi K3 and Fable 5 both successfully generated functional 3D environments with interactive data, but Fable 5 demonstrated superior UI design and high-fidelity lighting features. GPT 5.6 struggled to render the globe effectively, making it the least successful implementation of the group.

Efficiency, Cost, and Time Analysis

  • Kimi K3 used 21.5 million tokens to complete the task, making it the least token-efficient model tested.
  • The generation time for Kimi K3 was 93 minutes, while Fable 5 required only 17 minutes.
  • Actual project costs show that token inefficiency negates the benefit of low base pricing for Kimi K3.
  • High latency makes Kimi K3 a difficult choice for applications requiring rapid iteration.

Despite the theory that Kimi K3 is cheaper due to its $3 per million token rate, the high volume of tokens consumed during actual development tasks resulted in an $8.66 total cost. While this is less than the $11.64 spent on Fable 5, the time disparity is significant, with Kimi K3 taking over an hour longer to deliver the final output. The results confirm that raw model power does not always translate to operational efficiency in real-world coding scenarios.

Community Posts

View all posts