Transcript
00:00:00So Fable 5.1 is here, and it is better, cheaper,
00:00:02and today I'm gonna tell you
00:00:04everything you need to know about it.
00:00:05So first things first, price.
00:00:07Fable 5.1 will cost an estimated 25% less than Fable 5,
00:00:11and this is because they're reducing the price
00:00:13on cash reads.
00:00:14This is huge if you do a lot of long-running agentic work,
00:00:19and they specifically call this out,
00:00:21and your cost reduction might be up to 45%
00:00:23in those cases where you're doing work
00:00:24that has thousands and thousands of thousands of tokens,
00:00:27and you're working in context window lengths
00:00:28of 50, 60, 70%.
00:00:30They're updating their data retention policy,
00:00:32they're phasing that in for their enterprise customers,
00:00:34and in regards to safeguards,
00:00:35they're reducing the false positives,
00:00:37which was kind of a pain in the butt
00:00:38with the original Fable 5,
00:00:40because if you asked it questions
00:00:41related to things like cybersecurity,
00:00:43you would constantly get blocked,
00:00:44and in this case, it's now a 60% fewer chance
00:00:47of a false positive.
00:00:49Now let's take a look at the benchmarks.
00:00:50Across the board, it beats Fable 5, Opus 5,
00:00:53and GPT 5.6 at literally everything.
00:00:56Now, some of the jumps are pretty significant.
00:00:58When we look at agentic scientific research,
00:01:00it's double what Fable 5 was.
00:01:03Agentic coding also takes a huge leap.
00:01:05We're going from 42% with Fable 5 to 55.8 with Fable 5.1.
00:01:09I remember when we looked at the benchmarks for Opus 5,
00:01:11we were kind of blown away by what Opus 5 was able to give us
00:01:14in terms of raw numbers.
00:01:16In reality, I think most people still prefer Fable 5 to Opus 5,
00:01:19but I feel like the comparison on Fable 5 versus 5.1
00:01:23is probably a little bit more accurate
00:01:24in terms of what it's going to feel like
00:01:26when you're actually using it.
00:01:28Beyond that, most of the benchmark increases
00:01:29are relatively incremental.
00:01:31We're talking a few percentage points here and there,
00:01:32the exception being business workflows
00:01:34using the automation benchmark.
00:01:37Now, what's important isn't just the benchmark itself,
00:01:39it's how much are we paying for these sort of scores,
00:01:41and what sort of changes do we see
00:01:43when we mess with the effort level?
00:01:44Because that's always really interesting to look at,
00:01:46because oftentimes when we look at extra high
00:01:48or max effort levels,
00:01:49we really don't get that much bang for our buck
00:01:51compared to high or even medium.
00:01:53So what we have here is the Terminal Bench 4.0.
00:01:56We have Mythos 5.1 up here up top,
00:01:59Fable 5.1 in the middle,
00:02:00and then Mythos 5 down here,
00:02:02and you can imagine Fable 5 is slightly below Mythos.
00:02:05So Fable 5.1 and Mythos 5.1,
00:02:07obviously a big jump compared to Mythos 5,
00:02:10and it is cheaper.
00:02:12And what's really cool is when we look at something
00:02:13like high, for example,
00:02:15which is where I think most people sit at,
00:02:16oftentimes I'm using medium
00:02:17when it comes to my Fable use cases,
00:02:20you're essentially,
00:02:20if we look at medium,
00:02:21we are matching the max output on Mythos 5,
00:02:25but instead of it being $26,
00:02:27it would be $7.80.
00:02:30And when we go to high,
00:02:32and we compare high to say max,
00:02:34you're only looking at a,
00:02:35what, 6% change in terms of the output,
00:02:38yet $19.50 versus $10.50.
00:02:42So what's really cool about these new models
00:02:43is you can get extremely high performance
00:02:45at sort of these middling effort levels
00:02:47and really save on cost.
00:02:50And we see the same pattern play out
00:02:52in these other benchmarks, right?
00:02:53This is Humanity's last exam.
00:02:55Again, we don't see a huge increase
00:02:57when we jump from high all the way to max,
00:02:59yet we have really good performance
00:03:00at a pretty reasonable cost
00:03:02on these middle tier settings.
00:03:04The one exception is when we look at
00:03:05agentic coding on Cursor Bench,
00:03:07Fable 5.1,
00:03:08we do see a pretty high jump from high to extra high,
00:03:11but again, between extra high and max,
00:03:13not a ton.
00:03:14But, you know, if I'm on high,
00:03:15on Fable 5.1,
00:03:17it's the same as being on extra high
00:03:18with Fable 5,
00:03:20yet less than half the cost.
00:03:22So huge price savings here,
00:03:24even if you aren't someone
00:03:25who's necessarily doing like wildly complex work.
00:03:27If you're just quote unquote,
00:03:29your average solo dev,
00:03:30like your problems aren't that complicated,
00:03:32but you like using these frontier models,
00:03:33you now have the ability
00:03:34to just drop down the effort level
00:03:36and you're pretty much doing
00:03:37what you were before
00:03:38at again, half the cost,
00:03:40which can't be overstated.
00:03:41Now, scientific research
00:03:42has always been a pretty big selling point
00:03:44when it comes to the Fable and Mythos models.
00:03:46And so with the 5.1 release,
00:03:48no changes there.
00:03:49They talk a lot about molecular design,
00:03:51computational analysis and modeling,
00:03:53as well as computational biology.
00:03:55Again, these are more sort of niche problems
00:03:57that some people are solving,
00:03:58but it is always kind of interesting
00:04:00to see what's going on in these domains
00:04:01versus your standard,
00:04:03what is this doing
00:04:03on this agentic coding benchmark.
00:04:05And as I alluded to in the intro,
00:04:07there has been some changes
00:04:08when it comes to safety,
00:04:09security,
00:04:09and alignment.
00:04:11Big picture,
00:04:12less false positives.
00:04:13You can ask it more questions
00:04:14related to things like cybersecurity
00:04:15and it's smart enough
00:04:16not to just freak out
00:04:18and send you to Opus
00:04:19no matter what.
00:04:20If you want to go really depth,
00:04:21in depth on this,
00:04:21they have their system card
00:04:23that you always can take a look at.
00:04:24This is, you know,
00:04:25several hundred pages of stuff,
00:04:27212 to be exact,
00:04:28but they give us a good summary here.
00:04:30But all you need to care about
00:04:31is that these guardrails
00:04:32are more precise,
00:04:33so they're less likely
00:04:34to flag benign content
00:04:35and just send you
00:04:36to a lower model.
00:04:37Specifically for cybersecurity questions,
00:04:40Cloud Code users
00:04:40can expect an average
00:04:41of 60% fewer interventions
00:04:43per session
00:04:44from these safeguards,
00:04:46which is great
00:04:46if you're working
00:04:47on any sort of application
00:04:48where you need to ask questions
00:04:49about best cybersecurity practices
00:04:51for your particular use case.
00:04:53They do have
00:04:53an interesting paragraph here
00:04:54talking about
00:04:55anti-distillation mechanisms,
00:04:57which is a method
00:04:58to essentially extract
00:04:59the best out of Fable 5
00:05:01and Fable 5.1
00:05:02to use in other models.
00:05:03This is a really big
00:05:04sort of discussion
00:05:05in relation to
00:05:06open source models.
00:05:07You probably heard a ton
00:05:08about all these
00:05:08really good open source models
00:05:09coming out,
00:05:10and they are no doubt
00:05:11going to continue coming out
00:05:13as they distill
00:05:14from these frontier models.
00:05:15So this is the type
00:05:17of discussion
00:05:17you kind of want
00:05:18to keep an eye on.
00:05:18It doesn't really affect you,
00:05:20the end user,
00:05:21but it is, you know,
00:05:22sort of like this meta thing
00:05:23where we have all these
00:05:24open source models
00:05:25coming from China
00:05:26and a lot of the frontier
00:05:26sort of companies
00:05:27are in the U.S.,
00:05:28and there's always
00:05:29this sort of, you know,
00:05:31talk that all these
00:05:32open source models
00:05:32really are just pulling
00:05:33from things like
00:05:35Opus and Fable.
00:05:36They give an example
00:05:37of how this works
00:05:38where new API accounts
00:05:39can't like edit
00:05:40Claude's prior context
00:05:42to essentially extract
00:05:43the way Claude
00:05:44is thinking
00:05:44and coming up
00:05:45with its answers.
00:05:46And again,
00:05:46this is specifically
00:05:47for new accounts,
00:05:48so if you already have one,
00:05:48like again,
00:05:49this won't really affect you.
00:05:50Now, in terms of cost
00:05:51and availability,
00:05:52availability,
00:05:52it's available
00:05:53everywhere right now.
00:05:54Cost.
00:05:55Per token,
00:05:56if you look at it,
00:05:56it is the exact same
00:05:58as Fable 5.
00:05:58We're looking at $10
00:05:59per million input tokens
00:06:00and $50 per million
00:06:02output tokens,
00:06:03but in reality,
00:06:04when you actually use this
00:06:05in terms of,
00:06:06hey, I'm having it
00:06:07do this task,
00:06:08it is going to cost less
00:06:09and it's going to cost less
00:06:10because of cash reads.
00:06:12Now, if you have no idea
00:06:12what prompt caching is,
00:06:14you can check a video
00:06:14I did last week
00:06:16where I talk about this.
00:06:18Essentially,
00:06:18this is huge
00:06:19if you're someone
00:06:20who is doing
00:06:20long-running agentic tasks.
00:06:23So you have like
00:06:23very large sessions,
00:06:25tons of context,
00:06:26and you're using it
00:06:27like continuously.
00:06:28You're not like having
00:06:29it do something
00:06:29and coming back the next day.
00:06:30You're just like talking
00:06:31and talking and talking
00:06:32with it.
00:06:32So these cash reads
00:06:33cost 75% less,
00:06:35which is a massive change,
00:06:37which is what is resulting
00:06:39in these lower costs.
00:06:40And this graph spells
00:06:41that out here.
00:06:42This can be roughly
00:06:4225% less in terms
00:06:44of your cost on normal tasks,
00:06:45but again,
00:06:46highly agentic workloads
00:06:47up to 45% less.
00:06:49So really exciting
00:06:49to get an upgrade
00:06:50to the Fable model.
00:06:51I have loved Fable 5
00:06:53was kind of eh
00:06:53when it came to Opus 5,
00:06:54but for those of you
00:06:55who have been hands-on
00:06:56with Fable,
00:06:57I think you can definitely agree
00:06:58that there is
00:06:59a significant change
00:07:01in what this model brings us
00:07:02versus everything else
00:07:03we've gotten into
00:07:03from Anthropic.
00:07:04So to have something
00:07:05that's better
00:07:06and cheaper
00:07:07and we have to worry less
00:07:09in terms of the guardrails
00:07:10I think is a great thing.
00:07:11So definitely check it out.
00:07:12Let me know
00:07:13what you think.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video