Fable 5.1 Is The Greatest Model Ever (And Cheaper Than Fable 5)

CChase AI
Computing/Software

Transcript

00:00:00So Fable 5.1 is here, and it is better, cheaper,
00:00:02and today I'm gonna tell you
00:00:04everything you need to know about it.
00:00:05So first things first, price.
00:00:07Fable 5.1 will cost an estimated 25% less than Fable 5,
00:00:11and this is because they're reducing the price
00:00:13on cash reads.
00:00:14This is huge if you do a lot of long-running agentic work,
00:00:19and they specifically call this out,
00:00:21and your cost reduction might be up to 45%
00:00:23in those cases where you're doing work
00:00:24that has thousands and thousands of thousands of tokens,
00:00:27and you're working in context window lengths
00:00:28of 50, 60, 70%.
00:00:30They're updating their data retention policy,
00:00:32they're phasing that in for their enterprise customers,
00:00:34and in regards to safeguards,
00:00:35they're reducing the false positives,
00:00:37which was kind of a pain in the butt
00:00:38with the original Fable 5,
00:00:40because if you asked it questions
00:00:41related to things like cybersecurity,
00:00:43you would constantly get blocked,
00:00:44and in this case, it's now a 60% fewer chance
00:00:47of a false positive.
00:00:49Now let's take a look at the benchmarks.
00:00:50Across the board, it beats Fable 5, Opus 5,
00:00:53and GPT 5.6 at literally everything.
00:00:56Now, some of the jumps are pretty significant.
00:00:58When we look at agentic scientific research,
00:01:00it's double what Fable 5 was.
00:01:03Agentic coding also takes a huge leap.
00:01:05We're going from 42% with Fable 5 to 55.8 with Fable 5.1.
00:01:09I remember when we looked at the benchmarks for Opus 5,
00:01:11we were kind of blown away by what Opus 5 was able to give us
00:01:14in terms of raw numbers.
00:01:16In reality, I think most people still prefer Fable 5 to Opus 5,
00:01:19but I feel like the comparison on Fable 5 versus 5.1
00:01:23is probably a little bit more accurate
00:01:24in terms of what it's going to feel like
00:01:26when you're actually using it.
00:01:28Beyond that, most of the benchmark increases
00:01:29are relatively incremental.
00:01:31We're talking a few percentage points here and there,
00:01:32the exception being business workflows
00:01:34using the automation benchmark.
00:01:37Now, what's important isn't just the benchmark itself,
00:01:39it's how much are we paying for these sort of scores,
00:01:41and what sort of changes do we see
00:01:43when we mess with the effort level?
00:01:44Because that's always really interesting to look at,
00:01:46because oftentimes when we look at extra high
00:01:48or max effort levels,
00:01:49we really don't get that much bang for our buck
00:01:51compared to high or even medium.
00:01:53So what we have here is the Terminal Bench 4.0.
00:01:56We have Mythos 5.1 up here up top,
00:01:59Fable 5.1 in the middle,
00:02:00and then Mythos 5 down here,
00:02:02and you can imagine Fable 5 is slightly below Mythos.
00:02:05So Fable 5.1 and Mythos 5.1,
00:02:07obviously a big jump compared to Mythos 5,
00:02:10and it is cheaper.
00:02:12And what's really cool is when we look at something
00:02:13like high, for example,
00:02:15which is where I think most people sit at,
00:02:16oftentimes I'm using medium
00:02:17when it comes to my Fable use cases,
00:02:20you're essentially,
00:02:20if we look at medium,
00:02:21we are matching the max output on Mythos 5,
00:02:25but instead of it being $26,
00:02:27it would be $7.80.
00:02:30And when we go to high,
00:02:32and we compare high to say max,
00:02:34you're only looking at a,
00:02:35what, 6% change in terms of the output,
00:02:38yet $19.50 versus $10.50.
00:02:42So what's really cool about these new models
00:02:43is you can get extremely high performance
00:02:45at sort of these middling effort levels
00:02:47and really save on cost.
00:02:50And we see the same pattern play out
00:02:52in these other benchmarks, right?
00:02:53This is Humanity's last exam.
00:02:55Again, we don't see a huge increase
00:02:57when we jump from high all the way to max,
00:02:59yet we have really good performance
00:03:00at a pretty reasonable cost
00:03:02on these middle tier settings.
00:03:04The one exception is when we look at
00:03:05agentic coding on Cursor Bench,
00:03:07Fable 5.1,
00:03:08we do see a pretty high jump from high to extra high,
00:03:11but again, between extra high and max,
00:03:13not a ton.
00:03:14But, you know, if I'm on high,
00:03:15on Fable 5.1,
00:03:17it's the same as being on extra high
00:03:18with Fable 5,
00:03:20yet less than half the cost.
00:03:22So huge price savings here,
00:03:24even if you aren't someone
00:03:25who's necessarily doing like wildly complex work.
00:03:27If you're just quote unquote,
00:03:29your average solo dev,
00:03:30like your problems aren't that complicated,
00:03:32but you like using these frontier models,
00:03:33you now have the ability
00:03:34to just drop down the effort level
00:03:36and you're pretty much doing
00:03:37what you were before
00:03:38at again, half the cost,
00:03:40which can't be overstated.
00:03:41Now, scientific research
00:03:42has always been a pretty big selling point
00:03:44when it comes to the Fable and Mythos models.
00:03:46And so with the 5.1 release,
00:03:48no changes there.
00:03:49They talk a lot about molecular design,
00:03:51computational analysis and modeling,
00:03:53as well as computational biology.
00:03:55Again, these are more sort of niche problems
00:03:57that some people are solving,
00:03:58but it is always kind of interesting
00:04:00to see what's going on in these domains
00:04:01versus your standard,
00:04:03what is this doing
00:04:03on this agentic coding benchmark.
00:04:05And as I alluded to in the intro,
00:04:07there has been some changes
00:04:08when it comes to safety,
00:04:09security,
00:04:09and alignment.
00:04:11Big picture,
00:04:12less false positives.
00:04:13You can ask it more questions
00:04:14related to things like cybersecurity
00:04:15and it's smart enough
00:04:16not to just freak out
00:04:18and send you to Opus
00:04:19no matter what.
00:04:20If you want to go really depth,
00:04:21in depth on this,
00:04:21they have their system card
00:04:23that you always can take a look at.
00:04:24This is, you know,
00:04:25several hundred pages of stuff,
00:04:27212 to be exact,
00:04:28but they give us a good summary here.
00:04:30But all you need to care about
00:04:31is that these guardrails
00:04:32are more precise,
00:04:33so they're less likely
00:04:34to flag benign content
00:04:35and just send you
00:04:36to a lower model.
00:04:37Specifically for cybersecurity questions,
00:04:40Cloud Code users
00:04:40can expect an average
00:04:41of 60% fewer interventions
00:04:43per session
00:04:44from these safeguards,
00:04:46which is great
00:04:46if you're working
00:04:47on any sort of application
00:04:48where you need to ask questions
00:04:49about best cybersecurity practices
00:04:51for your particular use case.
00:04:53They do have
00:04:53an interesting paragraph here
00:04:54talking about
00:04:55anti-distillation mechanisms,
00:04:57which is a method
00:04:58to essentially extract
00:04:59the best out of Fable 5
00:05:01and Fable 5.1
00:05:02to use in other models.
00:05:03This is a really big
00:05:04sort of discussion
00:05:05in relation to
00:05:06open source models.
00:05:07You probably heard a ton
00:05:08about all these
00:05:08really good open source models
00:05:09coming out,
00:05:10and they are no doubt
00:05:11going to continue coming out
00:05:13as they distill
00:05:14from these frontier models.
00:05:15So this is the type
00:05:17of discussion
00:05:17you kind of want
00:05:18to keep an eye on.
00:05:18It doesn't really affect you,
00:05:20the end user,
00:05:21but it is, you know,
00:05:22sort of like this meta thing
00:05:23where we have all these
00:05:24open source models
00:05:25coming from China
00:05:26and a lot of the frontier
00:05:26sort of companies
00:05:27are in the U.S.,
00:05:28and there's always
00:05:29this sort of, you know,
00:05:31talk that all these
00:05:32open source models
00:05:32really are just pulling
00:05:33from things like
00:05:35Opus and Fable.
00:05:36They give an example
00:05:37of how this works
00:05:38where new API accounts
00:05:39can't like edit
00:05:40Claude's prior context
00:05:42to essentially extract
00:05:43the way Claude
00:05:44is thinking
00:05:44and coming up
00:05:45with its answers.
00:05:46And again,
00:05:46this is specifically
00:05:47for new accounts,
00:05:48so if you already have one,
00:05:48like again,
00:05:49this won't really affect you.
00:05:50Now, in terms of cost
00:05:51and availability,
00:05:52availability,
00:05:52it's available
00:05:53everywhere right now.
00:05:54Cost.
00:05:55Per token,
00:05:56if you look at it,
00:05:56it is the exact same
00:05:58as Fable 5.
00:05:58We're looking at $10
00:05:59per million input tokens
00:06:00and $50 per million
00:06:02output tokens,
00:06:03but in reality,
00:06:04when you actually use this
00:06:05in terms of,
00:06:06hey, I'm having it
00:06:07do this task,
00:06:08it is going to cost less
00:06:09and it's going to cost less
00:06:10because of cash reads.
00:06:12Now, if you have no idea
00:06:12what prompt caching is,
00:06:14you can check a video
00:06:14I did last week
00:06:16where I talk about this.
00:06:18Essentially,
00:06:18this is huge
00:06:19if you're someone
00:06:20who is doing
00:06:20long-running agentic tasks.
00:06:23So you have like
00:06:23very large sessions,
00:06:25tons of context,
00:06:26and you're using it
00:06:27like continuously.
00:06:28You're not like having
00:06:29it do something
00:06:29and coming back the next day.
00:06:30You're just like talking
00:06:31and talking and talking
00:06:32with it.
00:06:32So these cash reads
00:06:33cost 75% less,
00:06:35which is a massive change,
00:06:37which is what is resulting
00:06:39in these lower costs.
00:06:40And this graph spells
00:06:41that out here.
00:06:42This can be roughly
00:06:4225% less in terms
00:06:44of your cost on normal tasks,
00:06:45but again,
00:06:46highly agentic workloads
00:06:47up to 45% less.
00:06:49So really exciting
00:06:49to get an upgrade
00:06:50to the Fable model.
00:06:51I have loved Fable 5
00:06:53was kind of eh
00:06:53when it came to Opus 5,
00:06:54but for those of you
00:06:55who have been hands-on
00:06:56with Fable,
00:06:57I think you can definitely agree
00:06:58that there is
00:06:59a significant change
00:07:01in what this model brings us
00:07:02versus everything else
00:07:03we've gotten into
00:07:03from Anthropic.
00:07:04So to have something
00:07:05that's better
00:07:06and cheaper
00:07:07and we have to worry less
00:07:09in terms of the guardrails
00:07:10I think is a great thing.
00:07:11So definitely check it out.
00:07:12Let me know
00:07:13what you think.

Key Takeaway

Fable 5.1 outperforms previous frontier models across benchmarks while cutting long-running agentic workload costs by up to 45% through reduced cache read pricing.

Highlights

  • Fable 5.1 costs an estimated 25% less overall than Fable 5, with cost reductions reaching up to 45% for long-running agentic work.

  • Cache reads cost 75% less, dropping from standard pricing to drive significant savings during continuous interactions.

  • False positives for cybersecurity queries decrease by 60% in Fable 5.1 compared to the original model.

  • Agentic scientific research scores double the performance of Fable 5.

  • Agentic coding jumps from 42% on Fable 5 to 55.8% on Fable 5.1.

Timeline

Pricing structure and cost reductions

  • Fable 5.1 costs 25% less than Fable 5 overall.
  • Long-running agentic tasks see cost reductions of up to 45% due to cheaper cache reads.
  • Lower effort levels match previous maximum performance at a fraction of the price.

Price reductions target cash reads, which directly benefit long-running agentic workflows utilizing high context window lengths. Users can drop effort levels to medium or high while retaining near-maximum performance, resulting in substantial savings compared to previous generation models like Mythos 5.

Benchmark performance across domains

  • Fable 5.1 beats Fable 5, Opus 5, and GPT 5.6 across all standard benchmarks.
  • Agentic scientific research doubles in capability compared to Fable 5.
  • Agentic coding increases from 42% to 55.8%.

Performance gains span multiple specialized fields including molecular design, computational analysis, and computational biology. While most benchmark increases are incremental, agentic coding and scientific research demonstrate substantial leaps over previous iterations.

Safety guardrails and security updates

  • False positive interventions drop by 60% for cybersecurity queries.
  • Safeguards are more precise, reducing unnecessary blocks on benign content.
  • New API accounts incorporate anti-distillation mechanisms to protect model context.

Safety alignments reduce the frequency of false positives that previously blocked cybersecurity questions. Cloud Code users experience fewer session interventions, and new API accounts restrict context editing to prevent model extraction.

Token pricing and availability

  • Base token pricing remains at $10 per million input tokens and $50 per million output tokens.
  • Cache reads cost 75% less to drive down active session expenses.
  • Fable 5.1 is globally available across all regions immediately.

While nominal per-token pricing matches Fable 5, prompt caching changes alter real-world operational costs. Continuous conversational tasks and large context sessions benefit from the 75% reduction in cache read expenses.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video