Opus 5 Is Here And Its BETTER THAN FABLE?!

CChase AI
컴퓨터/소프트웨어경제 뉴스AI/미래기술

Transcript

00:00:00So Claude Opus 5 just dropped in, and Anthropic is claiming we now have a model that is getting
00:00:04close to Fable 5 outputs at half the cost. So let's take a look at what they're giving us.
00:00:09So let's take a look at some of the benchmarks, and right away, kind of wild numbers. We see that
00:00:14it beats Fable 5 and blows Opus 4.8 out of the water on a number of these tests, specifically
00:00:20agentic terminal coding, knowledge work, agentic search. Also, it's beating things like GPT's 5.6
00:00:27in these categories as well. The only areas that we see Opus 5 doing worse than Fable 5 is
00:00:33multidisciplinary reasoning, which isn't by much. In fact, it beats it when it uses tools, as well as
00:00:39legal and health. But every other spot, it's actually beating their top model. And if those numbers are to
00:00:46be believed, that's crazy. Right here, it's saying on cursor bench 3.2 at max effort, the model performed
00:00:51within 0.5% of Fable 5's peak score, but at half the cost per task. Nuts. And you can see that right
00:00:57here. We have Opus 5 in the red, and then Fable 5 in the orange, and down here in the blue is Opus 4.8.
00:01:03So yes, at max, max effort, Fable 5 does pull ahead, but most people aren't actually operating
00:01:08at max. They're either operating at high or extra high. And when we look at extra high and high
00:01:13compared to Fable, we're pretty much getting the same outputs with Opus 5, but again, at like
00:01:17half the cost. And in some of these other benchmarks, like artificial analysis coding agent index,
00:01:23Opus 5 is just straight up beating Fable 5. And same thing with the frontier bench. Like
00:01:27it's, it's actually whooping Fable and it's cheaper. Now let's take a look at some knowledge
00:01:32work and problem solving tasks. And again, I think what's really crazy isn't just this whole
00:01:37like Fable 5, Opus 5 dynamic. It's the jump from Opus 4.8 to Opus 5. Like all of these
00:01:43across the board, just a massive, massive leap forward. And again, if this was like a Sonnet
00:01:495 thing where we saw it doing better than the previous model, yet it was just like prohibitively
00:01:53expensive because Sonnet 5 is somehow the most expensive model out there when we break it down
00:01:57by task. That's not the case here. Opus looks to be extremely token efficient compared to these
00:02:03two other models. And again, like the fact it's actually beating Fable 5, I really can't get
00:02:07over that. Another interesting improvement is visual output. So here's a look at Opus 5 doing
00:02:13some wind tunnel work. And then over here, we have it building a 3D interactive animal cell
00:02:19artifact. Now, one of the downsides with Cloud is it like can't create its own images and stuff. So
00:02:24it being able to do better with these sort of like 3D created SVG, almost like HTML graphics is huge.
00:02:31Another big improvement is that Anthropic is saying it's much stronger in verifying its work
00:02:35and iterating carefully until it succeeds. What does this mean? This means when we use Opus 5
00:02:40on long horizon agentic tasks, things that just aren't one shot. These are like loops, right? This
00:02:45is like, you know, the whole graph engineering thing people are talking about. Anything that requires
00:02:49Claude to essentially check its work against some sort of goal we give it and keep doing that over and
00:02:53over and over again. Well, Opus 5 is much better than Opus 4.8 in that regard. And one interesting
00:02:58example here is Opus 5 was given a drawing of a machine part and told to write code to rebuild it as a
00:03:033D free CAD model. But the model wasn't given any way to actually view the drawing. So, you know,
00:03:09it was told to reach this goal. It wasn't told specifically how to do it. So it went ahead and
00:03:14wrote its own computer vision pipeline to pull the geometry from raw pixels and then reconstructed the
00:03:19whole machine part. Again, anything that is long horizon, given some sort of North Star to reach for,
00:03:25Opus 5 is much better than 4.8 in actually completing that. In terms of misaligned behavior,
00:03:29it's also a step above all these previous models. Lower is better here. And we see a huge increase in its capability
00:03:35from Opus 4.8 to find vulnerabilities and exploits in open source code. Now the other cool thing is the
00:03:40safeguards. This isn't a Fable scenario where if you mention anything about cybersecurity biology,
00:03:45it's going to get shut down. Those guardrails are much less strict with Opus 5 versus Fable 5.
00:03:50Now, while there are still some classifiers in there, they say it's 85% less likely that they actually
00:03:56trigger versus Fable. Now, lastly, the big thing is cost. Like I said, this is half the price of Fable.
00:04:01We're looking at $5 per million input and 25 per million output. And like we saw from the benchmarks
00:04:07in the charts, this is a relatively token efficient model. So I'm really excited to try this one out
00:04:12because if the numbers are true, we kind of almost have a better Fable for half the cost.

Key Takeaway

Claude Opus 5 matches or outperforms Fable 5 across most agentic coding and knowledge benchmarks at half the cost, priced at $5 per million input and $25 per million output tokens.

Highlights

  • Claude Opus 5 delivers output performance within 0.5% of Fable 5 on CursorBench 3.2 at maximum effort while cutting task costs in half.

  • Anthropic sets the price of Opus 5 at $5.00 per million input tokens and $25.00 per million output tokens.

  • On high and extra-high effort levels, Opus 5 matches or exceeds Fable 5 performance across coding, agentic search, and terminal agent benchmarks.

  • The model includes guardrail classifiers that trigger 85% less frequently than Fable 5 during cybersecurity and biology tasks.

  • When tasked with reconstructing a 3D CAD model from an unseen drawing, Opus 5 autonomously created its own computer vision pipeline to parse pixel geometry.

Timeline

Benchmark Performance and Pricing Structure

  • Opus 5 outperforms Fable 5 and Opus 4.8 in agentic terminal coding, knowledge work, and agentic search.
  • On CursorBench 3.2 at maximum effort, Opus 5 performs within 0.5% of Fable 5 while cutting task costs in half.
  • Fable 5 leads slightly in multidisciplinary reasoning, legal, and health benchmarks when tools are not used.

Comparative benchmark testing shows Opus 5 surpassing previous frontier models like GPT-5.6 and Opus 4.8 across multiple core operational metrics. While Fable 5 retains a 0.5% edge at absolute peak computation on CursorBench 3.2, Opus 5 produces equivalent results at standard high and extra-high effort settings. On the Artificial Analysis Coding Agent Index and FrontierBench, Opus 5 outperforms Fable 5 without requiring premium token pricing.

Token Efficiency and Visual Output Capabilities

  • Opus 5 achieves high token efficiency relative to Sonnet 5 and Opus 4.8.
  • The model generates interactive 3D SVG and HTML graphics directly to compensate for native image generation limits.
  • Visual outputs include rendered wind tunnel simulations and interactive 3D animal cell models.

Unlike Sonnet 5, which incurs high per-task costs, Opus 5 uses fewer tokens to complete complex queries. Because the Claude suite lacks native raster image generation, Opus 5 relies on programmatic SVG and HTML code rendering to generate precise visual artifacts. Demonstrated outputs include aerodynamic flow visualizers and manipulable 3D biology diagrams.

Long-Horizon Agentic Execution and Safeguard Adjustments

  • Opus 5 autonomously validates and iterates on long-horizon loop tasks until target criteria are met.
  • Without explicit instructions, the model wrote a custom vision pipeline to extract pixel geometry from an image and build a 3D FreeCAD model.
  • Safety classifiers in Opus 5 are 85% less likely to false-trigger compared to Fable 5 on technical domain prompts.

Self-verification and error correction allow Opus 5 to handle multi-step, goal-oriented engineering workflows without human intervention. During testing, the model successfully reconstructed a mechanical component in 3D CAD by writing code to analyze image pixels directly. Additionally, Anthropic reduced guardrail friction for technical topics such as cybersecurity and biology, lowering false-positive interventions by 85% relative to Fable 5 while improving open-source vulnerability detection.

Community Posts

View all posts