GPT 6 Sol & Luna Are Here (And 50% CHEAPER!)

CChase AI
Computing/SoftwareBusiness NewsInternet Technology

Transcript

00:00:00So we didn't just get one new model today.
00:00:02We got three.
00:00:03About an hour ago, Anthropic released Claude, Opus 5.5.
00:00:06And now OpenAI is giving us GPT-6 Sol and GPT-6 Luna.
00:00:12Now, Opus 5.5 looked really good
00:00:14when it came to the benchmarks and the cost.
00:00:16So let's see what these two models bring to the table.
00:00:19Now let's start by looking at the benchmarks.
00:00:21And remember, we have this GPT-6 family.
00:00:24At the top is Astra, which has been out for a few weeks now.
00:00:27That is the main event.
00:00:28Below Astra, we have Sol and we have Luna.
00:00:32These are not meant to be better than Astra.
00:00:35They are meant to give us pretty close capabilities
00:00:38and performance, yet at a much cheaper cost.
00:00:40So if we're dealing with problems
00:00:43that aren't nearly as complex
00:00:44as something that we would throw at Astra,
00:00:45that's where we bring in these two models.
00:00:48So you shouldn't see this and be like,
00:00:49oh, it's not as good as Astra, therefore it sucks.
00:00:51That's not the comparison we're making.
00:00:53So in Frontier Code, when we look at GPT-Astra,
00:00:55what are we seeing?
00:00:56We're seeing it max, a 53.3% score.
00:00:59And what I want you to pay attention to is the cost.
00:01:01$4.59.
00:01:03For Sol, which is the next step down,
00:01:06we have a 49.3% score, a little bit lower,
00:01:09but the cost per task is $2.14.
00:01:12So it's less than half the cost,
00:01:13but we only drop from 53 to 49, which is pretty awesome.
00:01:18And then when we look at Luna,
00:01:19which is the cheapest out of all of them,
00:01:21we drop down to 42.4%,
00:01:24but it's costing us 11 cents per task.
00:01:27And this is on the Frontier Code benchmark.
00:01:29Now on that same benchmark,
00:01:31when we look at Opus 5.5,
00:01:32we were doing 54.4 at 619.
00:01:36So 54.4 at 619.
00:01:38Compare that to Sol,
00:01:40it's about a third of the cost
00:01:41at a 5% drop in score.
00:01:44So this is still an awesome model,
00:01:47even if it's not on the Opus 5.5 or Astra 6 levels.
00:01:50Furthermore, when we go down the effort levels,
00:01:53you know, as we at max, we're at 49%.
00:01:55But if I go to say medium, I'm at 46%,
00:01:58but I'm at 80 cents.
00:02:00So these models, Sol and Luna,
00:02:03just give us the sweet spot
00:02:04that we don't always have with the Anthropic models,
00:02:06where you have like, you have Opus and you have Fable,
00:02:09and they're really, really good,
00:02:10but they're rather expensive.
00:02:11Anthropic really struggles at giving you
00:02:13those cheap, highly efficient models
00:02:14for those non-complicated tasks.
00:02:16And now we have Luna and Sol that can fill that gap.
00:02:20Now, in general, when we look at these benchmarks,
00:02:21where we compare the GPT-6 versions
00:02:23to the 5.6 version,
00:02:24so GPT-6 Sol to GPT-5.6 Sol
00:02:28is a small increase in essentially the performance,
00:02:31but a significant increase
00:02:33in terms of the cost efficiency.
00:02:35This is really significant when we look at Luna.
00:02:37So when we look at 5.6 Luna here to GPT-6 Luna,
00:02:40pretty much the same in regards to performance,
00:02:43but the cost difference is relatively dramatic.
00:02:46We're looking at 15 cents per task
00:02:48over here on the 6 version,
00:02:50and $2.57 cost per task on the 5.6.
00:02:55So just hyper-efficient models.
00:02:57In certain cases, we see GPT-6
00:02:59exceeding Claude Fable 5.1
00:03:01and even beating low-effort GPT-6 Astra.
00:03:05And a lot of this has to do with the pricing changes.
00:03:07Across the board, moving from 5.6 to the 6 series,
00:03:10we see a 50% reduction in both the input
00:03:13and the output costs.
00:03:15When it comes to Luna, the input's only 10 cents
00:03:17and the output is 50 cents, which is crazy.
00:03:20Another big upgrade with these new models
00:03:21is in regards to its factuality.
00:03:24How often is it giving us the wrong answer?
00:03:27And by and large, solid increases across the board.
00:03:30With Luna, back in the day on the 5.6,
00:03:32if I was on max effort, it was giving me answers
00:03:35with a factual error 12% of the time.
00:03:39Compare that with GPT-6, only 7.6% of the time.
00:03:42And if you were using Luna on low,
00:03:44like trying to get really, really cheap and efficient,
00:03:46it was giving you answers with the wrong answer
00:03:4836% of the time.
00:03:50Now it's only 27%.
00:03:51And with Sol, we've gone from 8.5%
00:03:54to only getting an answer with a factual error
00:03:574.6% of the time, which is really close
00:04:00to what we see with GPT-6 Astra, which is 3.9%.
00:04:03So when it comes to accuracy, Sol on the GPT-6 level
00:04:07is pretty much matched up with the best of the best,
00:04:10which is Astra.
00:04:11There's also been improvements to the caching system,
00:04:13and this is really important when we're talking about costs.
00:04:15If I have a warm cache, which means I've had
00:04:17a back and forth conversation with my OpenAI agent
00:04:20and I haven't walked away from my computer for an hour,
00:04:22same thing applies on the anthropic side.
00:04:24When I send it messages, that has a 90% discount.
00:04:29You know, if you walk away for a day,
00:04:30you come back for the same conversation
00:04:31and you send, you know, Astra or Sol or Luna a message,
00:04:36it sends back the entire conversation
00:04:38and it's getting charged the entire input cost,
00:04:41which can be pretty expensive.
00:04:43Also, there were other things that would mess with the cache.
00:04:44If you adjusted your reasoning efforts,
00:04:46let's say you were on low reasoning
00:04:48and you wanted to bump it up to high,
00:04:49well, that would reset your cache too
00:04:51and you would incur a pretty large charge.
00:04:53Stuff like that no longer resets the cache.
00:04:55So adjusting your reasoning effort keeps you in a warm cache state,
00:04:59which is going to use less usage
00:05:01and it's going to save you money,
00:05:02especially if you're on the API.
00:05:03They've actually added a prompt caching dashboard
00:05:06and if you're someone who's on the API
00:05:08and you want to get really, really sort of nitty gritty,
00:05:10you can actually optimize sort of break points
00:05:13when it comes to your cache.
00:05:14And lastly, in terms of availability,
00:05:15GPT-6 Sol in Luna is available essentially everywhere.
00:05:19What will be interesting will be
00:05:21if we see a GPT-6 Terra anytime soon.
00:05:24So definitely a lot to love
00:05:26with this GPT-6 Sol in Luna release.
00:05:29It looks like on paper,
00:05:30maybe it is a step below what we saw
00:05:32on the anthropic side with Opus 5.5,
00:05:35but I think they are trying to solve two different problems.
00:05:38Luna isn't meant to solve these Fable or Opus problems
00:05:41and really neither is Sol.
00:05:43It's meant to deal with these problems
00:05:45that are a step below,
00:05:45which frankly is where most people are operating
00:05:47and do it in a much more efficient manner.
00:05:50And we can't ever discount
00:05:52how important cost is to these equations.

Key Takeaway

GPT-6 Sol and Luna cut API processing costs by over 50% while retaining high factuality and performance levels within 4% to 11% of OpenAI's flagship Astra model.

Highlights

  • OpenAI released GPT-6 Sol and GPT-6 Luna as lower-cost alternatives to the flagship GPT-6 Astra model.

  • GPT-6 Sol scores 49.3% on Frontier Code at $2.14 per task, compared to Astra's 53.3% score at $4.59 per task.

  • GPT-6 Luna reduces cost per task to $0.11 on Frontier Code while delivering a 42.4% score.

  • The GPT-6 series cuts input and output API pricing by 50% compared to the GPT-5.6 generation, pricing Luna at $0.10 input and $0.50 output.

  • Factual error rates dropped to 4.6% for GPT-6 Sol and 7.6% for max-effort GPT-6 Luna, bringing Sol close to Astra's 3.9% error rate.

  • Modifying reasoning effort levels no longer invalidates prompt cache, preserving the 90% warm-cache discount during active API sessions.

Timeline

Model Hierarchy and Performance Benchmarks

  • OpenAI launched GPT-6 Sol and GPT-6 Luna to sit beneath the flagship GPT-6 Astra model.
  • GPT-6 Sol hits a 49.3% score on Frontier Code at $2.14 per task versus Astra's 53.3% score at $4.59 per task.
  • GPT-6 Luna lowers task costs to $0.11 while maintaining a 42.4% score on Frontier Code.

Sol and Luna trade minor accuracy drops for substantial cost reductions on routine computing tasks. Claude Opus 5.5 scores 54.4% on Frontier Code at $6.19 per task, making Sol roughly one-third the price of Opus 5.5 for a 5% score reduction. Medium effort settings on Sol drop accuracy slightly to 46% while reducing the task cost to $0.80.

Generational Cost Reductions and Pricing Changes

  • GPT-6 Luna costs $0.15 per task compared to $2.57 per task for 5.6 Luna on equivalent workloads.
  • Input and output API prices decreased by 50% across the board moving from the 5.6 series to the 6 series.
  • GPT-6 Luna input tokens cost $0.10 and output tokens cost $0.50.

Upgrading models from version 5.6 to 6 provides moderate benchmarks gains alongside steep efficiency improvements. Comparative tasks that ran at $2.57 on older Luna variants execute at $0.15 on the 6 architecture. Lower base token fees allow GPT-6 models to beat competitor variants such as Claude Fable 5.1 on cost-to-performance metrics.

Factuality Improvements and Prompt Caching Updates

  • Factual error rates for GPT-6 Sol decreased to 4.6%, aligning closely with GPT-6 Astra at 3.9%.
  • Max-effort GPT-6 Luna reduced factual errors from 12% to 7.6%, while low-effort Luna error rates dropped from 36% to 27%.
  • Reasoning effort adjustments no longer reset active prompt caches, preserving the 90% warm-cache discount.

Model halluncinations show measurable declines across all effort levels compared to previous iterations. Developer API costs drop further via updated prompt caching mechanics. Changing reasoning parameters mid-conversation maintains the cache state, preventing full context re-transmissions that trigger 100% input token rates. A new API dashboard provides breakpoint optimization for cached system prompts.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video