Combine Fable 5 & Sol 5.6 With This One Skill (Or Fall Behind)

CChase AI
Computing/SoftwareManagementInternet Technology

Transcript

00:00:00GPT 5.6 aka Sol is coming out tomorrow and the big question on everyone's mind is
00:00:05does this new model beat Claude Fable? Well I think that's the wrong question to ask because
00:00:11instead of trying to figure out which one of these models is better we should be asking how can we
00:00:16use these two powerful models together? And in today's video I'm going to be giving you a skill
00:00:20that does exactly that. This skill includes a supercharged plan mode based on Matt Pocock's
00:00:26grill me. We then have an adversarial planning session where Claude and Codex go head-to-head
00:00:30till they come to a conclusion. Then once we're ready we take that Fable driven plan and hand it
00:00:37off to Codex which starting tomorrow will include Sol 5.6. After Codex goes to work then we have Fable
00:00:44review the entire process. All in all we're not just getting the best of both of these models we're
00:00:49also saving tokens in the aggregate versus having Fable do everything. So I'm going to show you how
00:00:54this works we're going to do a quick demo and then I'll be giving you the skill. So why should we even
00:00:58care about creating some sort of skill where Fable does most of the planning and then we pass it off
00:01:02to Sol 5.6? Well first reason is Sol 5.6 is wildly powerful at least according to the benchmarks.
00:01:10Now grain of salt this is coming from OpenAI but when we look at Sol 5.6 Ultra and just standard
00:01:175.6 Sol we see numbers on Terminal Bench 2.1 that put it ahead of Claude mythos let alone Fable 5.
00:01:24The second reason is token efficiency. There is a real reason why you see so much content around
00:01:29hey how can we reduce Fable's usage and the types of things you see are like advisor mode you know
00:01:35essentially having Fable plan and have Opus execute. Well why have Opus execute if with at the same price
00:01:42I could have 5.6 do it or 5.5 do it. Point is we're doing that same construct but with a better model
00:01:49and arguably a cheaper model than Opus. When we look at 5.6 it is more token efficient than 5.5 which
00:01:56was more token efficient than Opus 4.6 and we can see that in the data. What we're looking at right here
00:02:02is GPT 5.5 on extra high. Its pass rate on this benchmark was 23% at $1.24. When I look at 5.6
00:02:12it's 25% score so higher score at 56 cents so way cheaper and therefore way more efficient and when we
00:02:20look at direct comparisons of 5.5 versus Opus 4.8 there's really no contest. Higher pass rates lower cost.
00:02:28So we're essentially taking that same idea and just ramping it up with this 5.6 improvement.
00:02:33So how does the skill actually work? Well I actually have a couple skills for you. In a vacuum we have
00:02:38the codex build skill. This is the idea that you created a plan with Fable and codex is just going
00:02:43to go ahead and build that particular feature or particular product. I also have included an updated
00:02:48grill me codex. Now I've done a video on this skill before and what we've done is we've added on
00:02:54this idea of GPT 5.6 actually going out there and building things for us and so when we look at the
00:03:01more comprehensive grill me codex which is the big skill it occurs in four stages. The idea is you have
00:03:08some sort of project some sort of feature you want to start and you kick it off with grill me codex and
00:03:12the first thing that happens is an interview. This interview is literally the grill me skill from
00:03:18Matt Pocock. So it is a plan mode on steroids. It goes way way deeper than cloud code normally would
00:03:23and we do this with Fable. All right so Fable's driving the ship here. Secondly we have adversarial
00:03:30planning. So Fable's come up with the plan. We then take that Fable plan and we push it over to codex.
00:03:37Now in today's video that's going to be 5.5 but tomorrow that will be 5.6 and Fable and codex go back
00:03:43and forth for a maximum of five iterations where Fable says hey here's the plan. Codex says okay looks good
00:03:49except x y and z. Then Fable says uh I agree I disagree and they go back and forth till they reach
00:03:54a consensus. Now once they reach that consensus and this is where the upgrades happen is we now push
00:04:01the actual build to codex to 5.5 today and 5.6 tomorrow. I think this is way better than passing
00:04:09things off to opus or to sonnet or using advisor mode inside of cloud code because these gpt models are
00:04:14just better than those smaller anthropic models and they are cheaper so it really is a scenario where
00:04:21unless you just are super anti-gpt and anti-codex it's hard to argue otherwise especially if we get
00:04:27to a place where Fable's like kind of off the market and lastly once codex finishes the build Fable is
00:04:32going to come in and it's going to review what it did and it's going to go through a maximum of two
00:04:37sort of iterations where let's say Fable thinks codex did something wrong it's going to say hey codex you did
00:04:42that wrong fix it it's going to do that twice if by the third time it's not complete well then Fable
00:04:47will clean it up itself so this is the process by which i think we get the best of open ai and
00:04:54anthropic now before we hop into the demo a quick word from today's sponsor me so i just released my
00:04:59cloud code master class inside of chase ai plus and it is the number one way to go from zero to ai dev
00:05:04especially if you don't come from a technical background i update this every single week we focus
00:05:09on real use cases and it also includes a codex master class as well so if you want to get a little
00:05:15bit more serious about ai and you have no idea where to begin this is the place for you there will be a
00:05:19link in the pin comment now installing and using the skill is pretty straightforward i will put a link to
00:05:24the github in the description now to use this we're just going to do forward slash grill me codex
00:05:28and we just give it a prompt what it is we're trying to build so we're trying to build trip atlas which
00:05:33is a stylized cinematic trip planner web app and i go into a little bit more details about what i
00:05:38want it to be right i want it to look kind of cool i can put in the different places i'm going all that
00:05:42and once i do this what's going to happen is it's going to kick off the grill me section
00:05:46of the plan which if you're familiar with matt pocock's work it essentially is just a plan mode on
00:05:51steroids it's going to ask me like eight nine ten different questions i go a lot deeper than your standard
00:05:56plan mode stuff so it's asking me what is this for we're going to say this is for a real personal
00:06:01tool not just a video demo and for each of these it also gives its recommendations so if you're
00:06:06confused about what i should choose and why that's all spelled out for you now it's asking about geocoding
00:06:11and it's going to continue to go down these series of questions until it's happy with what we're
00:06:15creating now i'm going to skip through the rest of the questions because you can imagine
00:06:19what the next seven or eight questions will look like and we'll move into the adversarial planning
00:06:24stage so you can see here it's written the plan and it also creates a markdown file where it logs
00:06:28all the back and forth between codex and cloud code and so right now we are on round one where it's
00:06:34passing it off to gpt and so you can see them kind of going back and forth here on the log but in this
00:06:41case it only took them two rounds before it was approved and so we can see what the two acts improved
00:06:47you know lock the identity a real person tool kind of what's going to be the actual sort of stack and
00:06:53then it had 12 findings in the second round related to like hardening the data core now once it's
00:06:59completed this back and forth you have a few options either codex is going to build it kind of what we've
00:07:05talked about from the beginning we have the option just having claude build it so for whatever reason
00:07:09like i don't want to bring gpt in you can keep it with fable or you can stop here but we're going to go ahead and
00:07:14like codex build this and again we can kind of go back and forth if you think well gpt 5.6 is going
00:07:20to be better than fable or 5.5 versus opus 4.8 at the end of the day the real value that can't really
00:07:26be argued with is going to be the token efficiency especially if 5.6 is even close to what the benchmarks
00:07:34are claiming so codex has finished up its build and you can see now what's happening is the review stage
00:07:40so now fable is going through everything codex has built and then it's going to go back to codex
00:07:44and say this is wrong this was right remember it'll do two iterations of that before it's like hey
00:07:49i want to drive it'll take the wheel and it will start writing the code itself now fable is done with
00:07:54its review it said there are a couple deviations which it felt were all reasonable goes over the
00:07:59files and all this and now it's asking hey do you want to commit or you want to take a look at it so
00:08:02let's take a look at what it actually built and so here's what we got so over here we have sort of
00:08:08a map of the world and it looks like it created some custom graphics using the gpt image generator
00:08:14so you're able to like name the trip you can add stops over here on the left you can put where
00:08:20you're going and then sort of like what you're going to do at those different locations it also
00:08:25has this cinematic replay and i'm just going to mute this so let's see what happens here
00:08:31so it looks like i'll move over here you can see sort of this weird plane hopping from spot to spot
00:08:38which it looks like it created as an svg there's a little passport stamps and boom there we go so
00:08:45you know there's a lot we could do here to kind of make it look i think better but in general it built
00:08:52what we said we wanted to right like everything actually works here you know if i delete things on
00:08:57here delete some i can move them up down i can change stuff let's say we added tokyo all of a sudden
00:09:04it actually shows how far away that stop is that's interesting if i add that to the route there we go
00:09:11so you know i actually built this out i think it'd be a good not bad i think for the first pass and
00:09:17what this really was about was just showing this workflow in action and you can also see down here
00:09:21in terms of our usage we only burned up about 130 000 tokens on the fable side to get this whole
00:09:26thing done so that's the skill in action hopefully you get a ton of use out of this one's 5.6 drops
00:09:31as always let me know what you thought about this video in the comments make sure to check out chase ai plus
00:09:37and i'll see you around

Key Takeaway

By offloading the build and review processes to a more token-efficient model like Codex while retaining Fable for planning, developers can significantly reduce costs and improve overall output quality.

Highlights

  • Integrating Fable and Codex models allows for a more token-efficient workflow than relying solely on Fable for all tasks.

  • The 5.6 model version achieves a 25% pass rate on Terminal Bench 2.1 at a cost of 56 cents, compared to the 5.5 version's 23% pass rate at $1.24.

  • The workflow uses Fable for initial planning via a 'grill me' interview, followed by an adversarial planning stage where Fable and Codex debate the plan for up to five iterations.

  • Codex handles the actual build phase, and Fable concludes by reviewing the generated code, performing up to two correction iterations before fixing any remaining issues itself.

  • This multi-model approach required only 130,000 tokens on the Fable side to fully develop a functional web application.

Timeline

Strategy for Multi-Model Integration

  • Focus on using models together rather than picking a single superior model.
  • 5.6 models demonstrate higher token efficiency and lower costs compared to 5.5 and earlier iterations.
  • Terminal Bench 2.1 data shows 5.6 models achieving higher pass rates at nearly half the cost of 5.5.

The approach centers on optimizing token usage by leveraging the distinct strengths of different models. Benchmarks indicate that newer iterations provide better performance at lower price points, making them ideal for execution tasks, while Fable remains highly effective for the initial planning phase.

Four-Stage Workflow Implementation

  • The process consists of an interview, adversarial planning, execution, and final review.
  • Fable conducts a detailed 'grill me' interview to generate an initial, deep project plan.
  • Adversarial planning forces Fable and Codex to negotiate until they reach a consensus, ensuring a robust plan before building begins.
  • Fable reviews the completed build from Codex for up to two cycles, providing corrections before finalizing the code.

This structured workflow starts with a specialized interview to define project requirements. Once the plan is established, it undergoes a critical negotiation phase between models to eliminate potential flaws. Codex then executes the build based on this refined plan, followed by an automated review process by Fable to ensure code quality.

Practical Application and Demo

  • The workflow is initiated by the command /grill me codex followed by project requirements.
  • The system generates a markdown log detailing the adversarial negotiation process between Fable and Codex.
  • A demo application, Trip Atlas, was successfully built using this workflow, demonstrating functional custom graphics and routing features.
  • The entire development process for the application consumed approximately 130,000 tokens on the Fable side.

Executing the skill involves entering project requirements to trigger the structured four-stage process. The system tracks all interactions in a markdown file, providing transparency into how the models arrived at the final build. The resulting web application confirms that this workflow effectively produces functional software while maintaining token efficiency.

Community Posts

View all posts