5 Things RUINING Your GPT 6 Astra Results (And How To Fix Them)
CChase AI
Computing/SoftwareInternet Technology
Transcript
00:00:00GPT-6 Astra is giving Anthropic a run for its money, and for good reason.
00:00:05This is the greatest AI model we have ever seen. But if you are using this model the same way
00:00:10you've used older models in the past, you are severely hamstringing your results. But today,
00:00:17in this video, I'm going to be going over the five mistakes you are making when it comes to
00:00:22your Astra usage, and more importantly, I'm going to show you how to fix them.
00:00:26Now, the first mistake you're making when it comes to Astra is you are using the wrong
00:00:30effort level. It is unfortunate how many people I have seen open up codecs, throw effort level to max,
00:00:36or even if they're complete freak shows, throw it to ultra and think this is going to automatically
00:00:41get you better outputs. Spoiler alert, it is not. We need to be very conscientious of what sort of
00:00:47effort level we are taking for our particular project. And unless you're doing some sort of
00:00:52wild, extremely complicated projects, you probably shouldn't ever be going above high. And in many
00:00:58cases, you should probably be on medium or even light. And the stats kind of prove this.
00:01:03This is illustrated very well here with the Deep Swee benchmark, which is a benchmark that's all about
00:01:08long running agentic tasks, the type of tasks that you would expect the much higher effort levels to
00:01:14really thrive in when we compare it to the lower effort levels. Yet, looking at Deep Swee,
00:01:19and I am on the max setting, I get a score of 73%. And my average cost per task is $12. However,
00:01:27on the exact opposite side of the spectrum, if I go to low, I'm at 67%, which is only a 6% drop off.
00:01:35Yet, I've gone from $12 per task to $2.19 per task. Now I know most of you are on a subscription plan,
00:01:43we're talking about usage. But the fact remains, the amount of weekly usage you will burn on something
00:01:47like low will be significantly less than on high, or on max rather. And as we go down the line from max
00:01:55to extra high, you see we actually got better results on extra high, yet the cost per task was
00:02:00almost half. We see that again, when we go to high, we are at the same exact score as max, yet we go
00:02:06from $12 to $5.72. And at medium, we're basically the same exact place. What's the point here? The point
00:02:13is Astra is extremely efficient, even at lower settings. And if we compare that to something
00:02:19like Fable 5, the low setting on GPT-6 Astra is basically right in between Fable 5 high and medium.
00:02:25So that would be about a $7 price point on the anthropic side. Again, $2.19 with Astra. And this
00:02:31idea is repeated across multiple benchmarks. Here's the artificial analysis coding agent index.
00:02:36At low setting, we're at $1.50 at a 62.6 score. On max, we're at $4.98 per cost and a 67 score. So a
00:02:48difference of 4.4 for the score, yet the price jumps up almost $3.50. Now, on some of these other
00:02:57benchmarks, we see there is a dip as we go from low to max. Like max will give you better outputs on these
00:03:04far ends. But even here, right on terminal bench, high gives us a better score than max. And again,
00:03:10much cheaper. So what's the takeaway here? When you are inside of codex, don't just automatically
00:03:16throw this bar for effort level all the way to the right. In fact, you can probably get away
00:03:21on average with high. Now, I went ahead and tested this out on front end design, which is a common use
00:03:26case for Astra. I gave Astra the same prompt with two different effort levels. This was the prompt
00:03:32on light. I said I wanted to create a homepage for Dune House, a fictional boutique desert hotel in
00:03:38Joshua Tree, California. And this is what it created for us. Honestly, pretty solid. Again, this was using
00:03:45the light effort level in terms of the time it took. I gave it the prompt at 341. And 13 minutes later,
00:03:52it was complete. Total tokens used was 91,000. And here's a look at the same result on max setting.
00:03:59This time it took 30 minutes to complete. We used 152,000 tokens. And this was the end result.
00:04:06And I mean, this looks pretty good, but is it significantly better than light? You can make
00:04:14an argument for both. Point being, is there a huge difference in the outcome here for this particular
00:04:19task? Not really. So when in doubt, when we're working with effort levels, less is more. If you
00:04:24aren't getting the output you like, go ahead and begin incrementally bumping up that effort level. But
00:04:29there is essentially zero reason why you should start at extra high, max, or ultra. Now, the second mistake
00:04:35you are making when it comes to Astra is you are completely under utilizing its browser and computer
00:04:40use. These two things allow us to sort of connect Astra to applications where we don't have a readily
00:04:48available CLI, MCP, or API. For example, let's say we want to improve on this website we created. This
00:04:55was the version we got with the max effort level. And instead of me manually going out on the web, finding
00:05:02new references and feeding it to Astra, why don't I have Astra do it on its own? Why don't I send it to
00:05:08a website like dribble, which doesn't have an API that I have access to at least, and have it find references
00:05:14for other hotel type websites, download the images or take screenshots of the images and bring it into the
00:05:20fold for its next iteration. It can do all that automatically. I don't have to touch anything.
00:05:25So that prompt will sound something like this. Hey, so can you in the browser on the right hand
00:05:31side, head to dribble, that's D-R-I-B-B-B-L-E.com, and then look up some websites, some references for
00:05:41hotel websites. Because what I want you to do is I want you to go onto dribble, I want you to search for
00:05:45hotel websites. And I want you to find references of other hotel websites that look very visually
00:05:50stunning. Take screenshots of them, do what you need to do. So then you can bring those reference
00:05:55images back into Astra, back into Codex, and then come up with a new version of our website.
00:06:02So we can see here over on the right is now pulled up dribble. And it has my login because it actually
00:06:08saves those when you log in at all. It's now searching for hotel website. It's then clicking on
00:06:14individual websites that were listed there. It's capturing screenshots. And then it's continuing
00:06:18this process with more websites that it thinks sort of fits the bill. It then put all that together to
00:06:23create this website. And I'll turn off my camera so you can see it better. So it generated the brand
00:06:28new website. And like Codex always does, it also made sure it worked on mobile, ran all the tests.
00:06:33But we got something a little bit different. And to be honest, I kind of like this version better. And I
00:06:39like the very prominent image it generated as well. And normally we would do this completely manually
00:06:46in terms of finding all these references. But you could see how it'd be so easy to scale this where
00:06:50maybe you want to just work at look at dribble, it could look at Pinterest, it could look at Twitter,
00:06:53or we could play this to really any scenario where we have some sort of application that we want to grab
00:06:58information from and bring in the Codex. But again, we don't have an API, we don't have an MCP.
00:07:03This is where these browser automations and computer use tools are so handy. Astra even created sort of a
00:07:09reference table so I can see which screenshots it actually used as well sort of its notes and links
00:07:14to the original source on dribble, which is also nice. Now, before we go into the third mistake,
00:07:19a quick word from today's sponsor, me. So inside of Chase AI plus, I have released not only a brand new
00:07:25cloud code masterclass, I also have a codex masterclass as well. So no matter your technical background or lack
00:07:32thereof, I will teach you how to master these AI tools, focus on real use cases, I post updates
00:07:38every single week. So if this is something you really want to dive into, Chase AI plus is the
00:07:43place for you. There is a link to it in the pinned comment. Now the third mistake you are making when
00:07:48it comes to Astra is that your skills are all wrong, your skills are holding you back. Because we are in
00:07:54the same place now with Codex that we were at with Cloud Code just a few weeks ago. You remember when Boris
00:07:59Cherney, the maker of Cloud Code came out and said, you need to delete your cloud.md, you need to delete
00:08:03your skills. Well, the idea there wasn't that we should just delete them for the sake of it.
00:08:08It was that these models, Astra and Fable have gotten so good that many of the skills that acted as
00:08:15scaffolding for older models we used to play around with just are irrelevant now. And in fact,
00:08:20in many cases, they are holding you back. Now this isn't black and white, although I will say a lot
00:08:24of the super heavy scaffolding skills, things like superpowers and GSD should probably go by the
00:08:29wayside. But even if you disagree with that statement, your context window is probably clogged
00:08:35with a ton of skills that you just don't even use. How many skills did you install nine, six,
00:08:40three months ago that you haven't touched since then? Have you gotten rid of them? If you haven't,
00:08:44they're still there. Like it's just their description, but these add up. You might have
00:08:4730, 40, 50 skills that have nothing to do with what you do anymore. And they're just clogging up
00:08:53your context window. So what do we need to do? Well, we need to do an audit. And I created a skill
00:08:58that does that for you. This is the skill audit. And this is actually based on Anthropic's skill
00:09:03creator skill because it has benchmarks and testing in place where if you pointed at a specific skill,
00:09:09it will run tests to see, does this skill make sense with the current model? Codex has its own
00:09:15skill creator skill, but it wasn't as robust as Anthropic. So that's why I created this and this
00:09:20is what it does. So I'll put a link down below to where you can find it. I show you the install,
00:09:26and then I also give you a prompt you can run. This prompt will then go through all of your skills
00:09:31and your logs and figure out which skills have you not been using at all, and that we should probably
00:09:35prune. And it's going to take a look at the front matter of your skills, see what
00:09:39descriptions make sense, and then also have a list of skills for you to take a look at, where if you
00:09:43want to, you can go one by one through these skills and actually run this skill audit benchmark test
00:09:48against them to see, do they make sense with Astra? And so group solve your skills into a bunch of
00:09:53different buckets, either fix now, review for retirement, text next, test later, or preserve. Again,
00:09:59this is super simple for you to use. You're just going to install the skill and then run this prompt.
00:10:03And this will buy you a few things. One, it's going to free up some of our context window that has been bloated
00:10:08through skills we don't even use. And then two, for those skills that are in the gray area of, well,
00:10:13we still want to use them, but we're not sure if they're legit or can we improve them? It's going to
00:10:17improve them. It's going to benchmark them. And you're no longer going to have a question of, do these
00:10:21actually help me? Now, the fourth mistake you are making when it comes to Astra is you are not using
00:10:25its voice mode and its voice mode is best in class. What we get here inside of Codex destroys its
00:10:32cousin inside of Anthropic's Cloud Code desktop. And it got improved with Astra. Before, when you use
00:10:38voice mode and you use it by just clicking this little thing right here and this bobble is going to pop up,
00:10:43before it was powered by GPT Terra. And I believe it was Terra on the light effort setting. Now you can
00:10:49use it with Astra and you can use it with Astra high or Astra low. Furthermore, if I wanted to use voice
00:10:54before, so let's say I want to start a new voice chat, it would be in an entirely different chat
00:10:59panel like you see here. And I'd have to use it as an orchestrator. So I would say, hey, go do X, Y,
00:11:05and Z. And now it's kind of listening to me right now. I would say, hey, go do X, Y, and Z. And it would
00:11:10open up a new chat and do that. Now what I can do is I can go into an individual chat, like the one we were
00:11:16just at, and I can use voice mode inside of here. So I have multiple options, orchestrator that controls
00:11:21multiple chats, or I can use it inside an individual chat. And I have Astra. So for example,
00:11:28if I wanted to use it inside of here, I'm just going to click this thing. And let's say I wanted to,
00:11:34let's say I wanted to adjust something, let's say I wanted it to add some sort of like form somebody
00:11:40could fill out on this web page over here on the right. And then I also wanted it to test it out.
00:11:45So let's try that. So I'm going to click on this. And I will say when it comes to how you should best use
00:11:50it, if you are using it to do things. So I'm inside a chat right now, I'm having it do something,
00:11:55which is add the form, I'm probably going to put it on high effort. If I'm using it as an orchestrator,
00:12:00or I just want to kind of talk to it back and forth, I'll probably just put it on light. So we're going
00:12:05to do voice. Throw this on high. And I'm just going to click it and talk. So what I want you to do right
00:12:17now is I want you to take a look at this website we've built on the right hand side of the browser.
00:12:21I kind of would like some place where people can fill out just a form if they want more information,
00:12:26probably at the bottom near the footer, and kind of figure out what best practices are for that,
00:12:32what sort of information they should put in there. Right now, we don't need the functionality to
00:12:36actually work. I don't need you to hook it up with resend or anything. I kind of just wanted to see what
00:12:40it will look like on the web page. Can you go ahead and do that for me?
00:12:44Sure, let me take a look.
00:12:48And also, I want you to notice kind of stop that for a second. I also want you to notice how quick
00:12:55and snappy that was when it said, Sure, let me take a look. If you haven't played with the voice
00:12:58mode here, it is extremely snappy and very responsive. And so while that voice mode is
00:13:03working, let me sort of demo what else we can do with it. So I'm here in a new voice chat,
00:13:07I'm going to set this to light because I'm going to have it act as an orchestrator. And I can pretty much
00:13:12tell it to like open up a new chat over here on the left, use Astra, and we'll say something like,
00:13:17Hey, can you figure out what the top five GBT voice, you know, use cases are?
00:13:25Can you go ahead and spin up a new chat window over on the left hand side doesn't need to be a new
00:13:31project. Get Astra working on doing some research for us to figure out what are the top five use cases for
00:13:38the new GPT six Astra voice mode, specifically looking at the voice mode to open up a new chat
00:13:43with that. And I also want you to tell that agent to write up a little document for us like HTML or
00:13:48something on it, setting that up now. So you can see here, it's now created that chat, I can open up
00:13:56that chat over here on the left hand side top five voice models. And you can see right here, it says,
00:14:01sent by a new task focused on the top Astra voice mode use cases, plus a short HTML work up with
00:14:08examples and sources. And so you can see here the prompt that was sent to this chat. And inside of
00:14:15this chat window, like I said, this can act as an orchestrator, I can have it create a bunch of
00:14:19different chats, I can control everything from this one voice pane. So this is super useful if you're
00:14:24someone who's like, has 23456 plus agents going on all at once. And instead of trying to track all
00:14:30those manually, again, have Astra track all them for you. And just talk to this one orchestrator. And
00:14:35we can see over here with our original command, it is now added, and it's working on our little signup
00:14:42sheet right here, more information sheet, you can see sort of the browser control in action with a little
00:14:47cursor, it's actually testing it out as if it was a person. Done. The HTML brief is ready to open.
00:14:53And it separates live voice handling from Astra's role doing the underlying work, which is a handy way
00:14:58to frame it. So here you can see the five use cases it found. And really, I think the big unlock with
00:15:04voice chat is using it as an orchestrator. Now the fifth and final mistake you are making with Astra is
00:15:10you are prompting it wrong. Now this model is extremely effective, but there are some differences with
00:15:16how GPT-6 Astra handles things compared to GPT-5.6 Sol. Specifically, when it comes to making
00:15:22assumptions, when it comes to forks in the road, this is coming from OpenAI themselves. And what
00:15:29they tell us is that Astra is more likely to ask for clarification where older, earlier models would
00:15:35make assumptions. What does this mean for you? Well, this means when you are coming up with your prompts,
00:15:40especially when we're talking about long running agentic complex tasks,
00:15:45you should probably put in some sort of verbiage there about how you want it to handle these forks
00:15:50in the road. Do you want it to always ask you questions? Or do you want to sort of carry the
00:15:56user's intended task to completion? Now OpenAI gives us a specific prompt you can use, which is listed
00:16:03right here. Or you can use the template I'm going to put on the screen right now to help guide you
00:16:07for some of these bigger tasks. The big thing here is that you just need to know how you want it to
00:16:12behave. Are you someone who likes it when AI is constantly checking in, which Astra will tend to
00:16:18do on its own? Or when we're doing longer stuff, do you want to just do its thing? I'm going to give
00:16:23you a North Star, I'm going to give you some sort of end state, go forth and conquer, don't ask me
00:16:28questions, figure out. There's pros and cons to both of these solutions. Just know when you take the
00:16:33route of just, Hey, go ahead and figure it out. That's when you tend to get sort of a regression
00:16:39to the mean, unless you set up a lot of scaffolding that sort of pointed in certain directions when it
00:16:43reaches said forks in the road. Another good option you have here is this prompt that OpenAI gives us
00:16:49where we're going to tell the model to ask for approval only after preparing a concrete reviewable
00:16:54result. This enjoys blocking the task. If it could already have done it, right? Don't just come to me
00:16:59with a problem, come to me with a problem and your potential solution. So I think some combination of
00:17:04all of these and what I showed earlier will get you to a good place. If you're having some problems with
00:17:09Astros or just like stuttering its way through tasks. So those are the five mistakes that are holding
00:17:15you back when it comes to GPT-6 after this model is extremely powerful, especially when we look at it
00:17:21inside of Codex because Codex has so many cool little features that I think Anthropic and the Cloud Code
00:17:26Desktop app are sort of falling behind on namely browser use, computer use, and the voice bone. So
00:17:33make sure to try those out. Let me know how they worked for you. As always, if you want to learn more
00:17:39about Codex and Cloud Code, make sure to check out Chase AI+. I'll put a link to that down in the pin
00:17:44comment and I'll see you around.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video