5 Things RUINING Your GPT 6 Astra Results (And How To Fix Them)

English
CChase AI
컴퓨터/소프트웨어AI/미래기술

스크립트

00:00:00GPT-6 Astra is giving Anthropic a run for its money, and for good reason.
00:00:05This is the greatest AI model we have ever seen. But if you are using this model the same way
00:00:10you've used older models in the past, you are severely hamstringing your results. But today,
00:00:17in this video, I'm going to be going over the five mistakes you are making when it comes to
00:00:22your Astra usage, and more importantly, I'm going to show you how to fix them.
00:00:26Now, the first mistake you're making when it comes to Astra is you are using the wrong
00:00:30effort level. It is unfortunate how many people I have seen open up codecs, throw effort level to max,
00:00:36or even if they're complete freak shows, throw it to ultra and think this is going to automatically
00:00:41get you better outputs. Spoiler alert, it is not. We need to be very conscientious of what sort of
00:00:47effort level we are taking for our particular project. And unless you're doing some sort of
00:00:52wild, extremely complicated projects, you probably shouldn't ever be going above high. And in many
00:00:58cases, you should probably be on medium or even light. And the stats kind of prove this.
00:01:03This is illustrated very well here with the Deep Swee benchmark, which is a benchmark that's all about
00:01:08long running agentic tasks, the type of tasks that you would expect the much higher effort levels to
00:01:14really thrive in when we compare it to the lower effort levels. Yet, looking at Deep Swee,
00:01:19and I am on the max setting, I get a score of 73%. And my average cost per task is $12. However,
00:01:27on the exact opposite side of the spectrum, if I go to low, I'm at 67%, which is only a 6% drop off.
00:01:35Yet, I've gone from $12 per task to $2.19 per task. Now I know most of you are on a subscription plan,
00:01:43we're talking about usage. But the fact remains, the amount of weekly usage you will burn on something
00:01:47like low will be significantly less than on high, or on max rather. And as we go down the line from max
00:01:55to extra high, you see we actually got better results on extra high, yet the cost per task was
00:02:00almost half. We see that again, when we go to high, we are at the same exact score as max, yet we go
00:02:06from $12 to $5.72. And at medium, we're basically the same exact place. What's the point here? The point
00:02:13is Astra is extremely efficient, even at lower settings. And if we compare that to something
00:02:19like Fable 5, the low setting on GPT-6 Astra is basically right in between Fable 5 high and medium.
00:02:25So that would be about a $7 price point on the anthropic side. Again, $2.19 with Astra. And this
00:02:31idea is repeated across multiple benchmarks. Here's the artificial analysis coding agent index.
00:02:36At low setting, we're at $1.50 at a 62.6 score. On max, we're at $4.98 per cost and a 67 score. So a
00:02:48difference of 4.4 for the score, yet the price jumps up almost $3.50. Now, on some of these other
00:02:57benchmarks, we see there is a dip as we go from low to max. Like max will give you better outputs on these
00:03:04far ends. But even here, right on terminal bench, high gives us a better score than max. And again,
00:03:10much cheaper. So what's the takeaway here? When you are inside of codex, don't just automatically
00:03:16throw this bar for effort level all the way to the right. In fact, you can probably get away
00:03:21on average with high. Now, I went ahead and tested this out on front end design, which is a common use
00:03:26case for Astra. I gave Astra the same prompt with two different effort levels. This was the prompt
00:03:32on light. I said I wanted to create a homepage for Dune House, a fictional boutique desert hotel in
00:03:38Joshua Tree, California. And this is what it created for us. Honestly, pretty solid. Again, this was using
00:03:45the light effort level in terms of the time it took. I gave it the prompt at 341. And 13 minutes later,
00:03:52it was complete. Total tokens used was 91,000. And here's a look at the same result on max setting.
00:03:59This time it took 30 minutes to complete. We used 152,000 tokens. And this was the end result.
00:04:06And I mean, this looks pretty good, but is it significantly better than light? You can make
00:04:14an argument for both. Point being, is there a huge difference in the outcome here for this particular
00:04:19task? Not really. So when in doubt, when we're working with effort levels, less is more. If you
00:04:24aren't getting the output you like, go ahead and begin incrementally bumping up that effort level. But
00:04:29there is essentially zero reason why you should start at extra high, max, or ultra. Now, the second mistake
00:04:35you are making when it comes to Astra is you are completely under utilizing its browser and computer
00:04:40use. These two things allow us to sort of connect Astra to applications where we don't have a readily
00:04:48available CLI, MCP, or API. For example, let's say we want to improve on this website we created. This
00:04:55was the version we got with the max effort level. And instead of me manually going out on the web, finding
00:05:02new references and feeding it to Astra, why don't I have Astra do it on its own? Why don't I send it to
00:05:08a website like dribble, which doesn't have an API that I have access to at least, and have it find references
00:05:14for other hotel type websites, download the images or take screenshots of the images and bring it into the
00:05:20fold for its next iteration. It can do all that automatically. I don't have to touch anything.
00:05:25So that prompt will sound something like this. Hey, so can you in the browser on the right hand
00:05:31side, head to dribble, that's D-R-I-B-B-B-L-E.com, and then look up some websites, some references for
00:05:41hotel websites. Because what I want you to do is I want you to go onto dribble, I want you to search for
00:05:45hotel websites. And I want you to find references of other hotel websites that look very visually
00:05:50stunning. Take screenshots of them, do what you need to do. So then you can bring those reference
00:05:55images back into Astra, back into Codex, and then come up with a new version of our website.
00:06:02So we can see here over on the right is now pulled up dribble. And it has my login because it actually
00:06:08saves those when you log in at all. It's now searching for hotel website. It's then clicking on
00:06:14individual websites that were listed there. It's capturing screenshots. And then it's continuing
00:06:18this process with more websites that it thinks sort of fits the bill. It then put all that together to
00:06:23create this website. And I'll turn off my camera so you can see it better. So it generated the brand
00:06:28new website. And like Codex always does, it also made sure it worked on mobile, ran all the tests.
00:06:33But we got something a little bit different. And to be honest, I kind of like this version better. And I
00:06:39like the very prominent image it generated as well. And normally we would do this completely manually
00:06:46in terms of finding all these references. But you could see how it'd be so easy to scale this where
00:06:50maybe you want to just work at look at dribble, it could look at Pinterest, it could look at Twitter,
00:06:53or we could play this to really any scenario where we have some sort of application that we want to grab
00:06:58information from and bring in the Codex. But again, we don't have an API, we don't have an MCP.
00:07:03This is where these browser automations and computer use tools are so handy. Astra even created sort of a
00:07:09reference table so I can see which screenshots it actually used as well sort of its notes and links
00:07:14to the original source on dribble, which is also nice. Now, before we go into the third mistake,
00:07:19a quick word from today's sponsor, me. So inside of Chase AI plus, I have released not only a brand new
00:07:25cloud code masterclass, I also have a codex masterclass as well. So no matter your technical background or lack
00:07:32thereof, I will teach you how to master these AI tools, focus on real use cases, I post updates
00:07:38every single week. So if this is something you really want to dive into, Chase AI plus is the
00:07:43place for you. There is a link to it in the pinned comment. Now the third mistake you are making when
00:07:48it comes to Astra is that your skills are all wrong, your skills are holding you back. Because we are in
00:07:54the same place now with Codex that we were at with Cloud Code just a few weeks ago. You remember when Boris
00:07:59Cherney, the maker of Cloud Code came out and said, you need to delete your cloud.md, you need to delete
00:08:03your skills. Well, the idea there wasn't that we should just delete them for the sake of it.
00:08:08It was that these models, Astra and Fable have gotten so good that many of the skills that acted as
00:08:15scaffolding for older models we used to play around with just are irrelevant now. And in fact,
00:08:20in many cases, they are holding you back. Now this isn't black and white, although I will say a lot
00:08:24of the super heavy scaffolding skills, things like superpowers and GSD should probably go by the
00:08:29wayside. But even if you disagree with that statement, your context window is probably clogged
00:08:35with a ton of skills that you just don't even use. How many skills did you install nine, six,
00:08:40three months ago that you haven't touched since then? Have you gotten rid of them? If you haven't,
00:08:44they're still there. Like it's just their description, but these add up. You might have
00:08:4730, 40, 50 skills that have nothing to do with what you do anymore. And they're just clogging up
00:08:53your context window. So what do we need to do? Well, we need to do an audit. And I created a skill
00:08:58that does that for you. This is the skill audit. And this is actually based on Anthropic's skill
00:09:03creator skill because it has benchmarks and testing in place where if you pointed at a specific skill,
00:09:09it will run tests to see, does this skill make sense with the current model? Codex has its own
00:09:15skill creator skill, but it wasn't as robust as Anthropic. So that's why I created this and this
00:09:20is what it does. So I'll put a link down below to where you can find it. I show you the install,
00:09:26and then I also give you a prompt you can run. This prompt will then go through all of your skills
00:09:31and your logs and figure out which skills have you not been using at all, and that we should probably
00:09:35prune. And it's going to take a look at the front matter of your skills, see what
00:09:39descriptions make sense, and then also have a list of skills for you to take a look at, where if you
00:09:43want to, you can go one by one through these skills and actually run this skill audit benchmark test
00:09:48against them to see, do they make sense with Astra? And so group solve your skills into a bunch of
00:09:53different buckets, either fix now, review for retirement, text next, test later, or preserve. Again,
00:09:59this is super simple for you to use. You're just going to install the skill and then run this prompt.
00:10:03And this will buy you a few things. One, it's going to free up some of our context window that has been bloated
00:10:08through skills we don't even use. And then two, for those skills that are in the gray area of, well,
00:10:13we still want to use them, but we're not sure if they're legit or can we improve them? It's going to
00:10:17improve them. It's going to benchmark them. And you're no longer going to have a question of, do these
00:10:21actually help me? Now, the fourth mistake you are making when it comes to Astra is you are not using
00:10:25its voice mode and its voice mode is best in class. What we get here inside of Codex destroys its
00:10:32cousin inside of Anthropic's Cloud Code desktop. And it got improved with Astra. Before, when you use
00:10:38voice mode and you use it by just clicking this little thing right here and this bobble is going to pop up,
00:10:43before it was powered by GPT Terra. And I believe it was Terra on the light effort setting. Now you can
00:10:49use it with Astra and you can use it with Astra high or Astra low. Furthermore, if I wanted to use voice
00:10:54before, so let's say I want to start a new voice chat, it would be in an entirely different chat
00:10:59panel like you see here. And I'd have to use it as an orchestrator. So I would say, hey, go do X, Y,
00:11:05and Z. And now it's kind of listening to me right now. I would say, hey, go do X, Y, and Z. And it would
00:11:10open up a new chat and do that. Now what I can do is I can go into an individual chat, like the one we were
00:11:16just at, and I can use voice mode inside of here. So I have multiple options, orchestrator that controls
00:11:21multiple chats, or I can use it inside an individual chat. And I have Astra. So for example,
00:11:28if I wanted to use it inside of here, I'm just going to click this thing. And let's say I wanted to,
00:11:34let's say I wanted to adjust something, let's say I wanted it to add some sort of like form somebody
00:11:40could fill out on this web page over here on the right. And then I also wanted it to test it out.
00:11:45So let's try that. So I'm going to click on this. And I will say when it comes to how you should best use
00:11:50it, if you are using it to do things. So I'm inside a chat right now, I'm having it do something,
00:11:55which is add the form, I'm probably going to put it on high effort. If I'm using it as an orchestrator,
00:12:00or I just want to kind of talk to it back and forth, I'll probably just put it on light. So we're going
00:12:05to do voice. Throw this on high. And I'm just going to click it and talk. So what I want you to do right
00:12:17now is I want you to take a look at this website we've built on the right hand side of the browser.
00:12:21I kind of would like some place where people can fill out just a form if they want more information,
00:12:26probably at the bottom near the footer, and kind of figure out what best practices are for that,
00:12:32what sort of information they should put in there. Right now, we don't need the functionality to
00:12:36actually work. I don't need you to hook it up with resend or anything. I kind of just wanted to see what
00:12:40it will look like on the web page. Can you go ahead and do that for me?
00:12:44Sure, let me take a look.
00:12:48And also, I want you to notice kind of stop that for a second. I also want you to notice how quick
00:12:55and snappy that was when it said, Sure, let me take a look. If you haven't played with the voice
00:12:58mode here, it is extremely snappy and very responsive. And so while that voice mode is
00:13:03working, let me sort of demo what else we can do with it. So I'm here in a new voice chat,
00:13:07I'm going to set this to light because I'm going to have it act as an orchestrator. And I can pretty much
00:13:12tell it to like open up a new chat over here on the left, use Astra, and we'll say something like,
00:13:17Hey, can you figure out what the top five GBT voice, you know, use cases are?
00:13:25Can you go ahead and spin up a new chat window over on the left hand side doesn't need to be a new
00:13:31project. Get Astra working on doing some research for us to figure out what are the top five use cases for
00:13:38the new GPT six Astra voice mode, specifically looking at the voice mode to open up a new chat
00:13:43with that. And I also want you to tell that agent to write up a little document for us like HTML or
00:13:48something on it, setting that up now. So you can see here, it's now created that chat, I can open up
00:13:56that chat over here on the left hand side top five voice models. And you can see right here, it says,
00:14:01sent by a new task focused on the top Astra voice mode use cases, plus a short HTML work up with
00:14:08examples and sources. And so you can see here the prompt that was sent to this chat. And inside of
00:14:15this chat window, like I said, this can act as an orchestrator, I can have it create a bunch of
00:14:19different chats, I can control everything from this one voice pane. So this is super useful if you're
00:14:24someone who's like, has 23456 plus agents going on all at once. And instead of trying to track all
00:14:30those manually, again, have Astra track all them for you. And just talk to this one orchestrator. And
00:14:35we can see over here with our original command, it is now added, and it's working on our little signup
00:14:42sheet right here, more information sheet, you can see sort of the browser control in action with a little
00:14:47cursor, it's actually testing it out as if it was a person. Done. The HTML brief is ready to open.
00:14:53And it separates live voice handling from Astra's role doing the underlying work, which is a handy way
00:14:58to frame it. So here you can see the five use cases it found. And really, I think the big unlock with
00:15:04voice chat is using it as an orchestrator. Now the fifth and final mistake you are making with Astra is
00:15:10you are prompting it wrong. Now this model is extremely effective, but there are some differences with
00:15:16how GPT-6 Astra handles things compared to GPT-5.6 Sol. Specifically, when it comes to making
00:15:22assumptions, when it comes to forks in the road, this is coming from OpenAI themselves. And what
00:15:29they tell us is that Astra is more likely to ask for clarification where older, earlier models would
00:15:35make assumptions. What does this mean for you? Well, this means when you are coming up with your prompts,
00:15:40especially when we're talking about long running agentic complex tasks,
00:15:45you should probably put in some sort of verbiage there about how you want it to handle these forks
00:15:50in the road. Do you want it to always ask you questions? Or do you want to sort of carry the
00:15:56user's intended task to completion? Now OpenAI gives us a specific prompt you can use, which is listed
00:16:03right here. Or you can use the template I'm going to put on the screen right now to help guide you
00:16:07for some of these bigger tasks. The big thing here is that you just need to know how you want it to
00:16:12behave. Are you someone who likes it when AI is constantly checking in, which Astra will tend to
00:16:18do on its own? Or when we're doing longer stuff, do you want to just do its thing? I'm going to give
00:16:23you a North Star, I'm going to give you some sort of end state, go forth and conquer, don't ask me
00:16:28questions, figure out. There's pros and cons to both of these solutions. Just know when you take the
00:16:33route of just, Hey, go ahead and figure it out. That's when you tend to get sort of a regression
00:16:39to the mean, unless you set up a lot of scaffolding that sort of pointed in certain directions when it
00:16:43reaches said forks in the road. Another good option you have here is this prompt that OpenAI gives us
00:16:49where we're going to tell the model to ask for approval only after preparing a concrete reviewable
00:16:54result. This enjoys blocking the task. If it could already have done it, right? Don't just come to me
00:16:59with a problem, come to me with a problem and your potential solution. So I think some combination of
00:17:04all of these and what I showed earlier will get you to a good place. If you're having some problems with
00:17:09Astros or just like stuttering its way through tasks. So those are the five mistakes that are holding
00:17:15you back when it comes to GPT-6 after this model is extremely powerful, especially when we look at it
00:17:21inside of Codex because Codex has so many cool little features that I think Anthropic and the Cloud Code
00:17:26Desktop app are sort of falling behind on namely browser use, computer use, and the voice bone. So
00:17:33make sure to try those out. Let me know how they worked for you. As always, if you want to learn more
00:17:39about Codex and Cloud Code, make sure to check out Chase AI+. I'll put a link to that down in the pin
00:17:44comment and I'll see you around.

핵심 요약

Optimizing GPT-6 Astra requires lowering effort levels to reduce costs, utilizing browser automation for reference collection, pruning bloated skills, leveraging voice mode as an orchestrator, and explicitly prompting for handling decision forks.

하이라이트

  • Setting GPT-6 Astra effort level to max yields 73% on DeepBench for $12 per task, while the low setting achieves 67% for only $2.19 per task.

  • Browser and computer use features allow Astra to independently navigate sites like Dribbble, capture reference screenshots, and incorporate visual assets without APIs.

  • Running a skill audit prunes unused skills from the context window and runs benchmark tests to verify model compatibility.

  • Astra voice mode supports both individual chat interactions and multi-chat orchestrator workflows across different effort levels.

  • GPT-6 Astra asks for clarification at decision forks rather than making assumptions, requiring explicit prompt instructions for autonomous task completion.

타임라인

Configuring Model Effort Levels

  • Maximum effort levels burn weekly usage and increase costs without delivering proportional output quality improvements.
  • The DeepBench score drops only 6% from max to low while cost decreases from $12 to $2.19 per task.
  • Front-end design tasks completed in 13 minutes on light effort use 91,000 tokens compared to 30 minutes and 152,000 tokens on max effort.

Using max or ultra effort levels for standard projects provides marginal gains while significantly increasing resource consumption. Lower settings such as high or medium maintain competitive benchmark scores across coding and agentic tasks while preserving weekly usage limits and reducing execution time.

Leveraging Browser and Computer Use

  • Browser and computer use tools connect Astra to applications lacking official CLIs, MCPs, or APIs.
  • Astra independently navigates external platforms like Dribbble to search for references, capture screenshots, and integrate assets.
  • Automated visual reference gathering eliminates manual downloading and scaling tasks during web development iterations.

Connecting the model directly to the browser enables autonomous multi-step workflows. Instead of manually gathering design references from external sites, users can prompt Astra to visit platforms, take screenshots, and apply those visual references directly to the active project iteration.

Auditing and Pruning Developer Skills

  • Outdated scaffolding skills from older models clutter the context window and degrade performance.
  • A dedicated skill audit prompt analyzes logs and front matter to categorize skills into retirement, fixing, or preservation buckets.
  • Benchmark tests evaluate whether installed skills remain compatible with advanced models like Astra.

Retained skills installed months prior consume valuable context window space with obsolete instructions. A systematic audit process scans usage logs and front matter descriptions, isolating unused components and running automated tests to ensure remaining skills enhance current model capabilities.

Deploying Voice Mode Workflows

  • Voice mode operates either inside individual chat threads or as an orchestrator controlling multiple parallel chats.
  • Configuring voice mode to light effort suits high-level orchestration, while high effort supports active in-chat editing.
  • Orchestrator voice commands spin up separate background tasks and generate structured HTML briefs automatically.

Voice integration provides a fast, snappy interaction layer that separates live conversation handling from background task execution. Users can manage multiple concurrent agents through a single voice session, directing background chats to conduct research or build assets without manually switching windows.

Refining Prompt Instructions for Decision Forks

  • GPT-6 Astra frequently requests user clarification at decision forks instead of making autonomous assumptions.
  • Long-running agentic tasks require explicit instructions detailing whether the model should pause for questions or drive tasks to completion.
  • Specific prompting templates instruct the model to prepare concrete reviewable results before seeking approval.

Behavioral differences between model generations mean Astra halts more often to ask clarifying questions. Adjusting prompts to define a clear North Star and setting rules for handling forks in the road prevents unnecessary interruptions and maintains momentum during complex execution phases.

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기