스크립트
00:00:00Good morning. We're super excited to be here at AI Engineer with all of you. I'm Caitlin,
00:00:25and I lead platform engineering at Anthropic. And I'm Angela, I lead platform product at
00:00:29Anthropic. And today we want to talk to you about a concept that we've been spending a lot of time
00:00:34thinking about and working on with our team, which is this idea that we think that tokens should have
00:00:39jobs. So if you're building an agentic system and you're trying to accomplish some specific outcome,
00:00:46you're trying to get something done with agents, there's one lever that everybody pulls in order to
00:00:50get a better outcome, and that's usually increasing your budget, which means you spend more tokens,
00:00:55or you spend more expensive tokens. But we've been wondering, is that all there is?
00:01:02Underlying this assumption of using the budget is this kind of implicit perspective that every single
00:01:08token is basically fungible. And we've been wondering, is that actually true? Are all these
00:01:14tokens actually fungible? And to test that, we've been thinking, what if we gave tokens jobs?
00:01:20So if you think about the way that you would normally set up an agent to go accomplish a task,
00:01:23you give it that task, you give it this token budget, and then all the tokens that are being
00:01:27spent are basically indiscriminate, in the sense that they're all doing one job, they're just executing.
00:01:34But what if you take some of those tokens, and they're not just executing, they're doing some other job.
00:01:40So for example, maybe you take some of your tokens, and they're advising the tokens that are executing,
00:01:45or maybe the tokens that are executing, try to get something done well, and you take some other
00:01:50tokens, and you actually grade how well the executor is doing so that it can iterate and try again.
00:01:56Or maybe you have tokens that are dreaming, they're reflecting back on the job that other executors
00:02:02have done, and writing learnings to memory so that they can do it again. And what we call each of
00:02:08these, if you take some tokens that are executing, and some tokens that are doing some other job,
00:02:12let's call this a strategy. So let's go take a look at the first strategy, the advising strategy.
00:02:19Here, we're splitting up an executor and an advisor. The executor obviously executes, but crucially,
00:02:24they can call out to an advisor for advice. And then they can take this advice and figure out if
00:02:28they're doing the next step correctly. This is really helpful in use cases, for example, if you're
00:02:33building a sales agent. In an ideal world, you'd have that sales agent be able to actually help the
00:02:38sales rep flag when a follow-up is overdue or a deal is stalling. In this construct, having an advisor
00:02:43to be able to kind of make sure that all the different pieces are actually working is really helpful.
00:02:49So another example is grading. Let's say you're executing, and you kind of know exactly what good
00:02:55really does look like. You can define this in a rubric. And then each time an executor tries to
00:03:00accomplish that outcome, you can have a greater provision that grades how well the executor did
00:03:06while looking at that rubric. And if the executor did a good job, then great, it can be done. But if
00:03:11it didn't do such a great job, you can iterate again until you get that good outcome.
00:03:17So an example in practice of when you might want to use this is let's say you have a customer service
00:03:21agent. And you're running a store, and your customers are writing in, and they're saying,
00:03:25oh, I should get a refund for this thing. And your customer service agent needs to be able to respond.
00:03:30You probably have some like pretty specific criteria on when you would give somebody a refund.
00:03:35And so what you can do is define a rubric that uses that criteria. You can have a grader that goes and
00:03:41looks at the work that the customer service agent is doing and decide is it getting it right and is it
00:03:46coming to the right outcome. And the last strategy we have is dreaming. So in dreaming, there's an
00:03:52executor who naturally executes, and then there's a dreamer. The dreamer is actually able to inspect
00:03:57the work and the transcripts of the executor, and then it takes any of the findings that it has,
00:04:02and it writes them to memory. This memory is re-picked up by the executor for the next round,
00:04:06so ideally would have improved. A great use case for this is if you're building a recruiting agent.
00:04:12Now, recruiting requires a lot of interaction with feedback on whether or not a candidate does or
00:04:16doesn't make sense and if it's a good fit between both parties. And so by taking all this type of
00:04:20data, if you build a dreaming type of strategy on this agent, it's actually able to kind of sharpen
00:04:24the next round so that it's more and more increasingly useful. So let's make this concrete with some
00:04:30experiments. So what we did was we created a bench of a bunch of tasks related to financial analysis.
00:04:37And what we were doing with each of these tasks is trying to replicate in the real world an expert
00:04:42human financial analyst, how well would they do on each of these various tasks. And so what we did
00:04:47was we started with a control that's just executing. Let's try each of these tasks and we'll eval them
00:04:51and we're literally just executing. But then we can experiment with each of our strategies and see how
00:04:56well we perform.
00:05:00So we start with a super basic experiment. Let's just one-shot it. Let's take each of our
00:05:04strategies and we'll go and just make an attempt to accomplish these tasks. And we'll see how accurate
00:05:09we are. And so you can see here with executing, it didn't do so well. 15% accuracy. But because it was just
00:05:16a one-shot, the strategy got to choose how many tokens it would actually spend on its own. And so you can
00:05:21actually see that execute decided not to spend that many tokens, only 39,000. And as we go into our larger
00:05:28strategies, our more complex strategies, we did choose to spend more tokens and we did a better job.
00:05:33So this isn't really telling us much because sure, Dream did really, really well, but it used a whopping
00:05:37600,000 tokens to get there. That's right. So in order to actually figure out if varying the jobs
00:05:44produces any alpha, what we need to do is hold the budget constant. And to do this, we're going to take
00:05:48Dreaming's budget, that 600,000 or so, as the maximum budget that is fixed across the board. And we give
00:05:53every single strategy this budget in order to analyze how well it's performing. And as expected,
00:05:59again, if you give a lot of strategies more budget, you are going to see performance increase
00:06:03across the board. So execute went from 0.15 to 0.76. Advise and grade went from the 60s to closer to the
00:06:0990s. And that's, again, expected given the fact that if you give things more test time compute,
00:06:14they should generally perform better. But if that was the only thing that mattered,
00:06:18we should actually expect to see execute, advise, grade, dream actually all be at the exact same
00:06:23level given the exact same token budget. But what we're actually seeing is that there is an alpha,
00:06:29or there is a difference, and therefore an alpha for us to exploit. If you look at execute at the exact
00:06:34same budget level, it gets to 0.76. But advise is at 0.89. So while a minimal, it does exist,
00:06:40and so there is alpha for us to take a look at. Now, we decided to take a look at this analysis from a
00:06:46completely different lens. And as Kayla mentioned, you know, we're doing this bench for a very complex
00:06:52set of financial tasks in the real world. And we wanted to analyze the usage of agents with actual
00:06:58experts. So if we look at a financial analyst expert, right, the kind of task they need to do with an
00:07:05agent is that they're giving it something very concrete, like let's say make a P&L, and they're
00:07:09getting the result back. Now, if that result is 80% accurate on a bench, that sounds great. But in
00:07:15reality, what that means for that expert is they have to go back and recompute that P&L themselves,
00:07:20or alternatively send it through another run. And that's because in this kind of domain, for this
00:07:25kind of task, if you're not 100% accurate, it's actually not useful. You cannot make up an income
00:07:32number or you can't make up a cost number. You have to make sure that it's 100% accurate. So with this
00:07:36lens of the real world consequence associated with this domain, we needed to recompute our experiments
00:07:41and score them a bit differently. Crucially, we needed to make sure that our experiments had this kind
00:07:46of construct, where if it was scored perfectly, we'd actually give it a pass. And if it scored anything
00:07:51less than 100% on that kind of task, we would actually mark it as a failure. So let's look at a different
00:07:57cut of our data from our experiments with this lens, where we're looking for this perfect run,
00:08:02100% accuracy pass. And let's look at what percent of the time each of these strategies was able to
00:08:08achieve a pass. So we've got execute down at 42%, and we've got our more complex strategies doing a bit
00:08:15better, up to 75% accuracy. And again, this doesn't necessarily tell us a ton, because, you know, each of
00:08:23these strategies might choose to use different budgets over time, right? So what we did here was
00:08:28we fixed the budget, and we said, within a fixed budget, how well do each of these strategies perform?
00:08:35And so what really matters to us, actually, is if you're trying to get this perfect answer,
00:08:40and you're in the real world, you're running a business, what matters to you is the cost to you
00:08:44to get to that perfect answer. And so one way we can think about this is we had our execute strategy,
00:08:49for example. The execute strategy around 40% of the time will give you that perfect answer. So
00:08:54on average, you can expect to have to run it three times, and you should hopefully, sometime in those
00:08:59three runs, get a perfect answer. And as we talked about earlier, we fixed our budget to that highest
00:09:04token budget strategy, which was 600,000 tokens. So if you spend 600,000 tokens in each individual
00:09:10run, you have to run approximately three times, you can expect, on average, to have to spend 1.8 million
00:09:16tokens with the execution strategy to get to your perfect answer.
00:09:21And so if we take this analysis and run it across the board against all these strategies,
00:09:25this is actually the true cost it took in this domain for that agent to be useful for that
00:09:30strategy. So as Caitlin mentioned, for execute, which is our baseline, this is going to be 1.8 million
00:09:36true total token cost for you. But advise, grade, and dream are showing us a bit of difference.
00:09:41Crucially, advise and grade are actually quite token efficient when you think about the actual usage
00:09:46of the end output of each of these agents. So what does this mean for you as a business? Well,
00:09:53it actually really depends on what kind of thing you want to optimize for. And it's going to vary,
00:09:58right? There's going to be businesses who say, actually, for me, the most important thing is to be
00:10:02really token efficient. In that case, you should probably pick the advise type of strategy in order to
00:10:07solve for that particular domain in which you want to optimize that. There's going to be other areas
00:10:11or other businesses where you're going to say, I'm not going to care so much about token efficiency
00:10:15because what I really care about is reliability of that answer. And so I need to maximize the
00:10:19percentage of runs in which I get that perfect answer. In which case, you would actually pick
00:10:23completely different strategies. You'd probably lean towards grade or dream.
00:10:28So if you take away one thing, the thing we want everyone to think about is this idea that tokens are not
00:10:33fungible. You can use your tokens to execute. You can brute force your way through your tasks and you
00:10:37can throw more budget at it. But if you get really smart about having your tokens do these different
00:10:42jobs and try these different strategies, you're very, very likely to be able to get a better outcome
00:10:47for the task at hand within a fixed budget. And so let's talk a little bit about how we actually build
00:10:53strategies and how we bring this to life. So we've done a lot of work to create a really excellent
00:10:58harness for individual agents. And if you see this picture at the bottom here, this is actually the
00:11:04architecture that we've used for cloud managed agents, which is our agentic solution that we give
00:11:09to you within the cloud platform. And what we do on top of this is we start to get into the meta
00:11:15harness level, like the multi-agent orchestration and execution level, where this strategy can go and be
00:11:21coordinated between our executor and our advisor or the other agents within our strategy. And some of
00:11:27these, like dreaming and outcomes, we actually give to you out of the box within cloud managed agents.
00:11:32So with those set of primitives, it's actually relatively trivial for us to construct
00:11:38this kind of, you know, architecture, where we're able to combine these different types of strategies
00:11:43and figure out how they should work together. So for example, it's relatively trivial for us to say,
00:11:48okay, now with this, I can take a task and I should be able to execute it, but also allow it to advise.
00:11:53And fable is back online. So we could actually say fable is the one that's actually advising
00:11:57the executor. And then I can take all these results and say, send them to a grader so that I can make
00:12:03sure that this is verifying in a loop that makes sense. And if it passes, that's awesome. I want to
00:12:07send all of that stuff to dreaming and make sure that my next run is better than ever.
00:12:12And of course, you don't have to stop there, right? If the right primitives are there and the right
00:12:16coordination is there, then you can actually construct really complex setups that fit for
00:12:21all the different types of dynamic problems that you have. You can invent these kinds of large-scale
00:12:26architectures, again, very trivially. And you could also invent completely new jobs,
00:12:30not just the ones of the pieces that Kaitlyn and I have presented in this conversation.
00:12:34Kaitlyn Lee: So a big goal that we have over time is to get our models better and better and our
00:12:40platform better and better at dynamically constructing these strategies for you as you're doing work. But
00:12:45in the meantime, as we're working our way there, we would love for you to continue to think about
00:12:50this idea that you should give your tokens jobs and you should use different novel strategies by
00:12:54combining these primitives in order to get the outcomes that you want for your tasks.
00:12:58Kaitlyn Lee: Thanks for joining us.
커뮤니티 글
아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!
이 영상에 대해 글쓰기