Tokens Should Have Jobs — Katelyn Lesse & Angela Jiang, Anthropic

English
AAI Engineer
Computing/SoftwareManagementInternet Technology

Transcript

00:00:00Good morning. We're super excited to be here at AI Engineer with all of you. I'm Caitlin,
00:00:25and I lead platform engineering at Anthropic. And I'm Angela, I lead platform product at
00:00:29Anthropic. And today we want to talk to you about a concept that we've been spending a lot of time
00:00:34thinking about and working on with our team, which is this idea that we think that tokens should have
00:00:39jobs. So if you're building an agentic system and you're trying to accomplish some specific outcome,
00:00:46you're trying to get something done with agents, there's one lever that everybody pulls in order to
00:00:50get a better outcome, and that's usually increasing your budget, which means you spend more tokens,
00:00:55or you spend more expensive tokens. But we've been wondering, is that all there is?
00:01:02Underlying this assumption of using the budget is this kind of implicit perspective that every single
00:01:08token is basically fungible. And we've been wondering, is that actually true? Are all these
00:01:14tokens actually fungible? And to test that, we've been thinking, what if we gave tokens jobs?
00:01:20So if you think about the way that you would normally set up an agent to go accomplish a task,
00:01:23you give it that task, you give it this token budget, and then all the tokens that are being
00:01:27spent are basically indiscriminate, in the sense that they're all doing one job, they're just executing.
00:01:34But what if you take some of those tokens, and they're not just executing, they're doing some other job.
00:01:40So for example, maybe you take some of your tokens, and they're advising the tokens that are executing,
00:01:45or maybe the tokens that are executing, try to get something done well, and you take some other
00:01:50tokens, and you actually grade how well the executor is doing so that it can iterate and try again.
00:01:56Or maybe you have tokens that are dreaming, they're reflecting back on the job that other executors
00:02:02have done, and writing learnings to memory so that they can do it again. And what we call each of
00:02:08these, if you take some tokens that are executing, and some tokens that are doing some other job,
00:02:12let's call this a strategy. So let's go take a look at the first strategy, the advising strategy.
00:02:19Here, we're splitting up an executor and an advisor. The executor obviously executes, but crucially,
00:02:24they can call out to an advisor for advice. And then they can take this advice and figure out if
00:02:28they're doing the next step correctly. This is really helpful in use cases, for example, if you're
00:02:33building a sales agent. In an ideal world, you'd have that sales agent be able to actually help the
00:02:38sales rep flag when a follow-up is overdue or a deal is stalling. In this construct, having an advisor
00:02:43to be able to kind of make sure that all the different pieces are actually working is really helpful.
00:02:49So another example is grading. Let's say you're executing, and you kind of know exactly what good
00:02:55really does look like. You can define this in a rubric. And then each time an executor tries to
00:03:00accomplish that outcome, you can have a greater provision that grades how well the executor did
00:03:06while looking at that rubric. And if the executor did a good job, then great, it can be done. But if
00:03:11it didn't do such a great job, you can iterate again until you get that good outcome.
00:03:17So an example in practice of when you might want to use this is let's say you have a customer service
00:03:21agent. And you're running a store, and your customers are writing in, and they're saying,
00:03:25oh, I should get a refund for this thing. And your customer service agent needs to be able to respond.
00:03:30You probably have some like pretty specific criteria on when you would give somebody a refund.
00:03:35And so what you can do is define a rubric that uses that criteria. You can have a grader that goes and
00:03:41looks at the work that the customer service agent is doing and decide is it getting it right and is it
00:03:46coming to the right outcome. And the last strategy we have is dreaming. So in dreaming, there's an
00:03:52executor who naturally executes, and then there's a dreamer. The dreamer is actually able to inspect
00:03:57the work and the transcripts of the executor, and then it takes any of the findings that it has,
00:04:02and it writes them to memory. This memory is re-picked up by the executor for the next round,
00:04:06so ideally would have improved. A great use case for this is if you're building a recruiting agent.
00:04:12Now, recruiting requires a lot of interaction with feedback on whether or not a candidate does or
00:04:16doesn't make sense and if it's a good fit between both parties. And so by taking all this type of
00:04:20data, if you build a dreaming type of strategy on this agent, it's actually able to kind of sharpen
00:04:24the next round so that it's more and more increasingly useful. So let's make this concrete with some
00:04:30experiments. So what we did was we created a bench of a bunch of tasks related to financial analysis.
00:04:37And what we were doing with each of these tasks is trying to replicate in the real world an expert
00:04:42human financial analyst, how well would they do on each of these various tasks. And so what we did
00:04:47was we started with a control that's just executing. Let's try each of these tasks and we'll eval them
00:04:51and we're literally just executing. But then we can experiment with each of our strategies and see how
00:04:56well we perform.
00:05:00So we start with a super basic experiment. Let's just one-shot it. Let's take each of our
00:05:04strategies and we'll go and just make an attempt to accomplish these tasks. And we'll see how accurate
00:05:09we are. And so you can see here with executing, it didn't do so well. 15% accuracy. But because it was just
00:05:16a one-shot, the strategy got to choose how many tokens it would actually spend on its own. And so you can
00:05:21actually see that execute decided not to spend that many tokens, only 39,000. And as we go into our larger
00:05:28strategies, our more complex strategies, we did choose to spend more tokens and we did a better job.
00:05:33So this isn't really telling us much because sure, Dream did really, really well, but it used a whopping
00:05:37600,000 tokens to get there. That's right. So in order to actually figure out if varying the jobs
00:05:44produces any alpha, what we need to do is hold the budget constant. And to do this, we're going to take
00:05:48Dreaming's budget, that 600,000 or so, as the maximum budget that is fixed across the board. And we give
00:05:53every single strategy this budget in order to analyze how well it's performing. And as expected,
00:05:59again, if you give a lot of strategies more budget, you are going to see performance increase
00:06:03across the board. So execute went from 0.15 to 0.76. Advise and grade went from the 60s to closer to the
00:06:0990s. And that's, again, expected given the fact that if you give things more test time compute,
00:06:14they should generally perform better. But if that was the only thing that mattered,
00:06:18we should actually expect to see execute, advise, grade, dream actually all be at the exact same
00:06:23level given the exact same token budget. But what we're actually seeing is that there is an alpha,
00:06:29or there is a difference, and therefore an alpha for us to exploit. If you look at execute at the exact
00:06:34same budget level, it gets to 0.76. But advise is at 0.89. So while a minimal, it does exist,
00:06:40and so there is alpha for us to take a look at. Now, we decided to take a look at this analysis from a
00:06:46completely different lens. And as Kayla mentioned, you know, we're doing this bench for a very complex
00:06:52set of financial tasks in the real world. And we wanted to analyze the usage of agents with actual
00:06:58experts. So if we look at a financial analyst expert, right, the kind of task they need to do with an
00:07:05agent is that they're giving it something very concrete, like let's say make a P&L, and they're
00:07:09getting the result back. Now, if that result is 80% accurate on a bench, that sounds great. But in
00:07:15reality, what that means for that expert is they have to go back and recompute that P&L themselves,
00:07:20or alternatively send it through another run. And that's because in this kind of domain, for this
00:07:25kind of task, if you're not 100% accurate, it's actually not useful. You cannot make up an income
00:07:32number or you can't make up a cost number. You have to make sure that it's 100% accurate. So with this
00:07:36lens of the real world consequence associated with this domain, we needed to recompute our experiments
00:07:41and score them a bit differently. Crucially, we needed to make sure that our experiments had this kind
00:07:46of construct, where if it was scored perfectly, we'd actually give it a pass. And if it scored anything
00:07:51less than 100% on that kind of task, we would actually mark it as a failure. So let's look at a different
00:07:57cut of our data from our experiments with this lens, where we're looking for this perfect run,
00:08:02100% accuracy pass. And let's look at what percent of the time each of these strategies was able to
00:08:08achieve a pass. So we've got execute down at 42%, and we've got our more complex strategies doing a bit
00:08:15better, up to 75% accuracy. And again, this doesn't necessarily tell us a ton, because, you know, each of
00:08:23these strategies might choose to use different budgets over time, right? So what we did here was
00:08:28we fixed the budget, and we said, within a fixed budget, how well do each of these strategies perform?
00:08:35And so what really matters to us, actually, is if you're trying to get this perfect answer,
00:08:40and you're in the real world, you're running a business, what matters to you is the cost to you
00:08:44to get to that perfect answer. And so one way we can think about this is we had our execute strategy,
00:08:49for example. The execute strategy around 40% of the time will give you that perfect answer. So
00:08:54on average, you can expect to have to run it three times, and you should hopefully, sometime in those
00:08:59three runs, get a perfect answer. And as we talked about earlier, we fixed our budget to that highest
00:09:04token budget strategy, which was 600,000 tokens. So if you spend 600,000 tokens in each individual
00:09:10run, you have to run approximately three times, you can expect, on average, to have to spend 1.8 million
00:09:16tokens with the execution strategy to get to your perfect answer.
00:09:21And so if we take this analysis and run it across the board against all these strategies,
00:09:25this is actually the true cost it took in this domain for that agent to be useful for that
00:09:30strategy. So as Caitlin mentioned, for execute, which is our baseline, this is going to be 1.8 million
00:09:36true total token cost for you. But advise, grade, and dream are showing us a bit of difference.
00:09:41Crucially, advise and grade are actually quite token efficient when you think about the actual usage
00:09:46of the end output of each of these agents. So what does this mean for you as a business? Well,
00:09:53it actually really depends on what kind of thing you want to optimize for. And it's going to vary,
00:09:58right? There's going to be businesses who say, actually, for me, the most important thing is to be
00:10:02really token efficient. In that case, you should probably pick the advise type of strategy in order to
00:10:07solve for that particular domain in which you want to optimize that. There's going to be other areas
00:10:11or other businesses where you're going to say, I'm not going to care so much about token efficiency
00:10:15because what I really care about is reliability of that answer. And so I need to maximize the
00:10:19percentage of runs in which I get that perfect answer. In which case, you would actually pick
00:10:23completely different strategies. You'd probably lean towards grade or dream.
00:10:28So if you take away one thing, the thing we want everyone to think about is this idea that tokens are not
00:10:33fungible. You can use your tokens to execute. You can brute force your way through your tasks and you
00:10:37can throw more budget at it. But if you get really smart about having your tokens do these different
00:10:42jobs and try these different strategies, you're very, very likely to be able to get a better outcome
00:10:47for the task at hand within a fixed budget. And so let's talk a little bit about how we actually build
00:10:53strategies and how we bring this to life. So we've done a lot of work to create a really excellent
00:10:58harness for individual agents. And if you see this picture at the bottom here, this is actually the
00:11:04architecture that we've used for cloud managed agents, which is our agentic solution that we give
00:11:09to you within the cloud platform. And what we do on top of this is we start to get into the meta
00:11:15harness level, like the multi-agent orchestration and execution level, where this strategy can go and be
00:11:21coordinated between our executor and our advisor or the other agents within our strategy. And some of
00:11:27these, like dreaming and outcomes, we actually give to you out of the box within cloud managed agents.
00:11:32So with those set of primitives, it's actually relatively trivial for us to construct
00:11:38this kind of, you know, architecture, where we're able to combine these different types of strategies
00:11:43and figure out how they should work together. So for example, it's relatively trivial for us to say,
00:11:48okay, now with this, I can take a task and I should be able to execute it, but also allow it to advise.
00:11:53And fable is back online. So we could actually say fable is the one that's actually advising
00:11:57the executor. And then I can take all these results and say, send them to a grader so that I can make
00:12:03sure that this is verifying in a loop that makes sense. And if it passes, that's awesome. I want to
00:12:07send all of that stuff to dreaming and make sure that my next run is better than ever.
00:12:12And of course, you don't have to stop there, right? If the right primitives are there and the right
00:12:16coordination is there, then you can actually construct really complex setups that fit for
00:12:21all the different types of dynamic problems that you have. You can invent these kinds of large-scale
00:12:26architectures, again, very trivially. And you could also invent completely new jobs,
00:12:30not just the ones of the pieces that Kaitlyn and I have presented in this conversation.
00:12:34Kaitlyn Lee: So a big goal that we have over time is to get our models better and better and our
00:12:40platform better and better at dynamically constructing these strategies for you as you're doing work. But
00:12:45in the meantime, as we're working our way there, we would love for you to continue to think about
00:12:50this idea that you should give your tokens jobs and you should use different novel strategies by
00:12:54combining these primitives in order to get the outcomes that you want for your tasks.
00:12:58Kaitlyn Lee: Thanks for joining us.

Key Takeaway

Assigning tokens specific jobs like advising, grading, and dreaming instead of treating them as fungible execution units yields higher accuracy and cost efficiency within fixed token budgets.

Highlights

  • Tokens in agentic systems are not fungible, and assigning specific jobs beyond pure execution drastically alters performance.

  • Advising, grading, and dreaming strategies improve accuracy compared to baseline execution under fixed token budgets.

  • Baseline execution achieves 42% accuracy on complex financial tasks, whereas more complex strategies reach up to 75% accuracy.

  • Cloud-managed agents at Anthropic incorporate native multi-agent orchestration primitives for combining execution, advising, grading, and dreaming.

Timeline

Token Fungibility and Alternative Strategies

  • Agentic systems typically rely on increasing token budgets to improve outcomes, assuming all tokens are fungible executors.
  • Tokens can be assigned distinct roles such as advising executing tokens, grading outputs against rubrics, or dreaming and writing findings to memory.

System performance can be optimized by splitting token allocation across different strategic roles rather than relying solely on brute-force execution. Advising provides real-time feedback during tasks like sales tracking, grading evaluates outputs against clear rubrics for customer service, and dreaming records transcript lessons for future rounds in recruiting.

Experimental Benchmarks in Financial Analysis

  • Experiments involving financial analysis tasks compare one-shot execution against advisor, grader, and dreamer strategies.
  • Fixing the maximum budget at 600,000 tokens reveals performance differences, with execution reaching 0.76 accuracy and advising reaching 0.89 accuracy.

Testing complex real-world financial analyst tasks shows that higher test-time compute improves performance across the board. However, holding token budgets constant highlights performance alpha between strategies, demonstrating that specialized token jobs yield superior accuracy at identical cost scales.

Real-World Consequences and True Token Cost

  • Real-world financial tasks require 100% accuracy to be useful to human experts, making partial correctness a failure.
  • Measuring the true cost required to achieve a 100% pass rate shows baseline execution consumes 1.8 million tokens on average due to repeated runs.

Evaluating strategies based on the cost to achieve a perfect pass rate shifts priorities between token efficiency and absolute reliability. Businesses focusing on token efficiency deploy advisor strategies, while those prioritizing reliability lean toward grade or dream configurations.

Cloud-Managed Agents and Multi-Agent Orchestration

  • Anthropic provides cloud-managed agents with native multi-agent orchestration primitives for strategies like dreaming and outcomes.
  • Combining execution, advising, grading, and dreaming architectures enables the construction of scalable solutions for dynamic problems.

Building multi-agent workflows allows developers to chain different token roles together into cohesive pipelines. Models and platforms continue to evolve toward dynamically constructing these strategies automatically based on workflow requirements.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video