Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori
AAI Engineer
Computing/SoftwareBusiness NewsManagement
Transcript
00:00:00Thanks a lot for your time, really appreciate you dropping by, and it's always a great honor to
00:00:18speak at the World's Fair, so I'll do my best to give you guys some valuable insights, and yeah,
00:00:24hopefully make it worth your time. So my name is Maximilian Puros, and today I'll be talking about
00:00:29mouse power, and this is a talk about measuring agents through mental models, but before I get
00:00:36into talking about measuring agents, I'm going to talk through a bit about how I use them every day,
00:00:41and it might seem familiar to you, but just a level set, we'll go through it, so I tend to background
00:00:46them like I'm sure a lot of you people are as well, so while my active attention is focusing on one
00:00:52thing, like perhaps giving this talk to you, I still want to make some progress on peripheral tasks,
00:00:57so I'll keep my attention focused on giving this talk while my agents can help me explore some
00:01:04designs in the background, because I think that my slides need a bit of work, so I've got my design
00:01:09system already set up, I've got some guidance given to my agents, so I'll kick off an agent to just try to
00:01:15explore some different directions on the type treatment, and the layout, and just try to get as many explorations
00:01:21as possible, but of course, one agent's never enough, so I like to kick off a bunch in parallel, you know, I've got a
00:01:28lot of slides to get through, so I need all my agents on exploring it in different directions, and hopefully I can
00:01:34get some interesting things to make my slides a bit better, and hopefully they can finish the job pretty soon, because we're
00:01:40obviously kind of up against the timeline, the deadline here, so this is generally how I work, I'm sure it's probably
00:01:49familiar to a lot of you, where we're just trying to kick off agents for as much as possible in parallel, because it
00:01:54always feels like there's just way more research to do, we want it to be as thorough as possible, there's way more design
00:02:00explorations to do, so whenever our main focus is on one thing, why not kick a bunch of agents off in parallel,
00:02:06and just try to maximize your time, and it's a lot of fun, of course, until you get the bill, and then
00:02:13you start to wonder, was it all worth it, right, did you vibe code too hard, were you token maxing too much,
00:02:21like, could you have been more efficient in how you approached your sequencing your agents, and so this is what I'm
00:02:29going to get into today, it's how do we value the token cost, and specifically how do we help our customers
00:02:34value it, so for the past year and a half, I've had the pleasure of working as the founding designer at a
00:02:40company called Utori, and we focus on computer use models, these are models that learn to use a computer
00:02:45like a human would, and the use case for them is when you can't get information from an API or an MCP, why
00:02:52not just send an agent out to use a computer like a human would, and then we can extract all types of data and
00:02:58manipulate it in ways that let us access all the stuff that wasn't accessible previously, so obviously less
00:03:04efficient than APIs and MCPs, but as a last resort, just have an agent go use the computer and try to get the
00:03:11information, here's the Utori agent using the Utori website, checking out its own benchmark, so kind of, it's
00:03:19admiring itself in a way, so yeah, it gets a bit weird, like, and a lot of what I do as a
00:03:25founding designer there is talk to customers, try to understand how can we make agents as intuitive as
00:03:29possible, how do we figure out the mental models they're using to value the use cases they want to
00:03:34send out agents for, and a lot of them do seem pretty confused so far, a lot of people are excited about
00:03:41agents, but the phrase that comes up quite often is that they feel like they're just scratching the
00:03:46surface, it seems like it's not quite intuitive how we can best use them yet, and so in a lot of my
00:03:52customer discussions, it always comes down to a question of, like, what is the best way to use
00:03:57agents, what are the best use cases for them, and how do I think about the trade-offs with
00:04:01regards to token cost relative to value, so I think we're still kind of building this muscle
00:04:06today, and this leads me to the thesis of the talk, which is that I think agents have a
00:04:11measurement problem, and as an example, here's me at work trying to measure some
00:04:16agents, and one of my coworkers took this photo and told me it looked like I was trying to
00:04:21solve the mystery of Pepe Silva, so as you can see, it's not an easy task to measure agents, but I'm sure
00:04:30some of you are saying, hold on a sec, like, what is this guy talking about, I've got a fleet of agents
00:04:35working for me right now, we're building our next million dollar app as we speak, and I'm having a
00:04:40totally fine time measuring my agents, to which I will agree with you, but then I will point you to
00:04:46the mandatory Upton Sinclair quote to remind us all that everybody in this room is very biased, and
00:04:52we're early adopters, and we're very excited to explore this new technology, but it doesn't mean that we
00:04:57represent the people that ultimately we're going to be trying to help adopt this technology, and so, you
00:05:03know, I think it's important to remind ourselves that in some way or another, we probably are selling
00:05:08tokens, whether it's indirectly or directly, and so when we think about our own token usage, is it really
00:05:13representative of all the people out there who have never touched an agent yet? Some people are still
00:05:18copy and pasting into ChatGPT, I may be married to one of these people, and despite how much I tried to get
00:05:23her to try out agents, she's not let me set her up with it yet, and so, as a reminder, when we think about
00:05:30helping people adopt agents, you know, all the people across the world that we think could get as much excitement
00:05:35and value as we do when we run off parallel agents, let's just remember this quote, and so it really boils down to the age-old
00:05:43problem of a new technology, and, of course, there's tons of history we can go to to study how people saw this in the past.
00:05:52We have this really exciting new thing, but we haven't quite figured out the right ways to communicate it, and so,
00:05:58for this talk, I'll go back to the 1700s, and we can take some notes from when James Watt was trying to sell steam engines,
00:06:05and at the time, he decided that a great use case for his steam engines was trying to replace a horse gin,
00:06:12and these are the, was the power source of a mill at the time, so when you're, for, let's say, a brewery,
00:06:18and you need some power source to grind your barley or whatever, I don't know, I'm not like a big brewery guy, so I don't know exactly how it's made,
00:06:25but you need a power source, and the power source at the time that was common was you hooked a horse up to a rotary arm,
00:06:31and the horse walked in a circle, and that's how he generated your power, and seems crazy today, maybe,
00:06:36but at the time, it was commonplace, and Watt thought, you know, it would be much better than a horse is like a very efficient machine,
00:06:42although he rightfully acknowledged that one of the barriers to adopting it would be this cognitive dissonance of trying to tell people who kind of think in horses,
00:06:52how do you adapt to this, to this old machine that's kind of intimidating and scary, and perhaps somebody's going to say it's going to solve all your problems,
00:07:01but you can't go back to see the vision yet, so perhaps that sounds familiar to any of us working in agents today,
00:07:07and Watt's solution was that he needed to understand the mental model of these people,
00:07:12and specifically to create a metric that would help him give some baseline of the relative improvement in efficiency,
00:07:18and so he literally studied horse gins and tried to get some kind of armchair measurements of how is, like,
00:07:28what are the mechanics and the average performance of it, and eventually came to a metric called horsepower,
00:07:35which may sound familiar, and he used this measure to, you know, this was to quantify the general power that the horses were creating at the time,
00:07:44and then he could use it as a basis to show the multiplier of efficiency that a steam engine could provide,
00:07:49and this metric was not very scientific at the time.
00:07:52It was not necessarily even accurate, you could say, but the main thing it did was it communicated an increase in value,
00:08:00and so this let people who love horses let them kind of calibrate the efficiency gains that they could get by attempting to adopt a steam engine,
00:08:12so not even necessarily what you would get when you use it, but what would get you over the limit of trying it out in the first place.
00:08:19And, you know, it's a pretty big feat because, like, although he had efficiency on his side with regards to this metric,
00:08:27you know, let's be honest, regardless of how efficient this was, horses just have great vibes,
00:08:34so, like, it's kind of hard to beat the vibes of horses, and so he knew.
00:08:37He had to kind of overcome the emotion and actually speak to something that gave him an ability to calculate the ROI.
00:08:45And, oh, sorry, skipped something.
00:08:48And so, yeah, the lesson being if we're not able to give something that is a tangible ROI for our customers,
00:08:55then it's very hard for us to communicate value.
00:08:58And I think we only need to look to our own industry to see all the examples where other people in the technology sector are failing to calculate good ROI as well.
00:09:07And so we might, in this room, think this is somewhat of a solved problem, but if you look to the other engineers in the world who are perhaps not as AI-pilled,
00:09:16they're theoretically very smart and should be able to figure out how to calculate this quite well,
00:09:21but then you get these scenarios where people are blowing through their entire token budget for a year,
00:09:27and they're blowing through it in a quarter, or they're, like, dealing with token leaderboards and such.
00:09:32And so obviously the incentives haven't quite aligned, and we haven't perhaps got the right measure of value in terms of the technology sector itself.
00:09:39And so how then do we end up scaling past that and talk to people who have no idea what we're talking about
00:09:45but still try to provide them a measure of, like, increased efficiency with agents?
00:09:51And so right now I think we're kind of in this doom loop where we're overspending and we're underusing.
00:09:57This is a term I borrowed from RAMP, and they have a great blog post on this.
00:10:00And so it's kind of this vicious cycle where we're just token maxing ourselves into austerity
00:10:05and then kind of dropping out of the loop until we get more FOMO to get activated enough to try it again.
00:10:11And so I think we have to break this loop, and I think the way we do that is by getting better measures that will communicate value.
00:10:18Some people are obviously on the right track.
00:10:21There was this chart floating around on X recently that the Coinbase CEO posted where they had internally started changing the defaults of what models they will start with
00:10:31and trying to only save the frontier models for the hardest tasks.
00:10:34And as a result, saw some good -- saw AI spend start to diverge from token usage.
00:10:40And this is a good start -- RAMP also, as I mentioned, has a great blog post about this.
00:10:44But I think the problem is still that it's too focused on tokens.
00:10:48And tokens are, of course, useful as a measurement of an internal system, but at the end of the day they're just an output.
00:10:56And so the tokens need to then be traced very cleanly to an outcome.
00:11:01So how many bugs did the tokens -- sorry, how many bugs squashed the tokens that we bought -- sorry, totally butchered that.
00:11:12How many bugs got squashed with our token spend?
00:11:15How many support requests got closed, et cetera?
00:11:18So clean outcomes and then cleanly tying those to progress on our objectives.
00:11:23And so without a very tight measure of ROI, this becomes very hard to do.
00:11:28And I think I'll take this further and say that it need not even be the broader technology industry
00:11:35that's encountering this problem, but also many of us in this room perhaps are.
00:11:40And although we're all probably enjoying coding with various agents and feeling like it's --
00:11:46it does feel like there's something there in terms of the increase in ability and efficiency,
00:11:51the problem, of course, is that we're all kind of dying by a thousand pull requests.
00:11:54And so even Anthropic, who has -- some people on the team have claimed to have solved coding,
00:12:01they have also admitted that they've not solved code review.
00:12:04And so as a result, the bottleneck has now shifted to human review where the efficiency gains from coding agents
00:12:11aren't quite -- aren't seen yet because we spend most of the time reviewing the code
00:12:14and we've not figured out how to scale that in tandem with the generation of the code itself.
00:12:20And so the bottleneck ends up shifting to the verification side and thus we don't have a way to measure value at scale
00:12:28and to judge quality at the same speed.
00:12:30And so again, going back to the ROI calculations, we generate all this code, but how do we know --
00:12:35we don't know that enough of it is good to justify the spend.
00:12:39And of course, maybe code review was always flawed, but it's just that agents are now exposing it for the --
00:12:46are exposing the actual problem.
00:12:48And I like this quote by Noah Hine who -- from a post about how to solve code review,
00:12:53where he's mentioning specifically that the assumptions underneath code review are what's now being --
00:12:57what needs to be revisited.
00:12:59So we have to check our priors to try to figure out a new basis for how we can code review in the age of agents.
00:13:06And I'm not going to go into how to solve code review.
00:13:09I think that's definitely a talk that's better given by somebody else and is a totally different subject.
00:13:15But what I think is important for this talk is why does code review feel like it is solvable?
00:13:19And I think that Noah is hitting on something important here, which is that as a culture,
00:13:24code review has a very good convergence on shared assumptions.
00:13:28And that lets you -- that lets you measure things at scale when we can all kind of converge on the measurement
00:13:35and it becomes somewhat of a clear rubric.
00:13:39And so the task at hand now is we have to adopt -- we have to -- sorry -- adapt those assumptions for the agentic age.
00:13:49And so we need to go -- if we're able to do that, then we can go from execution,
00:13:54at the speed of compute, the measurement at the speed of compute.
00:13:57And, of course, the measurements need to fit the mental models of the customers using it.
00:14:01And I think the lesson here being that if you're going to think of how to build an agent for something,
00:14:09you also have to think about how do you help the customers build or at least create a method for verifying that the output is good.
00:14:16And so it's not enough to build it.
00:14:18We also have to help them -- we also have to help them get to clear ROI calculations to justify their spend.
00:14:25And so this brings me to the idea of mouse power, which could be the equivalent of horsepower for the agentic age.
00:14:33Just as James Watt was able to show a measure of efficiency relative to the horses in the gins --
00:14:39in the horse gins that were the source of power at the time,
00:14:42we perhaps can also figure out how do we create a baseline of efficiency for the way we use computers today
00:14:48and can then demonstrate how much better or perhaps more performant on certain vectors an agent could be at that task.
00:14:55And, of course, it's not perhaps as easy a task as he had back then where he could just study the horse gin
00:15:02because it's not as if we can create some method to measure our cursor movements
00:15:07and, like, figure out the delta of how much more efficient an agent could move them,
00:15:11and thus we can say, yeah, agents are this much more performant than humans at these tasks.
00:15:15Trust me, I've tried.
00:15:17I had Claude vibe code me, this measurement device,
00:15:20and I thought maybe if I can figure out the movement,
00:15:23like the potential movement across the screen and measure how fast it went,
00:15:27I could get some clean measure of mouse power.
00:15:29But, of course, it's only joking.
00:15:31This is, of course, like a fool's errand because information space is just way too high dimensional,
00:15:37and so I think mouse power is never going to be a metric, of course,
00:15:41but it's more so an idea, which the idea being if you're going to sell somebody an agent,
00:15:45you also have to help them with the rubric of how do we actually verify that this agent is doing good work,
00:15:51and thus we can have a good measure of saying that these tokens are worth it.
00:15:57So how to do that, of course, is really up to you,
00:15:59and I won't be able to tell you how do you --
00:16:01I don't have any good frameworks for how do you figure out the right measurements
00:16:05to help provide anybody you're building an agent for,
00:16:08but what I can do is give an idea that I've been kicking around,
00:16:12which is based in information theory.
00:16:15So going back to Claude Shannon's ideas about measuring entropy in information,
00:16:20entropy being the uncertainty of a probability distribution,
00:16:24and, of course, very much the basis of how we train agents today,
00:16:28things like cross entropy and such being a big factor in determining how capable an agent is.
00:16:33I think that entropy is an interesting idea to think through with regards to not just the performance of an agent,
00:16:38but also the tasks that we're sending them out to perform on.
00:16:42And so I put together this matrix, which it maps on the x-axis,
00:16:48the uncertainty in the steps it takes to perform a task.
00:16:51And so when we're thinking of building an agent,
00:16:53I think it's not enough to just think what would be a valuable task for the agent to do,
00:16:57but also thinking about how much uncertainty are in the steps to perform that task itself.
00:17:03So an example would be booking a flight has much less uncertainty than, let's say, painting a masterpiece, right?
00:17:10Because you know there's certain information that has to happen in the flight purchase.
00:17:15There has to be a departing destination, arriving destination.
00:17:18There's going to be a seat chosen.
00:17:19It might be by the person.
00:17:21It might just be random.
00:17:22But these things have to happen for that task to be completed.
00:17:25And on the other hand, there is the task of like painting a masterpiece, right?
00:17:29And who knows what the steps are to that?
00:17:31And maybe you can get an agent to do it.
00:17:33But it would be very hard to figure out how we can actually create a relatively predictable pathway to that.
00:17:40But then on the other axis is the uncertainty in the acceptance criteria itself.
00:17:45So not just can the agent perform the task, but can we help somebody actually --
00:17:50or is there actually a clean rubric for how it's graded?
00:17:54And so thinking about ideas on these two axes and where they intersect perhaps gives us a better guide
00:18:00for how to build agents.
00:18:01And we can run through a few examples.
00:18:03So if we look at the -- at the left side, your right side -- yes.
00:18:09No, your left as well.
00:18:11Then -- I know the last speaker was also confused by that.
00:18:15So yeah, on the left side, when uncertainty in the task steps are low, then it's a very -- it's a very predictable outcome,
00:18:24or it's a very predictable pathway to achieve that goal.
00:18:27And so then, you know, why would you waste tokens?
00:18:30Just write a script.
00:18:31On the other side, when the steps to perform the task are very high in uncertainty, then you have very unpredictable information.
00:18:40And so it's probably at risk of being out of distribution and pre-training and probably has very sparse rewards for reinforcement learning.
00:18:47And so perhaps it's not a good task for an agent because it's just much harder to figure out how to actually model that data.
00:18:54And so obviously in the middle is -- I think I'm out of time, but I'm not getting kicked off yet.
00:19:01So I'll just finish this up quickly.
00:19:04So yeah, in the middle is probably the sweet spot.
00:19:06But then on the other axis, what's the uncertainty in verifying that this is actually valuable?
00:19:10So when you have high uncertainty in the acceptance criteria, you pretty much have a spot where verification is indistinguishable from execution.
00:19:17So why would you build an agent for something that to verify it was useful, a person pretty much has to do the work again?
00:19:24So like waste of tokens, obviously.
00:19:26And then it leaves that middle area where you have this interesting intersection of tasks that are -- they're not too uncertain in that they -- or they have a degree of uncertainty where they're not just a script or they're not out of distribution for training, but they have enough uncertainty to be interesting.
00:19:47But at the same time, they also have a property being relatively easy to validate the -- to check the value of them.
00:19:56And so they become in this place where they kind of become the shape of an MP style problem, which means they're easier to verify than to execute.
00:20:02And the reason I say that is because if you can figure out a pretty repeatable pattern for verifying their work, you can actually just throw agents at that problem as well.
00:20:10And so of course you don't just build the agent, you perhaps build the agent that verifies the work of the agent.
00:20:18And so yeah, this is perhaps -- this is a thought starter mostly, kind of still in the works, so happy to hear any thoughts on it.
00:20:27But if -- with this guidance, I hope when you're building your next agent, you can also figure out how to also build its mouse power.
00:20:34And thanks very much.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video