If we want them to do Knowledge Work, design them as Knowledge Agents — Benjamin Clavié, Mixedbread
AAI Engineer
Computing/SoftwareManagementInternet Technology
Transcript
00:00:00Okay, so hi everyone. I'm going to give the quickest introduction to myself. I'm Ben Clavier. I worked at Mixed Bread where we do retrieval. I'm French and I live in Tokyo. And today I'm going to talk to you about the fact that agents should do knowledge work and so we should design them like knowledge workers.
00:00:29Like we should design them like knowledge agents and not coding agents. And I'm going to explain the difference and why I think that's important. It's a bit of a hot tech talk, but let's start now. So the first thing is like first agents give us fun trivia. And I'm talking like early agents, 2022 agents back when all you had was, you know, rag, but agents couldn't even do tool call back then. So all you had is like you had the NIF statement, you did search, you got like cool, you could talk to your PDF. That was the very first form of agentic work. It was not very useful.
00:00:59We talked about that for long. What came next was programming agents and that's been all the rage. Like once agents started being able to search, actually properly search, actually properly understand things, carry tasks out, call tools, we started designing coding agents and coding agents are a big thing.
00:01:15And that's a form of knowledge work. And that's a form of knowledge work. But agents were not knowledge workers at the time. Like agents were coding agents. And now they're becoming knowledge worker. And by knowledge worker, I mean that knowledge work is a big superset and kind of every workflow you've thought of before is a form of knowledge agents.
00:01:37So it's like if you have an agent that's a lawyer that's looking for legal documents, if you go financial agent, if you're looking for, you know, medication information, like you've got a lot of people on Twitter that try to do those self-diagnosis and you've got just a lot of medical usage.
00:01:53All of that is coding agents. We're getting an agent. It's trying to find knowledge agents. It's trying to find knowledge. It's trying to make use of knowledge. And coding is part of that, of course. And even the small rag bit that we talked about in the first slide is part of that.
00:02:09But that's a very, very small proportion of the actual full thing. Like there's so much more to knowledge than any one domain. And why even is knowledge work? Because I'm saying that it's important. They do knowledge work. Coding is knowledge work.
00:02:23And I think there's two ways to define it in my opinion. One of them is knowledge work is work where your main input is information. Like your main input is not an actual physical material. It's not something that you can touch. It's knowledge. It's information.
00:02:38And the nature of knowledge work is that you process this information, which is by nature very ambiguous, very diffuse. And the main output you get from that is something actionable. It's a judgment. It's a decision. If it's a lawyer, you're going to get, you know, the findings on your case and they might plead for you. You're going to get an actual actionable thinking item, like something still not tangible, but that exists as knowledge.
00:03:03There's also a tautological definition, which makes sense here, is that if you need search, it's a knowledge problem. And if it's a knowledge problem, you need search. So it's very easy, self-defined.
00:03:14And in the real world, that's basically most of the work that we see in the service economy is a form of knowledge work, like lawyers, knowledge workers, academics, knowledge workers, actualized knowledge workers, software engineers, researchers, also knowledge workers.
00:03:29And the fact that there's so much knowledge work in society has contributed to a never-improving structuring of knowledge work. There's actually very, very well-defined workflows for how we should do knowledge work, for how knowledge works in itself and how we evolve that.
00:03:43But so far, agentics kind of focus on the special case, and I try to generalize from it. And that special case is coding and software engineering. And the thing is, code is knowledge, but not all knowledge is code. And code is a very, very unique form of knowledge because it has very durable cues.
00:04:01So in a code base, there's going to be a lot of references to an identifier or a file or a path. And, okay, when we vibe code, that can change, but most of the time, it's not going to change all that much.
00:04:12Like, all the things are very, very durable. Then you've got that surface, obviously, is now in grep-able. There's keywords, there's method definition, there's a lot of things that, by definition, you can grep in code.
00:04:24And the task, and this one's actually very important, and we don't talk about it a lot, but people are like, oh, why is grep good enough for programming, or why can an agent do programming, and then you're telling me it can do, like, deep research for a legal question?
00:04:37And one of those is because we don't realize it, but, like, when we interact with coding agents, we are giving them extremely narrow tasks. Like, we're not actually expecting that much from them.
00:04:47Like, everything's always kind of about a feature or about, like, a given ticket. Like, there's a task at hand. You're not going to tell the agent, like, discover a new programming paradigm and then implement it in this new app.
00:04:58I don't know why it's going to do good luck. But in knowledge work, that's often the case. Like, first of all, you don't have those durable queues. Like, the meaning is always implicit.
00:05:08And more importantly, the same queue can mean a lot of different things. We don't have function definitions in knowledge work. Like, if you see 30 days, if your agent's looking for 30 days,
00:05:17is it a deadline? Is it a grace period? Is it a retention rule? Is it even in the same domain? Like, are you searching for 30 days on a contract and you're getting medication?
00:05:25There's a lot of contextual information here, but more importantly, the search starts from an intent. Like, even if you're doing, again, legal example,
00:05:33if you're asking about, like, a specific rule that you want to apply to a specific domain, you're going to need to look at the international norms that apply,
00:05:40and then do they apply in this case? There's a lot of conditional information that is not predefined in the task. Like, that's all up for the agent to find.
00:05:47And so non-code knowledge is very contextual and meaning-driven, which is much harder than code.
00:05:54And that's led to the fact that none of what I'm saying is new. Like, people have been doing knowledge work for a very, very long time,
00:06:01and that's resulted in, like, two endless loops. So you have a tool loop, which is, at the start, we were, like, talking.
00:06:10Then at some point, some guy was like, "We should write stuff down." Then in Alexandria, we had the Pina case,
00:06:14which was the curator of the library of Alexandria, came up with an idea that maybe we should have a way
00:06:19to catalog all of the books we have. Then we developed writing. Then we developed bibliographies.
00:06:24Then we ended up with, like, the current version of the, like, Dewey system for libraries.
00:06:30And nowadays we have search engines. But we also had an organization loop,
00:06:34which is joined, but also disjoint from the tool one, which is it used to be the one gifted expert.
00:06:40Like, we've all heard of the polymath of the past, the person who just knew everything about one domain or all domains,
00:06:46and you just went to them if you had information. But that doesn't scale, so we ended up with, like, monasteries,
00:06:51which were, like, guardians of knowledge. And then we had universities, and then we ended up creating the bureaucracies.
00:06:55And now we ended up creating the modern organization of work where we have very specialized firms.
00:07:01Like, at hospitals, you've got the doctor, you've got the senior doctor, you've got the nurse practitioner,
00:07:05the nurses, the healthcare assistants, and all of them kind of, like, specialize on different levels of tasks.
00:07:10And that's a really good form of optimization. But the thing is that it's actually just the one loop.
00:07:17Like, I'm showing two loops here, but they're actually just the one loop,
00:07:20which is we have new knowledge, and new knowledge means that we need better tools.
00:07:24And better tools mean that we end up creating new workflows, new roles.
00:07:28Like, we need people that are trained to use those tools, people that understand what the new tool does.
00:07:33If you have a guy that knows how to go to the library, and you're like, okay, use Google,
00:07:37you need the knowledge of what Google is. Like, that person needs to be taught that's a search engine.
00:07:42You can just type stuff in it. There's no need to physically go there.
00:07:44And that means you retrain. You get new knowledge workers who are more efficient,
00:07:47so they create more knowledge. So we need new tools, and so on, and so on.
00:07:51So both the tool loop and the organizational loop are actually just this one self-optimizing loop
00:07:56that kind of, like, triggers the other endlessly.
00:08:00And the thing about, like, tooling and optimization is that they're not neutral add-ons.
00:08:06Like, I say that we keep optimizing tools, and things come up, and we create new things out of those tools.
00:08:11But that's never actually a neutral thing. Like, tooling is not just, oh, my search is 5% better.
00:08:16The fact that we have a tool or the fact that we don't have a tool is what decides if a task --
00:08:20not if the task is possible because you can do things without the right tool,
00:08:23but if the task is actually scalable and can be carried out cheaply because something being cheap means it can scale.
00:08:29And it's like, yes, of course, if you go to the Library of Alexandria before the Pinakis,
00:08:34you can find your manuscript somewhere. Like, whatever you're looking for is there.
00:08:37It's probably going to take two or three weeks.
00:08:39So you're going to really, really, really need that knowledge.
00:08:42But if there's a library catalog, it's going to take you 10 minutes, and now, like, it's way easier to just --
00:08:46oh, okay, I need to know something more about this, so I'm going to search for it.
00:08:50Likewise, if you have a map directory or if you even have a map in the first place,
00:08:54which in itself is a tool for information, then exploring the world is a much better idea.
00:08:59Like, you're not going to rely on randomly discovering America on your way to the Indies.
00:09:03You know, like, you know where you're going.
00:09:05And likewise, if you have, like, a multimodal search platform,
00:09:08then you can search millions of PDF in a way that we couldn't before.
00:09:11So now there's a lot of use cases where you're like, oh, it's in the archives.
00:09:14I'm not going to touch that, but become actually useful.
00:09:19And in practice, this kind of looks like that.
00:09:21And I'm getting into the more technical stuff here,
00:09:23which is on a simple deep research task.
00:09:25So this is the BrowseCom Plus leaderboard,
00:09:27which is made to evaluate the quality of search tools on a very bonded deep research task.
00:09:32You have 200 found documents, and you have, like, specific queries.
00:09:35We all talked about this this morning, and it's a really useful benchmark to, like, analyze queries.
00:09:40And what we see is that, like, okay, a bad tool --
00:09:43so that's the thing that people often rant about.
00:09:46You will see that there's two BM25 here.
00:09:48There's two, like, optimized and unoptimized.
00:09:51And that's because quite often people will tell you BM25 is not great.
00:09:55And the reason they'll tell you BM25 is not great is because there's not one BM25.
00:09:59There's hundreds of them.
00:10:00It's a way to do a lexical search.
00:10:01You should always optimize your baselines.
00:10:03You should always, like, optimize what you're beating.
00:10:06And so what you see here is, like, a badly optimized tool is useless, like 60% accuracy.
00:10:11You're not going to trust someone that's right 60% of the time.
00:10:14You're just going to do it yourself.
00:10:16When you start optimizing the tools, you can see we go up to 70, 80, and then the actual best is a hybrid harness.
00:10:22It gets to 90.
00:10:23But that's maybe not the most interesting part because we start kind of plateauing at one point.
00:10:28Like, the jump from 89.8 to 90.2 is in-run variance.
00:10:34That doesn't matter.
00:10:35What matters here, however, is that 90.2% accuracy, you reach it with 20% fewer tool calls.
00:10:42And that's huge because, in practice, that's 20% fewer tokens, 20% fewer resources that you use.
00:10:48That's basically 20% free cash.
00:10:50And if you compare it to the unoptimized baseline, you're spending, like, 5% of what you were spending in the first place.
00:10:57So the tool is actually what makes the task worth doing.
00:10:59Nobody would keep using that tool if it takes 25 calls.
00:11:02But if it takes 8 calls, you're like, oh, yeah, cool.
00:11:04That's a workflow I can introduce.
00:11:08And the second part, which goes with tooling, and I think is just as important, because BrowseCom Plus, in the previous slide, is interesting, but it's easy.
00:11:15It's 100,000 documents.
00:11:16It's just text.
00:11:18It's just the one question.
00:11:19And it's not really that open-ended.
00:11:21It's just a bit convoluted.
00:11:22But when you're actually doing, like, real-life knowledge work, there is that workflow that you see that I try to do, which is you have a client.
00:11:30They come to the big shot.
00:11:31They come to the lawyer.
00:11:32That's the partner of the agency.
00:11:34And they're like, okay, this is my situation.
00:11:36That's my problem.
00:11:37And they meet together.
00:11:38But then what the partner does is they're not going to be the ones doing all the legal research.
00:11:43They're not going to be the ones, like, doing every single step of the problem.
00:11:46What they'll do is kind of understand that, like, okay, this person has this problem.
00:11:50That's going to cause them that.
00:11:51Those are the facts.
00:11:53I need the relevant laws to this, that, and so on aspects.
00:11:57And then they've got party goals.
00:11:58They've got assistants.
00:11:59And the assistants are going to be doing this research.
00:12:01They're going to be using, like, the size of tools they've been trained to use.
00:12:04And they're going to produce memo and notes.
00:12:06And then they're going to give that back to the big shots.
00:12:08Maybe they'll research, like, one clarification point.
00:12:10But they mostly rely on what their searcher agents, if you want.
00:12:13Like, their assistants have fun for them.
00:12:15And that's the response that you're going to get.
00:12:18And that's echoing the point I made before, which is in code, when you're using code code,
00:12:22you're kind of doing that work yourself.
00:12:24You've already broken down the query.
00:12:25You know what you want to do.
00:12:27You've got a linear ticket.
00:12:28You've got something that you're giving the agent.
00:12:30In the real world, you've got a client that's got a very open-ended problem.
00:12:33And you need to break it down yourself.
00:12:35And your agent needs to break it down himself and then needs to use sub-agents that do this research.
00:12:41And this is how better tools and organization work together.
00:12:45Because this one is MatQA, which is another form of knowledge benchmark.
00:12:49MatQA is something that Hugging Face and Snowflake jointly released.
00:12:53And it's PDF-based enterprise task.
00:12:55It's got PDFs and it's got OCR versions of the PDF.
00:12:58And the current state of the leaderboard really showed that both tools and organizations are necessary.
00:13:04And you can see that in the fact that with BM25, however optimized it gets,
00:13:10the human in Gemini 3 reached the same setting.
00:13:13And that doesn't mean that Gemini 3 is as good as a human.
00:13:16That means that even a human cannot get the right information given unlimited searches with BM25.
00:13:22So you have the tool setting and you need better tools to go forward.
00:13:26That's the tool optimization part of the loop.
00:13:28Thankfully, we've got better tools.
00:13:30We've got models that can handle PDFs.
00:13:33We've got vision.
00:13:34We don't need to rely on OCR text.
00:13:36And what we see with that is that we get another jump, which is Gemini and the mixed bread search tool,
00:13:41which is fully multimodal so it can read the PDF.
00:13:44You get the tables.
00:13:45You get all that nice stuff in your own search.
00:13:47And that gets us a big jump in accuracy.
00:13:50But the interesting part is that, like, that doesn't work.
00:13:53Well, that doesn't work.
00:13:54That does work.
00:13:55But that doesn't work as well as we would like.
00:13:56Because why is my agent getting 88.9 if the human is getting 99.4?
00:14:00Like, that's 10% I'm living on the table here.
00:14:02But I don't understand that's an agent.
00:14:04It gets, I think, 10 tons in the benchmark.
00:14:06So it's a fully agentic system.
00:14:08It gets to think about its results.
00:14:09And yet, it's missing performance.
00:14:11And that's where we introduce the mixed bread search agent,
00:14:14which is exactly that breaking down of work we saw earlier,
00:14:17where we basically tell the main agent, the one that's answering the question,
00:14:21being like, okay, that's a big topic.
00:14:23There's, like, thousands of PDFs.
00:14:25You're not going to, like, search yourself.
00:14:26Just please break down the problem for me.
00:14:28Please write queries about the aspects that you think are important to answer.
00:14:31They are important to answer the actual query.
00:14:33And we get searchers that go off on their own.
00:14:35And they find the right results.
00:14:36And they bring you, like, a little memo to your agent.
00:14:38And then your agent actually answers that.
00:14:40And that gets the accuracy by 3.5 points.
00:14:43And that doesn't sound like a lot.
00:14:45But I like to think of it as, like, an Oracle gap.
00:14:47And the Oracle gap is the difference between perfect documents
00:14:51and your search system.
00:14:53And the Oracle gap here is about 10 points before using the agents.
00:14:58And it goes down to six points after using the agents.
00:15:01So that means that we have about a 40% reduction in mistakes.
00:15:04Like, the gap between humans and agents goes down by 40% just by, like,
00:15:08having a better architecture to search through it.
00:15:12And I think this is the end of my slide because I'm running out of time.
00:15:16And that's perfect because that's my takeaway slide.
00:15:18And what I want you to, like, get from this talk is that we know how to design better
00:15:23knowledge work for humans.
00:15:24And AI agents really benefit from this pattern.
00:15:26Like, we have designed this.
00:15:28We know how to do this.
00:15:29Humans have worked on this for centuries.
00:15:30Like, people have always needed more knowledge.
00:15:33The empires used to have librarians.
00:15:34We have paralegal.
00:15:35We've got legal firms.
00:15:36We know exactly how the legal industry has figured out paralegal.
00:15:41The medical industry has figured it out.
00:15:43And none of it looks like programming.
00:15:45Programming has a very different system because it's a very specific use case.
00:15:49And we should really learn from the knowledge world to know how to design agents
00:15:54that will do work for the knowledge world.
00:15:56And then you must not, like, overfeed on tools because tools don't exist as a way to do things
00:16:02by themselves.
00:16:03Tools exist as a way to overcome ceilings.
00:16:05You want a better tool when you see that you're hitting a ceiling, that your performance
00:16:08is not where you want it to be.
00:16:09So we designed better tools to overcome that ceilings.
00:16:12And more importantly, the tools need to be co-designed with the agents.
00:16:16Like, the agents need to know how to use tools because one thing you would often see is
00:16:20agents will try to write grep queries because grep's everywhere in the training data.
00:16:25BM250 and the query in the data.
00:16:27And that's not always what you need.
00:16:28Sometimes you need semantic search of our PDF.
00:16:30And you can grep a PDF.
00:16:31You can BM25 a PDF.
00:16:32You need to, like, write a better query.
00:16:34So it's very important that your agentic harnesses or even your agentic models know
00:16:39that they have got more than one tool.
00:16:41And it's about primitives.
00:16:42Grep's a primitive.
00:16:43BM25's a primitive.
00:16:44And semantic search is a primitive.
00:16:46And all of those need to be, like, very well trained.
00:16:49Like, the models need to know about all of them.
00:16:52And the last one is that the right orchestration of search will get you much better results
00:16:58because context is a finite resource.
00:17:01And even if we get to a model that's got, like, 100 million token context,
00:17:04A, that's going to cost you a lot of money.
00:17:06And B, that's still nothing.
00:17:08You're not even getting half of, like, one state's legal code,
00:17:11let alone the US, let alone the international law, let alone the specialized course, et cetera.
00:17:15So you need to have a way to break down your task.
00:17:18And you need to have your orchestrator, your main agents,
00:17:21and people that can actually organize the knowledge for them.
00:17:25And, yeah.
00:17:27So we've got two minutes for questions.
00:17:30Thank you.
00:17:31Thank you.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video