스크립트
00:00:00Let's just do a couple of quick questions and then we'll jump right in. How many of us in the room
00:00:19here have built voice AI agents? Okay, that's a pretty good audience here. And how many of you
00:00:28guys have built AI agents that have been deployed in production? Not bad. Okay, cool. So we'll talk
00:00:38about what typically happens, right? Like everyone's talking about voice AI agents. The, you know,
00:00:47one pill solution to pretty much everything in the world today is voice AI agents. So everyone's
00:00:51building one and trying to deploy that. They sound great when you're sort of building that in your
00:00:57dev sort of landscape. And then the moment you take this to, from a proof of concept to production,
00:01:03things start failing. So we'll walk through these five different angles of like how, or what we have
00:01:09seen at PLEVO with voice AI agents. But just before that, a quick intro from, from my side. I am Venki,
00:01:19the founder and CAO. She used the title agent engineering manager. I'm calling myself chief agent officer from,
00:01:28from a title standpoint. Okay, so what is, what is, you know, why, why are we even qualified for this,
00:01:35this discussion and like, what, what are we seeing that a lot of companies don't get to see? So I'll,
00:01:42I'll talk a bit about our journey in terms of like how we have come along so far and then jump right
00:01:48in. You know, we were, we've been around for about 14 years. Our journey has been a developer API platform.
00:01:54And then now an, you know, an AI agent business. We started with voice and SMS APIs back in the day 2011.
00:02:01And then, you know, now we are primarily focused on our AI agent offering the full stack.
00:02:08The, the full stack on our platform. We, we see over a billion voice calls each month across the globe.
00:02:16Uh, and which is where we have seen a lot of these, uh, you know, patterns emerge in terms of like how,
00:02:20when we work with our customers, what happens on their voice AI agents in, in production.
00:02:25Uh, we're a, uh, 90 member team and, uh, we've, we've, uh, 50 million funding in the bank. Fun fact,
00:02:33this is not from external VC investors. This is all from being a profitable company,
00:02:38having put that cash in the bank over, over these years. Uh,
00:02:42some customers we, we power across the globe. Uh, you know, we've just left some, some logos in there,
00:02:49uh, but primarily from an offering standpoint, uh, I would sort of cohort this into three different
00:02:54buckets. One is a programmable, uh, AI agent offering. We call it, uh, I mean, it's a speech
00:03:00pipeline, not a true speech to speech product yet, but that's a, that's a programmable offering.
00:03:05We also have an AI agent studio. It's a no code visual, uh, builder. And then we, like I said,
00:03:11we started with voice APIs. So we obviously have built this out over the last 14 years,
00:03:16the SIP trunking and the audio streaming layer. So we don't rely on other folks for the telephony
00:03:22or the carrier layer. Like that's the bread and butter business we've built over all these years.
00:03:26And that's on top of which our AI, uh, agent platform sits.
00:03:32Okay. With that, uh, let's get into this, right? Which I'm, I'm sure since you guys have all built
00:03:39AI agents, you've all seen this or, you know, build this in, in one manner or another. And we'll spend
00:03:44more time on this in terms of like how, uh, the entire pipeline looks right. Uh, what we see with
00:03:51customers is, and, and I'm sure you guys cannot relate to this is, you know, anyone thinking about
00:03:56AI agents, what they do is they pick up a bunch of these orchestration frameworks and, uh, they do a
00:04:01pretty good job, live kit or a pipe cat, you know, build their AI agent on top of that. Uh, they think
00:04:06they can just sort of orchestrate these different four layers, speech to text, LLM, uh, and, and TTS
00:04:13with turn detection in between and we're off to the races. Like my AI agent works in a, in a POC
00:04:19and it's good to work in production. Uh, typically that's what happens. They sort of measure their
00:04:24latencies and you can see some indicative latencies on, on this slide at, at each layer. And they're like,
00:04:30yeah, this, this, uh, seems good for me for what I need. So let's, let's position production. And then
00:04:36the, the production was, uh, sort of start to kick in and, and you see all sort of failure modes, which
00:04:41we are going to spend, you know, most of the time on, on, in this talk, at least. Uh, I've kept some time
00:04:48at the end for Q and A if you guys want to have, uh, you know, questions, but we'll jump right in from,
00:04:53from this to, uh, you know, different failure modes we see. Let's start with, you know, the first one which
00:05:01everyone talks about. Like this is the most spoken about failure mode, which is latency. Uh, I think we
00:05:07have a few AI agent talks today or voice AI agent talks today. Um, I'm pretty sure like everyone,
00:05:12everyone's going to touch upon this specific failure mode, which is why I'm bringing this right up, uh,
00:05:17in, in terms of, uh, you know, some like how this entire experience is for, uh, users, right? Uh, typically
00:05:26most folks measure this by time to first audio. So the time when your users stop speaking
00:05:34to your agent starts speaking, right? And I think you've, you've probably seen this if you guys have
00:05:39built YZ agents on, you know, what, uh, good or natural feels like, what, uh, sort of annoying
00:05:46feels like or noticeable feels like, and then what annoying feels like, which is, you know, different tiered
00:05:52steps. Uh, we notice, you know, most people want to be under 550 because that's what's advertised by,
00:06:00you know, platforms or, uh, you know, solutions or, or, or, or layers, but I think most end up between 750
00:06:07to 1.2. Uh, that's where most of the folks end up at. Uh, the really bad performing ones end up,
00:06:13you know, more than 1.2, and then you start to see users, uh, hang up. Uh, now I, I'll share with
00:06:19you like what we've seen practically in, in, uh, uh, production with, uh, customers using this with, at,
00:06:26at different layers, and then, you know, solutions to, uh, some of these. The way we want to think
00:06:33about this layer is sort of a balance between these three, which is cost, intelligence, and latency,
00:06:42right? And, and, and why do I bring these three up? Because they're sort of interrelated. I think one
00:06:48of the things I was just chatting with, uh, you know, a couple of folks outside. One of the things,
00:06:51last one year, we've seen a lot of innovations, a lot of intelligence spike on the LLM side of things,
00:06:58right? And most of the, uh, you know, intelligence has come in, in terms of thinking or, uh, you know,
00:07:05reinforcement learning and, and so on and so forth. The irony with voice agents is like almost always
00:07:11your, the, the LLM or the agent that's talking has to have thinking turned off, right? So all the
00:07:19advancements we've had in the LLM layer in the last one year, like none of that even apply here now,
00:07:25right? You, obviously you have, you know, better models that can do, you know, better instruction
00:07:29following or tool calling, but pretty much all of your intelligence that's been built in on the
00:07:34thinking layer is all off by default if you want it to be fast enough. So, so that's one of the
00:07:39ironies that we come up with. So then how do you sort of balance intelligent cost and latency?
00:07:44Let's, let's look at some of these, uh, you know, options, uh, that are out there in the market,
00:07:48right? So, and I'm specifically picking LLM because if you looked at the previous chart, LLM is, uh,
00:07:55you know, sort of your highest latency bucket that adds to this, right? And, uh, if you look at,
00:08:02you know, frontier models, which I think most folks start by default, your, your OpenAI, your clouds,
00:08:09your Gemini's, uh, you know, P50, TTFT is roughly around 450 to 500 on, on a good day and it can get
00:08:17spiky, right? It can, it can, uh, you know, P90, P95 can go easily upwards of 1.2, 1.3 seconds even,
00:08:25uh, and, and that's not good for the overall agent experience. So, so that, so that's your frontier
00:08:31model. Now, there's another option, which is your, your Cerebris or the Grok, uh, that is famous and
00:08:38popular for spitting out a lot of tokens or, or tokens very fast, right? Uh, these work, but for you to get
00:08:45dedicated latency or time to first token on these, you need dedicated capacity and that is really expensive.
00:08:51That's where I spoke about the cost, uh, as, as being one of the things to balance, right? It's really
00:08:56expensive. And then like you talk to anyone from the Grok team or the Cerebris team, they'll tell you,
00:09:01you need to book 12 months in advance for dedicated capacity. They're booked out for the next 12 months.
00:09:05So, so that's, that's a pretty expensive option. And then you really need to be sure that the model
00:09:11you're deploying on some of these infra layers, uh, will be here 12 months from now. And, and it's a,
00:09:17it's a big investment and a big unknown. So, so what's a realistic option for production grade, uh,
00:09:27agents that are, that are good quality and end up balancing, uh, three of these? Uh, this is what has
00:09:34worked for us, uh, which is the open source models. Uh, there are obviously a lot of them in terms of
00:09:41like the variety and, and, and variations you can pick. I'm specifically talking about the two we work with,
00:09:47uh, Quinn 3.5 and Gemma 4. These are, uh, you know, kind of cutting edge open source models, right, uh,
00:09:56out in the market right now. And we've done a lot of benchmarking around this, uh, in how they work.
00:10:02It, it can be scary to think like, okay, I have the models, now I have to host them, you know, run them
00:10:10on my own GPUs and so on and so forth. But if you are consistently targeting under 300 MS, uh, this,
00:10:17we've seen this to be a, uh, a great option to balance between latency, cost, and intelligence.
00:10:23Now, uh, some more deep dive here. If you're doing only English, uh, Quinn 3.5 or Gemma both work fine.
00:10:29But if you're doing multilingual, uh, right, international audiences, different languages, uh, Gemma 4 is a
00:10:35much better model for that. Uh, we have seen, uh, token, uh, fertility evals. Essentially what that means is,
00:10:44if, if I were to de-jargonize that, is like how many tokens does it take to generate one word in
00:10:48that language? Okay, so Gemma is much, much better, at least 2.5 to 3x better than Quinn 3.5 from that
00:10:56perspective. So your time to words is much faster on Gemma 4, uh, everything else equal, right, on, on a
00:11:04multilingual basis. Now, what sizes do you pick, uh, at the LLM layer? Uh, the mixture of expert usually
00:11:12works fine. Uh, the three or four billion mixture of expert usually works fine. The, the problem with
00:11:17mixture expert is like if anyone goes down, wants to go down the direction of fine tuning, that can be a
00:11:22challenge, uh, because fine tuning mixture of experts models are not easy. Uh, you can end up breaking the
00:11:28model, uh, a lot of times. So, so that's one challenge we see with mixture experts, but usually out of the box,
00:11:35it gets you 90% closer to where you want to be, like even without any fine tuning or, or, or custom
00:11:42work done on the model. Uh, so that's the advantage of mixer experts. Uh, now if you want to fine tune,
00:11:48and, and you, you want to go deeper and say like, look, I'm working for a specific domain, healthcare,
00:11:53what have you, right, and I want to make sure I'm able to fine tune my model, you want to start at least
00:11:58with, uh, the 8 billion, 12 billion, at least, uh, from where we are today, maybe, maybe six months from now,
00:12:04a 4 billion, uh, 4 billion model beats the 8 billion model, uh, hands down, but for today, uh, what we've
00:12:11seen is, uh, you minimum need a 8 billion or 12 billion model, uh, because you're looking for two
00:12:16things in these models. One, obviously, fast tokens, but, uh, good instruction following, okay? And the
00:12:23second thing is like very high, uh, success ratio in tool calling, because if you can do these two things
00:12:29well, then you are on to like 70, 80% there for not even having to fine tune it, fine tune any model,
00:12:35like models will work out of the box, right? Uh, so, so that's, uh, been our recipe. We've actually,
00:12:42uh, we run two flavors, one a fine tune model for specific industries, and then for, uh, you know,
00:12:50most generic use cases, uh, uh, MOE model just works out of the box. Uh, there are a few more
00:12:56tips and tricks we'll talk about in the upcoming slides where we see failure models, but, but that's
00:13:00where we, uh, stand from a, from a latency LLM standpoint. Um, all right, I'm running
00:13:07tired on time, so I'm gonna fast track this. Uh, now there are a couple of other flavors in this. Uh,
00:13:12people build agents with, uh, a mixture of models. What they do is, you know, for, uh, the, the talking
00:13:19part of it, they have a conversational model, which is a much lower, uh, smaller model, and then, uh,
00:13:24you know, maybe even a three billion model, and then for tool calling, they have a much larger model,
00:13:27so that they have, uh, improved tool calling success ratio there. Uh,
00:13:35sorry. The second one is, uh, assume your transcriptions are going to be brittle. Like, that's,
00:13:42that's, uh, something you want to sort of, uh, live by when you're building AI agents, even if you have
00:13:49the best transcription engine out there, and I'll, I'll show you why, right? Like, the, the, the state-of-the-art
00:13:55transcription engines out, out in the market, uh, you know, sort of get you to, uh, four to six percent
00:14:02word error rate, right? Uh, and this is on known eval sets. Uh, on real world, noisy calls with,
00:14:10you know, sort of, uh, accents, like people having different sort of accents, uh, domain vocabulary,
00:14:15and so on and so forth. Like, those usually end up in the double digits from a word error rate perspective,
00:14:21right? Uh, now you, obviously you can fine-tune, you know, pick up an open source model and fine-tune,
00:14:25uh, but we see typically, like, what breaks here often, and there are patterns here in terms of what
00:14:31breaks. So, proper nouns, jargons, uh, phone numbers, like random missing digits with phone numbers,
00:14:38uh, wrong substitutions. I'll, I'll walk through some examples of, like, how you solve for these
00:14:43addresses. When you're trying to collect a long address, uh, you know, the, the transcription engine
00:14:48could just end up missing some parts of it. Code switch languages. I, I'll just take an example of a
00:14:55language I speak, 'cause that's, was easy for me to put on the slide, uh, where, you know, like, if you
00:15:01were to sort of take English, but written in a different script, uh, that's what's used for Hindi,
00:15:06right? Like, this is English written in that script, right? Whereas, like, the actual English version
00:15:11of this is, hello, how are you? So, if, if I'm addressing an audience in a different country,
00:15:16where I have code switched languages, and I start getting my English in a different, uh, sort of
00:15:21script, everything starts breaking from the transcription engine to the LLM layer, and then
00:15:26beyond, 'cause your LLM starts then producing output in that sort of script a lot of times, and then your TTS
00:15:32messes up. Okay, so, so this is, uh, very important to be careful about, and if you want to build your
00:15:38agent independent of the transcription engine, you need to build a layer that normalizes all of this,
00:15:44right? We'll talk about solutions in a minute. And there is the other case, which is, uh, Hindi,
00:15:48in just, in just Latin or, or, or, you know, Roman, right? Which is, like, this is Hindi, but it reads
00:15:54English, which again messes up everything, uh, you know, downstream. Those are just examples. This applies to,
00:15:59you know, Arabic, Mandarin, uh, Japanese, what have you, uh, pretty much any language. So,
00:16:04what actually moves the needle with, uh, at the transcription layer? Uh, for prop, proper nouns,
00:16:11we recommend, uh, you using not just keyword boosting. I think a lot of transcription engine,
00:16:16engines provide you keyword boosting, where you can put in specific words into their engine,
00:16:21but doing dynamic keyword boosting. What that means is, don't keep the keyword for the entire state of the
00:16:26call. Just add that dynamically when you think you need that as an answer, so that you get the highest
00:16:32accuracy, meaning at different states of the call, the transcription engine will have different, uh,
00:16:38keywords boosted during different phases, right? And that's what we've seen works best, because if you
00:16:43just pollute your context of the transcription engine with tons of keywords, it'll start hallucinating again,
00:16:49right? So, so that's what we see typically working best. Uh, yeah, post process, post process your
00:16:55transcripts with an LLM, right? Because your LLM has domain context, your transcription engine does not.
00:17:02So, a lot of words that it would say, uh, I'll give you some examples, may not make sense. This is
00:17:07transcription, like a phone number from a transcription engine, right? Like, what do you think that E is?
00:17:13Right? If you give it to an LLM, it knows that's a three. Similarly, like what that one is, it's a digit
00:17:18one. So, so your transcription engine a lot of times could mess that up, but when you post process it
00:17:24with a LLM layer, it'll instantly correct that from a collection standpoint. I mean, uh, and, and the last
00:17:30one, like I said, uh, transliteration is your STT output that's sort of, uh, you know, multilingual also gets
00:17:37normalized, uh, using either an LLM you first transliterated or, you know, use some kind of a neural, uh,
00:17:46transliteration engine. There are a lot of them open source. You can just pick one of them,
00:17:49right? Uh, that will do all of that work for you. Send cleaned transcripts consistently,
00:17:55independent of the transcription engine to your LLM.
00:18:00All right. The third one we typically see is collecting data. This is where I think 50 to 60
00:18:05percent of AI agents mess up pretty badly. Uh, and like, we like to think of it as a UX problem,
00:18:14but just for voice. So think data models, uh, and not a transcript coming into an LLM and trying to
00:18:21figure out what the transcript said. So let's take some inspiration from, uh, I'm assuming most of us are
00:18:27developers here. Um, you know, take inspiration from Python's data classes, Pydantic,
00:18:32Zod from TypeScript or form fields in the UI, right? Like if you start thinking of it from that problem
00:18:39statement, uh, we have seen accuracy grow up from, uh, grow from 30% to like 95% from a data collection
00:18:46standpoint when you start thinking in that manner. So like decide your shape before you ask, right? Like,
00:18:52instead of keeping it open ended, can you keep it constrained? So can, can a phone number be a
00:18:58phone number type field? The moment you do that, right? You know, like how many digits it needs to
00:19:03have. You can do validation on, on top of that, right? And then what sort of allowed values can even be
00:19:10there. So in the previous example we saw, if an E comes in, in middle of a phone number and you know
00:19:15it's a phone number, you instantly know like either you smart guess that to three and confirm that with
00:19:20the user, or you know that's an error and then you validate that and ask the user to repeat again,
00:19:26right? So, so that's, I think, one of the common patterns we've seen here from, from a
00:19:31collection pattern. Name, I think, is the, is the interesting one. I've just picked a, you know,
00:19:36a hard to pronounce name. Like there is no way a human is going to get this right. And, and no way
00:19:43our transcription engine will get this right. How many ever times you do this, right? So the moment
00:19:47you start thinking of this as fields and then have rules and then confirmation mechanisms on, on, on
00:19:53spelling this, you know, sort of letter by letter, only then you kind of get it right. Otherwise,
00:19:58it's going to mess up pretty badly in terms of how you collect this on a voice call. And, and that's just
00:20:03an example of, you know, what I'm talking about in terms of the, the data collection piece of it.
00:20:11Another place where it goes badly dramatically is relative values, date being one of the examples.
00:20:18If somebody says next week, uh, Wednesday eight, it could mean 8:00 AM, 8:00 PM, and then figuring out
00:20:25what that date actually is, again, now becomes a very constrained problem. If you knew this was a
00:20:30date time field and I need, I'm collecting a date time field and then you take the current date and then
00:20:35figure out what this value would be based as that, right? So, so that's how you want to make sure
00:20:39like, uh, you do this with a combination of the LLM with the tool calling and the tool calling is
00:20:44doing a lot of this heavy lifting for you from a, from a field standpoint.
00:20:50Yeah. And then you make, you, you run like this from a unit test perspective. So all of your, uh, evals need to
00:20:58start treating these fields as unit tests. And as long as your unit tests, uh, sort of validate and pass,
00:21:06you know, your agent is going to be, uh, sort of reliable and repeatable. You don't,
00:21:10you know, run, uh, hundreds of end to end agent test cases just to find out, you know, one field
00:21:15collection is broken. You do your evals at a, at a field level and a unit test level.
00:21:24And then, yeah, I, like I said, I think, uh, you, this, this mindset makes everything more structured
00:21:30instead of hoping I'll put a ton of prompt, keep changing, you know, the prompt by a few, uh,
00:21:36characters every time. And somehow my prompt engineering is going to make LLM much more instruction
00:21:41tuned and sort of magically start following some of these things. So in fact, uh, like I said,
00:21:47right, like we have seen us get to 95, 97% accuracy without having to fine tune a model.
00:21:53All right. And then, and the trick is basically like just breaking down your context of
00:21:57what the agent is doing at that point with specific, uh, states of what the agent is going through.
00:22:04All right. Uh, I'm just going to quickly, uh, skip through this from a
00:22:10time standpoint. I just see, I got three more minutes. Um, hopefully that's a bug, but, but we'll
00:22:16leave it at that. Okay. Um, so, so this is the fourth area where we see issues coming in. Most folks take
00:22:24the LLM output and then we send it to a TTS. Obviously, I think there are a lot of good TTSs in the market
00:22:30that take care of a lot of heavy lifting, but a lot of times it, it messes up. Uh, what we recommend and
00:22:37what we've seen is you usually want to have a normalization layer between your LLM and what
00:22:43is fed to a TTS. You don't send your LLM output directly to a TTS, right? And, and we'll just walk
00:22:50through some examples. The basics, which is strip emojis, uh, markdown, uh, before, before any synthesis
00:22:58into the TTS. Most orchestration pipelines do this, like, you know, a LiveCAD or a PipeCAD would do that
00:23:02for you if you just set a few flags. So, I, but, but just make sure if you're not using them or
00:23:08buildings from scratch that you've said this explicitly because you don't want an emoji showing up
00:23:12on, on, on, on something read out or, you know, markdown showing up there.
00:23:17Okay. I think, I think some more common ones, uh, custom, uh, dictionaries. Most TTS engines provide
00:23:23this to you, uh, like how to pronounce custom words, whether it's, uh, you know, proper nouns, brands,
00:23:30uh, acronyms, and so on and so forth. So, set those in, uh, when you go from your LLM to your TTS output,
00:23:36because if you don't, you're going to mess that up. And I'll, I'll show you an example of like how we test
00:23:40that. Uh, the, the other one is like most engines also give you speed. So, if you know you're pronouncing
00:23:47an entity, slow down, have your agent slow down. So, at 0.8x or 0.7x, so that it, it's able to like
00:23:54enunciate on that specific entity and, and doesn't mess up how it's pronouncing an email or a phone
00:24:00number or a name letter by letter. And yeah, just normalize all the messy stuff, right? Like emails,
00:24:09currency, dates. Don't leave it to the TTS to do it. Uh, most of them do it, but don't leave it to
00:24:15the TTS to do it. Like build your normalization layer at your end so that tomorrow you think you
00:24:21need to switch TTS or, you know, for whatever reason the first one's down and you want to use
00:24:25another TTS, you're able to sort of not rely natively on the TTS's engine, but you're building
00:24:32this in-house, uh, for, for this to be managed. And then, yeah, uh, I think I, I don't have my batch
00:24:39here, but I, I don't have my last name on that. So, my first test is if it cannot pronounce my last name
00:24:44or my company's name, it's already dropping the ball. So, my last name is, uh, Balasobramanian,
00:24:50and if you cannot pronounce that using a voice AI agent, uh, like that's a check for me. I, I know,
00:24:56like, uh, uh, you know, the agent will mess up a lot of words that, uh, you know, need to be spelled
00:25:04out, uh, day by day. The second one is our company named Pliwo. So, a lot of engines pronounce it
00:25:09Pliwo or, uh, Pliwo and so on and so forth. But, but I think specifically being able to control this
00:25:16in your pipeline is super critical. And then if you're building a, if you're building a customer
00:25:21facing product, then, then, um, you know, sort of give this option to your customers. All right,
00:25:26I'm just going to skim through the, the, the last two slides. Uh, I'm, I'm running badly over time.
00:25:32Uh, end-of-turn detection, I think this is a separate topic, but I'm just going to
00:25:36quickly pull up all the points so you guys can skim through that. And if, if you need a chat, uh,
00:25:41after this we can, we can talk about this. Right, uh, I'm just going to leave that for like five
00:25:49seconds and then, and then we can chat about this offline. I'm, I'm quite over time. And then the,
00:25:53the, the last one is, uh, barging and, and back channeling. I think there's a lot of talk around
00:25:58speech-to-speech models that do some of this, but we've been able to see how we could do all of this
00:26:03in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this up.
00:26:07Uh, again, I'll just, I just put up, put this up on the slide and, and sort of close at that. Um,
00:26:15all right. I don't think we have time for questions. We can take them offline if you have any time, but,
00:26:19uh, hopefully this was helpful and gave you some insights on, uh, what we are seeing in productions,
00:26:24uh, with billions of calls at scale. All right. Thanks.
00:26:29We'll see you next time.
커뮤니티 글
아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!
이 영상에 대해 글쓰기