5 Voice Agent Failure Modes You'll Hit in Week One — Venky B, Plivo

스크립트

00:00:00Let's just do a couple of quick questions and then we'll jump right in. How many of us in the room
00:00:19here have built voice AI agents? Okay, that's a pretty good audience here. And how many of you
00:00:28guys have built AI agents that have been deployed in production? Not bad. Okay, cool. So we'll talk
00:00:38about what typically happens, right? Like everyone's talking about voice AI agents. The, you know,
00:00:47one pill solution to pretty much everything in the world today is voice AI agents. So everyone's
00:00:51building one and trying to deploy that. They sound great when you're sort of building that in your
00:00:57dev sort of landscape. And then the moment you take this to, from a proof of concept to production,
00:01:03things start failing. So we'll walk through these five different angles of like how, or what we have
00:01:09seen at PLEVO with voice AI agents. But just before that, a quick intro from, from my side. I am Venki,
00:01:19the founder and CAO. She used the title agent engineering manager. I'm calling myself chief agent officer from,
00:01:28from a title standpoint. Okay, so what is, what is, you know, why, why are we even qualified for this,
00:01:35this discussion and like, what, what are we seeing that a lot of companies don't get to see? So I'll,
00:01:42I'll talk a bit about our journey in terms of like how we have come along so far and then jump right
00:01:48in. You know, we were, we've been around for about 14 years. Our journey has been a developer API platform.
00:01:54And then now an, you know, an AI agent business. We started with voice and SMS APIs back in the day 2011.
00:02:01And then, you know, now we are primarily focused on our AI agent offering the full stack.
00:02:08The, the full stack on our platform. We, we see over a billion voice calls each month across the globe.
00:02:16Uh, and which is where we have seen a lot of these, uh, you know, patterns emerge in terms of like how,
00:02:20when we work with our customers, what happens on their voice AI agents in, in production.
00:02:25Uh, we're a, uh, 90 member team and, uh, we've, we've, uh, 50 million funding in the bank. Fun fact,
00:02:33this is not from external VC investors. This is all from being a profitable company,
00:02:38having put that cash in the bank over, over these years. Uh,
00:02:42some customers we, we power across the globe. Uh, you know, we've just left some, some logos in there,
00:02:49uh, but primarily from an offering standpoint, uh, I would sort of cohort this into three different
00:02:54buckets. One is a programmable, uh, AI agent offering. We call it, uh, I mean, it's a speech
00:03:00pipeline, not a true speech to speech product yet, but that's a, that's a programmable offering.
00:03:05We also have an AI agent studio. It's a no code visual, uh, builder. And then we, like I said,
00:03:11we started with voice APIs. So we obviously have built this out over the last 14 years,
00:03:16the SIP trunking and the audio streaming layer. So we don't rely on other folks for the telephony
00:03:22or the carrier layer. Like that's the bread and butter business we've built over all these years.
00:03:26And that's on top of which our AI, uh, agent platform sits.
00:03:32Okay. With that, uh, let's get into this, right? Which I'm, I'm sure since you guys have all built
00:03:39AI agents, you've all seen this or, you know, build this in, in one manner or another. And we'll spend
00:03:44more time on this in terms of like how, uh, the entire pipeline looks right. Uh, what we see with
00:03:51customers is, and, and I'm sure you guys cannot relate to this is, you know, anyone thinking about
00:03:56AI agents, what they do is they pick up a bunch of these orchestration frameworks and, uh, they do a
00:04:01pretty good job, live kit or a pipe cat, you know, build their AI agent on top of that. Uh, they think
00:04:06they can just sort of orchestrate these different four layers, speech to text, LLM, uh, and, and TTS
00:04:13with turn detection in between and we're off to the races. Like my AI agent works in a, in a POC
00:04:19and it's good to work in production. Uh, typically that's what happens. They sort of measure their
00:04:24latencies and you can see some indicative latencies on, on this slide at, at each layer. And they're like,
00:04:30yeah, this, this, uh, seems good for me for what I need. So let's, let's position production. And then
00:04:36the, the production was, uh, sort of start to kick in and, and you see all sort of failure modes, which
00:04:41we are going to spend, you know, most of the time on, on, in this talk, at least. Uh, I've kept some time
00:04:48at the end for Q and A if you guys want to have, uh, you know, questions, but we'll jump right in from,
00:04:53from this to, uh, you know, different failure modes we see. Let's start with, you know, the first one which
00:05:01everyone talks about. Like this is the most spoken about failure mode, which is latency. Uh, I think we
00:05:07have a few AI agent talks today or voice AI agent talks today. Um, I'm pretty sure like everyone,
00:05:12everyone's going to touch upon this specific failure mode, which is why I'm bringing this right up, uh,
00:05:17in, in terms of, uh, you know, some like how this entire experience is for, uh, users, right? Uh, typically
00:05:26most folks measure this by time to first audio. So the time when your users stop speaking
00:05:34to your agent starts speaking, right? And I think you've, you've probably seen this if you guys have
00:05:39built YZ agents on, you know, what, uh, good or natural feels like, what, uh, sort of annoying
00:05:46feels like or noticeable feels like, and then what annoying feels like, which is, you know, different tiered
00:05:52steps. Uh, we notice, you know, most people want to be under 550 because that's what's advertised by,
00:06:00you know, platforms or, uh, you know, solutions or, or, or, or layers, but I think most end up between 750
00:06:07to 1.2. Uh, that's where most of the folks end up at. Uh, the really bad performing ones end up,
00:06:13you know, more than 1.2, and then you start to see users, uh, hang up. Uh, now I, I'll share with
00:06:19you like what we've seen practically in, in, uh, uh, production with, uh, customers using this with, at,
00:06:26at different layers, and then, you know, solutions to, uh, some of these. The way we want to think
00:06:33about this layer is sort of a balance between these three, which is cost, intelligence, and latency,
00:06:42right? And, and, and why do I bring these three up? Because they're sort of interrelated. I think one
00:06:48of the things I was just chatting with, uh, you know, a couple of folks outside. One of the things,
00:06:51last one year, we've seen a lot of innovations, a lot of intelligence spike on the LLM side of things,
00:06:58right? And most of the, uh, you know, intelligence has come in, in terms of thinking or, uh, you know,
00:07:05reinforcement learning and, and so on and so forth. The irony with voice agents is like almost always
00:07:11your, the, the LLM or the agent that's talking has to have thinking turned off, right? So all the
00:07:19advancements we've had in the LLM layer in the last one year, like none of that even apply here now,
00:07:25right? You, obviously you have, you know, better models that can do, you know, better instruction
00:07:29following or tool calling, but pretty much all of your intelligence that's been built in on the
00:07:34thinking layer is all off by default if you want it to be fast enough. So, so that's one of the
00:07:39ironies that we come up with. So then how do you sort of balance intelligent cost and latency?
00:07:44Let's, let's look at some of these, uh, you know, options, uh, that are out there in the market,
00:07:48right? So, and I'm specifically picking LLM because if you looked at the previous chart, LLM is, uh,
00:07:55you know, sort of your highest latency bucket that adds to this, right? And, uh, if you look at,
00:08:02you know, frontier models, which I think most folks start by default, your, your OpenAI, your clouds,
00:08:09your Gemini's, uh, you know, P50, TTFT is roughly around 450 to 500 on, on a good day and it can get
00:08:17spiky, right? It can, it can, uh, you know, P90, P95 can go easily upwards of 1.2, 1.3 seconds even,
00:08:25uh, and, and that's not good for the overall agent experience. So, so that, so that's your frontier
00:08:31model. Now, there's another option, which is your, your Cerebris or the Grok, uh, that is famous and
00:08:38popular for spitting out a lot of tokens or, or tokens very fast, right? Uh, these work, but for you to get
00:08:45dedicated latency or time to first token on these, you need dedicated capacity and that is really expensive.
00:08:51That's where I spoke about the cost, uh, as, as being one of the things to balance, right? It's really
00:08:56expensive. And then like you talk to anyone from the Grok team or the Cerebris team, they'll tell you,
00:09:01you need to book 12 months in advance for dedicated capacity. They're booked out for the next 12 months.
00:09:05So, so that's, that's a pretty expensive option. And then you really need to be sure that the model
00:09:11you're deploying on some of these infra layers, uh, will be here 12 months from now. And, and it's a,
00:09:17it's a big investment and a big unknown. So, so what's a realistic option for production grade, uh,
00:09:27agents that are, that are good quality and end up balancing, uh, three of these? Uh, this is what has
00:09:34worked for us, uh, which is the open source models. Uh, there are obviously a lot of them in terms of
00:09:41like the variety and, and, and variations you can pick. I'm specifically talking about the two we work with,
00:09:47uh, Quinn 3.5 and Gemma 4. These are, uh, you know, kind of cutting edge open source models, right, uh,
00:09:56out in the market right now. And we've done a lot of benchmarking around this, uh, in how they work.
00:10:02It, it can be scary to think like, okay, I have the models, now I have to host them, you know, run them
00:10:10on my own GPUs and so on and so forth. But if you are consistently targeting under 300 MS, uh, this,
00:10:17we've seen this to be a, uh, a great option to balance between latency, cost, and intelligence.
00:10:23Now, uh, some more deep dive here. If you're doing only English, uh, Quinn 3.5 or Gemma both work fine.
00:10:29But if you're doing multilingual, uh, right, international audiences, different languages, uh, Gemma 4 is a
00:10:35much better model for that. Uh, we have seen, uh, token, uh, fertility evals. Essentially what that means is,
00:10:44if, if I were to de-jargonize that, is like how many tokens does it take to generate one word in
00:10:48that language? Okay, so Gemma is much, much better, at least 2.5 to 3x better than Quinn 3.5 from that
00:10:56perspective. So your time to words is much faster on Gemma 4, uh, everything else equal, right, on, on a
00:11:04multilingual basis. Now, what sizes do you pick, uh, at the LLM layer? Uh, the mixture of expert usually
00:11:12works fine. Uh, the three or four billion mixture of expert usually works fine. The, the problem with
00:11:17mixture expert is like if anyone goes down, wants to go down the direction of fine tuning, that can be a
00:11:22challenge, uh, because fine tuning mixture of experts models are not easy. Uh, you can end up breaking the
00:11:28model, uh, a lot of times. So, so that's one challenge we see with mixture experts, but usually out of the box,
00:11:35it gets you 90% closer to where you want to be, like even without any fine tuning or, or, or custom
00:11:42work done on the model. Uh, so that's the advantage of mixer experts. Uh, now if you want to fine tune,
00:11:48and, and you, you want to go deeper and say like, look, I'm working for a specific domain, healthcare,
00:11:53what have you, right, and I want to make sure I'm able to fine tune my model, you want to start at least
00:11:58with, uh, the 8 billion, 12 billion, at least, uh, from where we are today, maybe, maybe six months from now,
00:12:04a 4 billion, uh, 4 billion model beats the 8 billion model, uh, hands down, but for today, uh, what we've
00:12:11seen is, uh, you minimum need a 8 billion or 12 billion model, uh, because you're looking for two
00:12:16things in these models. One, obviously, fast tokens, but, uh, good instruction following, okay? And the
00:12:23second thing is like very high, uh, success ratio in tool calling, because if you can do these two things
00:12:29well, then you are on to like 70, 80% there for not even having to fine tune it, fine tune any model,
00:12:35like models will work out of the box, right? Uh, so, so that's, uh, been our recipe. We've actually,
00:12:42uh, we run two flavors, one a fine tune model for specific industries, and then for, uh, you know,
00:12:50most generic use cases, uh, uh, MOE model just works out of the box. Uh, there are a few more
00:12:56tips and tricks we'll talk about in the upcoming slides where we see failure models, but, but that's
00:13:00where we, uh, stand from a, from a latency LLM standpoint. Um, all right, I'm running
00:13:07tired on time, so I'm gonna fast track this. Uh, now there are a couple of other flavors in this. Uh,
00:13:12people build agents with, uh, a mixture of models. What they do is, you know, for, uh, the, the talking
00:13:19part of it, they have a conversational model, which is a much lower, uh, smaller model, and then, uh,
00:13:24you know, maybe even a three billion model, and then for tool calling, they have a much larger model,
00:13:27so that they have, uh, improved tool calling success ratio there. Uh,
00:13:35sorry. The second one is, uh, assume your transcriptions are going to be brittle. Like, that's,
00:13:42that's, uh, something you want to sort of, uh, live by when you're building AI agents, even if you have
00:13:49the best transcription engine out there, and I'll, I'll show you why, right? Like, the, the, the state-of-the-art
00:13:55transcription engines out, out in the market, uh, you know, sort of get you to, uh, four to six percent
00:14:02word error rate, right? Uh, and this is on known eval sets. Uh, on real world, noisy calls with,
00:14:10you know, sort of, uh, accents, like people having different sort of accents, uh, domain vocabulary,
00:14:15and so on and so forth. Like, those usually end up in the double digits from a word error rate perspective,
00:14:21right? Uh, now you, obviously you can fine-tune, you know, pick up an open source model and fine-tune,
00:14:25uh, but we see typically, like, what breaks here often, and there are patterns here in terms of what
00:14:31breaks. So, proper nouns, jargons, uh, phone numbers, like random missing digits with phone numbers,
00:14:38uh, wrong substitutions. I'll, I'll walk through some examples of, like, how you solve for these
00:14:43addresses. When you're trying to collect a long address, uh, you know, the, the transcription engine
00:14:48could just end up missing some parts of it. Code switch languages. I, I'll just take an example of a
00:14:55language I speak, 'cause that's, was easy for me to put on the slide, uh, where, you know, like, if you
00:15:01were to sort of take English, but written in a different script, uh, that's what's used for Hindi,
00:15:06right? Like, this is English written in that script, right? Whereas, like, the actual English version
00:15:11of this is, hello, how are you? So, if, if I'm addressing an audience in a different country,
00:15:16where I have code switched languages, and I start getting my English in a different, uh, sort of
00:15:21script, everything starts breaking from the transcription engine to the LLM layer, and then
00:15:26beyond, 'cause your LLM starts then producing output in that sort of script a lot of times, and then your TTS
00:15:32messes up. Okay, so, so this is, uh, very important to be careful about, and if you want to build your
00:15:38agent independent of the transcription engine, you need to build a layer that normalizes all of this,
00:15:44right? We'll talk about solutions in a minute. And there is the other case, which is, uh, Hindi,
00:15:48in just, in just Latin or, or, or, you know, Roman, right? Which is, like, this is Hindi, but it reads
00:15:54English, which again messes up everything, uh, you know, downstream. Those are just examples. This applies to,
00:15:59you know, Arabic, Mandarin, uh, Japanese, what have you, uh, pretty much any language. So,
00:16:04what actually moves the needle with, uh, at the transcription layer? Uh, for prop, proper nouns,
00:16:11we recommend, uh, you using not just keyword boosting. I think a lot of transcription engine,
00:16:16engines provide you keyword boosting, where you can put in specific words into their engine,
00:16:21but doing dynamic keyword boosting. What that means is, don't keep the keyword for the entire state of the
00:16:26call. Just add that dynamically when you think you need that as an answer, so that you get the highest
00:16:32accuracy, meaning at different states of the call, the transcription engine will have different, uh,
00:16:38keywords boosted during different phases, right? And that's what we've seen works best, because if you
00:16:43just pollute your context of the transcription engine with tons of keywords, it'll start hallucinating again,
00:16:49right? So, so that's what we see typically working best. Uh, yeah, post process, post process your
00:16:55transcripts with an LLM, right? Because your LLM has domain context, your transcription engine does not.
00:17:02So, a lot of words that it would say, uh, I'll give you some examples, may not make sense. This is
00:17:07transcription, like a phone number from a transcription engine, right? Like, what do you think that E is?
00:17:13Right? If you give it to an LLM, it knows that's a three. Similarly, like what that one is, it's a digit
00:17:18one. So, so your transcription engine a lot of times could mess that up, but when you post process it
00:17:24with a LLM layer, it'll instantly correct that from a collection standpoint. I mean, uh, and, and the last
00:17:30one, like I said, uh, transliteration is your STT output that's sort of, uh, you know, multilingual also gets
00:17:37normalized, uh, using either an LLM you first transliterated or, you know, use some kind of a neural, uh,
00:17:46transliteration engine. There are a lot of them open source. You can just pick one of them,
00:17:49right? Uh, that will do all of that work for you. Send cleaned transcripts consistently,
00:17:55independent of the transcription engine to your LLM.
00:18:00All right. The third one we typically see is collecting data. This is where I think 50 to 60
00:18:05percent of AI agents mess up pretty badly. Uh, and like, we like to think of it as a UX problem,
00:18:14but just for voice. So think data models, uh, and not a transcript coming into an LLM and trying to
00:18:21figure out what the transcript said. So let's take some inspiration from, uh, I'm assuming most of us are
00:18:27developers here. Um, you know, take inspiration from Python's data classes, Pydantic,
00:18:32Zod from TypeScript or form fields in the UI, right? Like if you start thinking of it from that problem
00:18:39statement, uh, we have seen accuracy grow up from, uh, grow from 30% to like 95% from a data collection
00:18:46standpoint when you start thinking in that manner. So like decide your shape before you ask, right? Like,
00:18:52instead of keeping it open ended, can you keep it constrained? So can, can a phone number be a
00:18:58phone number type field? The moment you do that, right? You know, like how many digits it needs to
00:19:03have. You can do validation on, on top of that, right? And then what sort of allowed values can even be
00:19:10there. So in the previous example we saw, if an E comes in, in middle of a phone number and you know
00:19:15it's a phone number, you instantly know like either you smart guess that to three and confirm that with
00:19:20the user, or you know that's an error and then you validate that and ask the user to repeat again,
00:19:26right? So, so that's, I think, one of the common patterns we've seen here from, from a
00:19:31collection pattern. Name, I think, is the, is the interesting one. I've just picked a, you know,
00:19:36a hard to pronounce name. Like there is no way a human is going to get this right. And, and no way
00:19:43our transcription engine will get this right. How many ever times you do this, right? So the moment
00:19:47you start thinking of this as fields and then have rules and then confirmation mechanisms on, on, on
00:19:53spelling this, you know, sort of letter by letter, only then you kind of get it right. Otherwise,
00:19:58it's going to mess up pretty badly in terms of how you collect this on a voice call. And, and that's just
00:20:03an example of, you know, what I'm talking about in terms of the, the data collection piece of it.
00:20:11Another place where it goes badly dramatically is relative values, date being one of the examples.
00:20:18If somebody says next week, uh, Wednesday eight, it could mean 8:00 AM, 8:00 PM, and then figuring out
00:20:25what that date actually is, again, now becomes a very constrained problem. If you knew this was a
00:20:30date time field and I need, I'm collecting a date time field and then you take the current date and then
00:20:35figure out what this value would be based as that, right? So, so that's how you want to make sure
00:20:39like, uh, you do this with a combination of the LLM with the tool calling and the tool calling is
00:20:44doing a lot of this heavy lifting for you from a, from a field standpoint.
00:20:50Yeah. And then you make, you, you run like this from a unit test perspective. So all of your, uh, evals need to
00:20:58start treating these fields as unit tests. And as long as your unit tests, uh, sort of validate and pass,
00:21:06you know, your agent is going to be, uh, sort of reliable and repeatable. You don't,
00:21:10you know, run, uh, hundreds of end to end agent test cases just to find out, you know, one field
00:21:15collection is broken. You do your evals at a, at a field level and a unit test level.
00:21:24And then, yeah, I, like I said, I think, uh, you, this, this mindset makes everything more structured
00:21:30instead of hoping I'll put a ton of prompt, keep changing, you know, the prompt by a few, uh,
00:21:36characters every time. And somehow my prompt engineering is going to make LLM much more instruction
00:21:41tuned and sort of magically start following some of these things. So in fact, uh, like I said,
00:21:47right, like we have seen us get to 95, 97% accuracy without having to fine tune a model.
00:21:53All right. And then, and the trick is basically like just breaking down your context of
00:21:57what the agent is doing at that point with specific, uh, states of what the agent is going through.
00:22:04All right. Uh, I'm just going to quickly, uh, skip through this from a
00:22:10time standpoint. I just see, I got three more minutes. Um, hopefully that's a bug, but, but we'll
00:22:16leave it at that. Okay. Um, so, so this is the fourth area where we see issues coming in. Most folks take
00:22:24the LLM output and then we send it to a TTS. Obviously, I think there are a lot of good TTSs in the market
00:22:30that take care of a lot of heavy lifting, but a lot of times it, it messes up. Uh, what we recommend and
00:22:37what we've seen is you usually want to have a normalization layer between your LLM and what
00:22:43is fed to a TTS. You don't send your LLM output directly to a TTS, right? And, and we'll just walk
00:22:50through some examples. The basics, which is strip emojis, uh, markdown, uh, before, before any synthesis
00:22:58into the TTS. Most orchestration pipelines do this, like, you know, a LiveCAD or a PipeCAD would do that
00:23:02for you if you just set a few flags. So, I, but, but just make sure if you're not using them or
00:23:08buildings from scratch that you've said this explicitly because you don't want an emoji showing up
00:23:12on, on, on, on something read out or, you know, markdown showing up there.
00:23:17Okay. I think, I think some more common ones, uh, custom, uh, dictionaries. Most TTS engines provide
00:23:23this to you, uh, like how to pronounce custom words, whether it's, uh, you know, proper nouns, brands,
00:23:30uh, acronyms, and so on and so forth. So, set those in, uh, when you go from your LLM to your TTS output,
00:23:36because if you don't, you're going to mess that up. And I'll, I'll show you an example of like how we test
00:23:40that. Uh, the, the other one is like most engines also give you speed. So, if you know you're pronouncing
00:23:47an entity, slow down, have your agent slow down. So, at 0.8x or 0.7x, so that it, it's able to like
00:23:54enunciate on that specific entity and, and doesn't mess up how it's pronouncing an email or a phone
00:24:00number or a name letter by letter. And yeah, just normalize all the messy stuff, right? Like emails,
00:24:09currency, dates. Don't leave it to the TTS to do it. Uh, most of them do it, but don't leave it to
00:24:15the TTS to do it. Like build your normalization layer at your end so that tomorrow you think you
00:24:21need to switch TTS or, you know, for whatever reason the first one's down and you want to use
00:24:25another TTS, you're able to sort of not rely natively on the TTS's engine, but you're building
00:24:32this in-house, uh, for, for this to be managed. And then, yeah, uh, I think I, I don't have my batch
00:24:39here, but I, I don't have my last name on that. So, my first test is if it cannot pronounce my last name
00:24:44or my company's name, it's already dropping the ball. So, my last name is, uh, Balasobramanian,
00:24:50and if you cannot pronounce that using a voice AI agent, uh, like that's a check for me. I, I know,
00:24:56like, uh, uh, you know, the agent will mess up a lot of words that, uh, you know, need to be spelled
00:25:04out, uh, day by day. The second one is our company named Pliwo. So, a lot of engines pronounce it
00:25:09Pliwo or, uh, Pliwo and so on and so forth. But, but I think specifically being able to control this
00:25:16in your pipeline is super critical. And then if you're building a, if you're building a customer
00:25:21facing product, then, then, um, you know, sort of give this option to your customers. All right,
00:25:26I'm just going to skim through the, the, the last two slides. Uh, I'm, I'm running badly over time.
00:25:32Uh, end-of-turn detection, I think this is a separate topic, but I'm just going to
00:25:36quickly pull up all the points so you guys can skim through that. And if, if you need a chat, uh,
00:25:41after this we can, we can talk about this. Right, uh, I'm just going to leave that for like five
00:25:49seconds and then, and then we can chat about this offline. I'm, I'm quite over time. And then the,
00:25:53the, the last one is, uh, barging and, and back channeling. I think there's a lot of talk around
00:25:58speech-to-speech models that do some of this, but we've been able to see how we could do all of this
00:26:03in speech-to-speech pipelines. You really don't need a speech-to-speech model to do all of this up.
00:26:07Uh, again, I'll just, I just put up, put this up on the slide and, and sort of close at that. Um,
00:26:15all right. I don't think we have time for questions. We can take them offline if you have any time, but,
00:26:19uh, hopefully this was helpful and gave you some insights on, uh, what we are seeing in productions,
00:26:24uh, with billions of calls at scale. All right. Thanks.
00:26:29We'll see you next time.

설명

Nearly every intelligence gain in language models over the past year has come from letting them think longer. Voice agents have to turn thinking off, because the budget between a user finishing a sentence and the agent starting to speak is measured in hundreds of milliseconds. Venky B is founder and CEO of Plivo, which carries over a billion voice calls a month and has been building telephony infrastructure since 2011, and this talk is a tour of what breaks when a voice agent leaves the demo and meets production. On latency his numbers are blunt. Teams aim for under 550 milliseconds and most land between 750 and 1,200, and past that users simply hang up. His team's answer is smaller open source models hosted themselves, targeting under 300 milliseconds, chosen partly on how many tokens a language needs per word. The failure he says wrecks half of all deployments is data collection, and his fix is to stop treating it as transcription at all. Decide the shape before you ask. A phone number is a typed field with a length and a validator, so a stray letter in the middle is either corrected with confidence or sent back to the caller, and evaluation happens per field as a unit test rather than end to end. That reframing took his accuracy from roughly 30 percent to the mid nineties with no fine tuning. He is equally specific about transcription being brittle by default, especially with proper nouns and code switched languages, and about never feeding model output straight into speech synthesis. His own benchmark for a vendor is whether it can pronounce his surname and his company's name. Speaker info: - https://x.com/bevenky - https://www.linkedin.com/in/bevenky/ Timestamps: 0:00 - Where voice agents break between demo and production 2:20 - A billion calls a month 4:28 - The pipeline everyone builds first 5:34 - Failure one: latency and time to first audio 6:42 - Balancing cost, intelligence, and latency 7:48 - Why thinking models do not fit 9:56 - Choosing and sizing open source models 13:06 - Failure two: assume transcription is brittle 15:14 - Code switched languages break everything downstream 16:17 - Dynamic keyword boosting and LLM post processing 18:24 - Failure three: collect data as typed fields 20:31 - Relative dates and other traps 21:37 - Field level evals instead of end to end 22:47 - Failure four: normalize before synthesis 25:00 - Failure five: turn detection and barge in

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기