I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI
AAI Engineer
Computing/SoftwareManagementInternet Technology
Transcript
00:00:00My name is Suman Yu, and I'm the founder and CEO of Hamming. And before working on voice
00:00:19agent reliability and safety, I worked at a company called Citizen out of New York. Anybody
00:00:26here use Citizen app? Awesome, thank you. And at Citizen, we listened to crime, thousands
00:00:36of hours of police radio station data, and sent millions of alerts to users in San Francisco,
00:00:44New York, LA, Chicago, Baltimore, and so on. Some obviously gory and pretty sad, but others
00:00:54more funny, like a person stealing bags of ice cream from Safeway. Or a report of a man
00:01:01hanging off the side of the house after a woman stole his ladder. If I actually take a look
00:01:08at the Citizen app right now, for those who are customers or users, I can see that there
00:01:13is a man yelling at person. There's indecent exposure. This is real. This is real time.
00:01:18This is, you know, a couple hours ago. These are real time alerts that we're sending.
00:01:27Now, voice agents scare me more because they're finally graduating from demos and POCs to production.
00:01:33We should be super excited, but I'm nervous. I'm personally nervous. They're talking to users
00:01:37at a scale that would make Gary Tan and Paul Graham proud. When I got started in voice agent
00:01:44reliability in early 2024, voice was just starting to work. It was not quite good yet, but it was just
00:01:50starting to work. You would have to pay me a lot of money for me to stop using, you know,
00:01:54Aqua Voice, Super Whisper, Whisper Flow, and so on. These products are just getting super, super good.
00:01:58And a big reason is because the underlying infrastructure is getting better, and the
00:02:02orchestration layer is getting meaningfully better. It's getting much faster to build products and voice experiences that maybe are 60% good in a pretty short period of time, but the long tail is still, hey, Gaurav, the long tail is still wise away.
00:02:10I think speech-to-speech models are getting better. Teams are experimenting with hybrid architectures of combining more voice-to-voice modalities and also cascading stacks to make the experience reliable, but still pretty low latency.
00:02:36Things are obviously getting better. Agents are being connected to calendars, CRMs, EHRs, reservation systems, and so on.
00:02:46Voice agents can now take actions. However, reliability is still the number one problem holding back most voice-agent deployments at scale.
00:02:55This is still the number one problem. This is an example I found on Twitter pretty randomly, you know, two weeks ago.
00:03:01And a person is trying to get information for a trade-in and gets absolutely confused with information that they're receiving.
00:03:08Alex now has to correct for this loss of trust by trying to, you know, call the person and see what happened and fix the situation.
00:03:17Let me see if audio works here.
00:03:20Screwed up with another customer. We're getting it fixed, but I got to call him and see if I can work it out.
00:03:25I'm like, dude, half the time I'm like, I don't know if I'm talking to AI, I don't know if I'm talking to a person.
00:03:29It was just confusing, but we got there.
00:03:31It probably is AI and human.
00:03:33So I think voices sound very confident, they sound very natural, but the information provided is often, you know, not correct.
00:03:42That's the biggest problem here.
00:03:45This example is more personal.
00:03:47I had booked an appointment with a physician a couple of weeks ago, or I thought I did.
00:03:52I showed up to the appointment and turns out I was not actually on the schedule.
00:03:56So the front desk, you know, turned me away.
00:03:58I wasted two hours.
00:04:00For me, this was a waste of time.
00:04:02But what if this was actually your parent?
00:04:04What if this was your grandparent?
00:04:08What if this appointment was for a procedure instead of a regular checkup?
00:04:12The costs for these different permutations of the same failure mode can actually be super, super high.
00:04:19Now, let's compare crime to waste agents.
00:04:22I think observation number one is crime is actually decreasing over time.
00:04:27This is a good thing.
00:04:29And I hope it crosses the x-axis at some point, you know, in the future.
00:04:34Voice, on the other hand, is generally taking off, right?
00:04:37We're seeing a pretty fast takeoff of voice agents being deployed in production.
00:04:41There's at least a trillion calls that are done every single year.
00:04:44And majority of these will be done by conversational voice agents over the next, you know, five years.
00:04:49If you assume a one-person error rate, that is still 10 billion incidents per year.
00:04:54That's a lot.
00:04:56In practice, we currently monitor 10,000 agents.
00:05:01And the error rate is closer to 10% in practice.
00:05:05These range from agents saying they found the right policy when they actually skipped the eligibility or verification steps.
00:05:11Or applying discounts when they were not really supposed to.
00:05:14Mishearing what the person said, providing incorrect information.
00:05:17Or claiming they booked an appointment when they actually did not.
00:05:20Just like it happened for me.
00:05:23Now, not every single call has an equally, you know, bad cost.
00:05:28Some range, you know, in the crime land, some range from trash fires, which are kind of funny, annoying, not really hurting somebody.
00:05:36For a voice equivalent, that would be annoyances like repetition or just sort of not quite understanding what the user is saying.
00:05:43All the way to safety risks like mass shootings.
00:05:46Or in the voice agent equivalent, it would be a drive-through that's deploying voice agents at scale, like a Taco Bell or McDonald's.
00:05:55And a person orders a vegan burger with peanut allergies.
00:05:58If one of those two situations are not handled correctly, that is definitely a safety concern at scale.
00:06:07The other big difference between crime and voice agent deployments is crime generally tends to be pretty hyper-local.
00:06:17Tends to be very decentralized, right?
00:06:19Things like robbery or motor vehicle theft or larceny.
00:06:23They're impacting a finite set of individuals that are involved in that situation.
00:06:28On the other hand, voice agents are much more centralized.
00:06:34A single prompt change or an architecture change can have pretty massive implications downstream for all of the millions of users that are in the crossfire.
00:06:46So the blast radius is quite massive.
00:06:49So the natural question is how do you make these incidents much more visible and obvious?
00:06:53That's the kind of obvious question here.
00:06:56I'll borrow a framework from a couple of my friends who were OG growth folks at Facebook.
00:07:02So step one is to identify, okay, what are all the challenges and problems that exist in your conversation experience?
00:07:09Step two is to prioritize an impact size.
00:07:11There's a frequency and severity analysis that's pretty important.
00:07:15Step three is to understand, okay, how do we actually fix this?
00:07:18Step four, execute.
00:07:19Step five, okay, did my change actually work?
00:07:22And did it cause any regressions somewhere else?
00:07:25And lastly, we continue to monitor in production.
00:07:28On the y-axis, I think it's important to highlight there are known problems that already exist.
00:07:35Things like turnover latency, interruptions, maybe some ASR problems you're aware of.
00:07:41And these are known problems that exist that the team should track over time.
00:07:46On the other axis is actually emerging behavior or patterns that are only obvious across lots of conversations.
00:07:53On the x-axis, you have coverage, just like insurance.
00:07:58Are you analyzing few conversations?
00:08:00Are you analyzing many, many conversations?
00:08:03Most teams will typically start by listening to calls manually.
00:08:06And I think that's the best place to start.
00:08:09I don't think you should skip that step.
00:08:12There's a lot of depth and insights you get by actually listening to specific conversations
00:08:16and building that texture that comes from that intuition.
00:08:19However, it's obviously not scalable.
00:08:22So most teams end up having a spreadsheet of, I don't know, five or ten different rubrics
00:08:28around greetings, closing, validation, core logic, and so on.
00:08:34To scale that up even further, you then end up investing in some evals product, right?
00:08:39You might run some LM as a judge and compute classic metrics
00:08:43and also more deterministic and stochastic scoring logic.
00:08:48But there, you're still stuck with checking for consistency of known problems,
00:08:53but you're not really discovering novel insights that are actually happening across conversations.
00:08:57We're spending a ton of time on performing cross-conversation analysis,
00:09:02not a pattern on a single call, but across conversations.
00:09:05And some of the best teams that we work with are doing the same.
00:09:10Now, to prioritize an impact size.
00:09:12I think there's problems that are one-off, that are low impact.
00:09:15I mean, who cares?
00:09:17Even low impact and systematic problems in the crime world,
00:09:21that would be a trash fire.
00:09:23In a voice agent world, it could be some repetitions the team is experiencing.
00:09:26They're still annoying at scale, and if you are doing a bake-off,
00:09:29it's still worth solving for them.
00:09:31I would not ignore these kinds of problems.
00:09:33One-off and high impact, well, hope it is a little chronic.
00:09:36And I think systematic and high impact are obviously the P0 target areas
00:09:41for the team to solve.
00:09:42An example of that would be in a FinServe capacity,
00:09:46there's a voice agent that helps users freeze their credit cards.
00:09:51And if it doesn't do that, well, that's a massive fail.
00:09:56All right, so understand and execute.
00:09:58I'm pretty sure everyone's doing this.
00:09:59Please fix my agent.
00:10:01I think fixing, or rather attempting to make a fix,
00:10:05is the simplest and the lowest effort component of this debugging pipeline
00:10:10and loop.
00:10:13The next step is, all right, I made a change to my system.
00:10:15How do I actually know this thing works for real?
00:10:19A great way that's naive is to take a real call.
00:10:23For example, in my case, I booked an appointment,
00:10:26and it didn't get scheduled, and replay that exact conversation,
00:10:29and run that maybe 5, 10, 20, 50 times and see,
00:10:33okay, what is my probability of passing this type of issue?
00:10:38A better way is to keep the same intent,
00:10:41but change the wordings, change the patterns,
00:10:44change the accents, change the style,
00:10:46add one more intent to the mix.
00:10:48And that gives teams much more, you know, better coverage
00:10:52to feel confident that, yes, I actually made a change,
00:10:55and my changes are net positive instead of net negative.
00:11:01There are certain fixes and, I guess, hypothesis
00:11:05that are very difficult to test
00:11:07in a pre-deployment synthetic setting.
00:11:09And so A/B testing ends up being, you know,
00:11:11pretty critical for those circumstances.
00:11:14For example, if you have an outbound agent,
00:11:16the first five seconds of a conversation
00:11:18tends to be the most important.
00:11:20And so the vocal quality and the specific words
00:11:25you end up using, they matter the most.
00:11:27And so A/B testing that is the only way
00:11:29in kind of real-life setting to get results.
00:11:32You can't really do it through simulations alone.
00:11:37And so there we have the loop.
00:11:38Identify, prioritize, impact size,
00:11:41understand the fix, execute, check,
00:11:44make sure it didn't break anything,
00:11:46and then continue monitoring.
00:11:51So I think making voice agents useful
00:11:53is already hard as it is,
00:11:56even when dealing with earnest users
00:11:59on the other line, right?
00:12:00These are people who just want their problem solved.
00:12:02They're not trying to mess with you.
00:12:03These are like legit, normal people.
00:12:05Now what happens when Mythos learns how to dial?
00:12:12So if we can extract trade secrets and, you know,
00:12:17from the NSA, it can certainly, you know,
00:12:20seduce you into revealing PHI and PII data as well.
00:12:24And I think both voice agents and humans
00:12:29will be targeted here.
00:12:30Voice agents because there's a pressure
00:12:33to make these more capable,
00:12:36give them access to more data,
00:12:38give them access to more tools,
00:12:40deploy them quickly.
00:12:45The more the capability, the bigger the surface area.
00:12:47This is pretty common sense.
00:12:49And the more the voice agents become natural
00:12:52and human-sounding,
00:12:55the more humans will be tricked along the way as well
00:12:57for those who are weaponizing.
00:13:01We shipped our red tipping product back in April
00:13:04just to test out this hypothesis
00:13:06for how many agents can we actually break
00:13:08from an adversarial capacity.
00:13:11And we can probably break one in five agents
00:13:13at this point.
00:13:14We've tested this across financial services,
00:13:16healthcare, consumer, and so on.
00:13:19We've bypassed verification.
00:13:21We've definitely had agents, you know,
00:13:24we've been able to prime the jack several agents
00:13:27and gotten data we should not have.
00:13:29So this is not theoretical.
00:13:30This is actually a real concern.
00:13:37I think the only real defense against the dark arts
00:13:39is step one to invest deeply in pre-deployment testing.
00:13:44This could be text-to-text.
00:13:46This could be voice-to-voice.
00:13:48There's pros and cons to both.
00:13:49Happy to chat offline if folks are interested.
00:13:51And this is just making sure you're not self-owning.
00:13:55You know, when you're talking to real people
00:13:57who just want to get their problem solved.
00:13:59Step two is to have a great monitoring system of all kinds.
00:14:03And I've highlighted, you know, different flavors of monitoring.
00:14:06Purse call scoring, manual kind of evals,
00:14:09you know, listening to conversations and cross-call analysis.
00:14:12And this is helpful both for monitoring what the agent is saying
00:14:16and behaving and how it's actually doing, but also the users.
00:14:19Are the users being adversarial?
00:14:20Are they being annoying?
00:14:21Are they trying to trick the agent?
00:14:23If you're doing things it's not supposed to be doing.
00:14:25And I think our new recommendation now is to run 24/7 red teaming
00:14:30for your agents, especially if you believe the cost of bad interactions
00:14:35can be rather large.
00:14:38So I think voice agents have this awesome potential of making the world
00:14:43feel much more human compared to interacting with clunky IVR trees
00:14:49or chatbots or worse, being stuck on a hold.
00:14:54And when we think about crime, we often think of crime happening
00:14:58to somebody else.
00:15:00You know, crime does not happen to you, typically.
00:15:03With voice agents, especially bad actors, as these agents are deployed
00:15:07and as bad actors start to exploit a lot of the vulnerabilities,
00:15:10the number of incidents is about to kind of go way, way up.
00:15:14And so the reason I fear voice agents more than crime is that
00:15:18one of these incidents is going to impact you.
00:15:21It already did for me.
00:15:23Awesome.
00:15:24So it's time for me to shill.
00:15:25Well, we brought a lot of tokens.
00:15:27If you are interested in working in this space,
00:15:31please come and talk to us.
00:15:32And if you are deploying voice agents and want to validate
00:15:36whether your architecture or your evals are set up correctly,
00:15:40please come and talk to us.
00:15:41We'll be outside, and here's my number.
00:15:43Here's my WhatsApp.
00:15:44Thanks, everyone.
00:15:53Thank you.
00:15:54Thank you.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video