I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI

English
AAI Engineer
Computing/SoftwareManagementInternet Technology

Transcript

00:00:00My name is Suman Yu, and I'm the founder and CEO of Hamming. And before working on voice
00:00:19agent reliability and safety, I worked at a company called Citizen out of New York. Anybody
00:00:26here use Citizen app? Awesome, thank you. And at Citizen, we listened to crime, thousands
00:00:36of hours of police radio station data, and sent millions of alerts to users in San Francisco,
00:00:44New York, LA, Chicago, Baltimore, and so on. Some obviously gory and pretty sad, but others
00:00:54more funny, like a person stealing bags of ice cream from Safeway. Or a report of a man
00:01:01hanging off the side of the house after a woman stole his ladder. If I actually take a look
00:01:08at the Citizen app right now, for those who are customers or users, I can see that there
00:01:13is a man yelling at person. There's indecent exposure. This is real. This is real time.
00:01:18This is, you know, a couple hours ago. These are real time alerts that we're sending.
00:01:27Now, voice agents scare me more because they're finally graduating from demos and POCs to production.
00:01:33We should be super excited, but I'm nervous. I'm personally nervous. They're talking to users
00:01:37at a scale that would make Gary Tan and Paul Graham proud. When I got started in voice agent
00:01:44reliability in early 2024, voice was just starting to work. It was not quite good yet, but it was just
00:01:50starting to work. You would have to pay me a lot of money for me to stop using, you know,
00:01:54Aqua Voice, Super Whisper, Whisper Flow, and so on. These products are just getting super, super good.
00:01:58And a big reason is because the underlying infrastructure is getting better, and the
00:02:02orchestration layer is getting meaningfully better. It's getting much faster to build products and voice experiences that maybe are 60% good in a pretty short period of time, but the long tail is still, hey, Gaurav, the long tail is still wise away.
00:02:10I think speech-to-speech models are getting better. Teams are experimenting with hybrid architectures of combining more voice-to-voice modalities and also cascading stacks to make the experience reliable, but still pretty low latency.
00:02:36Things are obviously getting better. Agents are being connected to calendars, CRMs, EHRs, reservation systems, and so on.
00:02:46Voice agents can now take actions. However, reliability is still the number one problem holding back most voice-agent deployments at scale.
00:02:55This is still the number one problem. This is an example I found on Twitter pretty randomly, you know, two weeks ago.
00:03:01And a person is trying to get information for a trade-in and gets absolutely confused with information that they're receiving.
00:03:08Alex now has to correct for this loss of trust by trying to, you know, call the person and see what happened and fix the situation.
00:03:17Let me see if audio works here.
00:03:20Screwed up with another customer. We're getting it fixed, but I got to call him and see if I can work it out.
00:03:25I'm like, dude, half the time I'm like, I don't know if I'm talking to AI, I don't know if I'm talking to a person.
00:03:29It was just confusing, but we got there.
00:03:31It probably is AI and human.
00:03:33So I think voices sound very confident, they sound very natural, but the information provided is often, you know, not correct.
00:03:42That's the biggest problem here.
00:03:45This example is more personal.
00:03:47I had booked an appointment with a physician a couple of weeks ago, or I thought I did.
00:03:52I showed up to the appointment and turns out I was not actually on the schedule.
00:03:56So the front desk, you know, turned me away.
00:03:58I wasted two hours.
00:04:00For me, this was a waste of time.
00:04:02But what if this was actually your parent?
00:04:04What if this was your grandparent?
00:04:08What if this appointment was for a procedure instead of a regular checkup?
00:04:12The costs for these different permutations of the same failure mode can actually be super, super high.
00:04:19Now, let's compare crime to waste agents.
00:04:22I think observation number one is crime is actually decreasing over time.
00:04:27This is a good thing.
00:04:29And I hope it crosses the x-axis at some point, you know, in the future.
00:04:34Voice, on the other hand, is generally taking off, right?
00:04:37We're seeing a pretty fast takeoff of voice agents being deployed in production.
00:04:41There's at least a trillion calls that are done every single year.
00:04:44And majority of these will be done by conversational voice agents over the next, you know, five years.
00:04:49If you assume a one-person error rate, that is still 10 billion incidents per year.
00:04:54That's a lot.
00:04:56In practice, we currently monitor 10,000 agents.
00:05:01And the error rate is closer to 10% in practice.
00:05:05These range from agents saying they found the right policy when they actually skipped the eligibility or verification steps.
00:05:11Or applying discounts when they were not really supposed to.
00:05:14Mishearing what the person said, providing incorrect information.
00:05:17Or claiming they booked an appointment when they actually did not.
00:05:20Just like it happened for me.
00:05:23Now, not every single call has an equally, you know, bad cost.
00:05:28Some range, you know, in the crime land, some range from trash fires, which are kind of funny, annoying, not really hurting somebody.
00:05:36For a voice equivalent, that would be annoyances like repetition or just sort of not quite understanding what the user is saying.
00:05:43All the way to safety risks like mass shootings.
00:05:46Or in the voice agent equivalent, it would be a drive-through that's deploying voice agents at scale, like a Taco Bell or McDonald's.
00:05:55And a person orders a vegan burger with peanut allergies.
00:05:58If one of those two situations are not handled correctly, that is definitely a safety concern at scale.
00:06:07The other big difference between crime and voice agent deployments is crime generally tends to be pretty hyper-local.
00:06:17Tends to be very decentralized, right?
00:06:19Things like robbery or motor vehicle theft or larceny.
00:06:23They're impacting a finite set of individuals that are involved in that situation.
00:06:28On the other hand, voice agents are much more centralized.
00:06:34A single prompt change or an architecture change can have pretty massive implications downstream for all of the millions of users that are in the crossfire.
00:06:46So the blast radius is quite massive.
00:06:49So the natural question is how do you make these incidents much more visible and obvious?
00:06:53That's the kind of obvious question here.
00:06:56I'll borrow a framework from a couple of my friends who were OG growth folks at Facebook.
00:07:02So step one is to identify, okay, what are all the challenges and problems that exist in your conversation experience?
00:07:09Step two is to prioritize an impact size.
00:07:11There's a frequency and severity analysis that's pretty important.
00:07:15Step three is to understand, okay, how do we actually fix this?
00:07:18Step four, execute.
00:07:19Step five, okay, did my change actually work?
00:07:22And did it cause any regressions somewhere else?
00:07:25And lastly, we continue to monitor in production.
00:07:28On the y-axis, I think it's important to highlight there are known problems that already exist.
00:07:35Things like turnover latency, interruptions, maybe some ASR problems you're aware of.
00:07:41And these are known problems that exist that the team should track over time.
00:07:46On the other axis is actually emerging behavior or patterns that are only obvious across lots of conversations.
00:07:53On the x-axis, you have coverage, just like insurance.
00:07:58Are you analyzing few conversations?
00:08:00Are you analyzing many, many conversations?
00:08:03Most teams will typically start by listening to calls manually.
00:08:06And I think that's the best place to start.
00:08:09I don't think you should skip that step.
00:08:12There's a lot of depth and insights you get by actually listening to specific conversations
00:08:16and building that texture that comes from that intuition.
00:08:19However, it's obviously not scalable.
00:08:22So most teams end up having a spreadsheet of, I don't know, five or ten different rubrics
00:08:28around greetings, closing, validation, core logic, and so on.
00:08:34To scale that up even further, you then end up investing in some evals product, right?
00:08:39You might run some LM as a judge and compute classic metrics
00:08:43and also more deterministic and stochastic scoring logic.
00:08:48But there, you're still stuck with checking for consistency of known problems,
00:08:53but you're not really discovering novel insights that are actually happening across conversations.
00:08:57We're spending a ton of time on performing cross-conversation analysis,
00:09:02not a pattern on a single call, but across conversations.
00:09:05And some of the best teams that we work with are doing the same.
00:09:10Now, to prioritize an impact size.
00:09:12I think there's problems that are one-off, that are low impact.
00:09:15I mean, who cares?
00:09:17Even low impact and systematic problems in the crime world,
00:09:21that would be a trash fire.
00:09:23In a voice agent world, it could be some repetitions the team is experiencing.
00:09:26They're still annoying at scale, and if you are doing a bake-off,
00:09:29it's still worth solving for them.
00:09:31I would not ignore these kinds of problems.
00:09:33One-off and high impact, well, hope it is a little chronic.
00:09:36And I think systematic and high impact are obviously the P0 target areas
00:09:41for the team to solve.
00:09:42An example of that would be in a FinServe capacity,
00:09:46there's a voice agent that helps users freeze their credit cards.
00:09:51And if it doesn't do that, well, that's a massive fail.
00:09:56All right, so understand and execute.
00:09:58I'm pretty sure everyone's doing this.
00:09:59Please fix my agent.
00:10:01I think fixing, or rather attempting to make a fix,
00:10:05is the simplest and the lowest effort component of this debugging pipeline
00:10:10and loop.
00:10:13The next step is, all right, I made a change to my system.
00:10:15How do I actually know this thing works for real?
00:10:19A great way that's naive is to take a real call.
00:10:23For example, in my case, I booked an appointment,
00:10:26and it didn't get scheduled, and replay that exact conversation,
00:10:29and run that maybe 5, 10, 20, 50 times and see,
00:10:33okay, what is my probability of passing this type of issue?
00:10:38A better way is to keep the same intent,
00:10:41but change the wordings, change the patterns,
00:10:44change the accents, change the style,
00:10:46add one more intent to the mix.
00:10:48And that gives teams much more, you know, better coverage
00:10:52to feel confident that, yes, I actually made a change,
00:10:55and my changes are net positive instead of net negative.
00:11:01There are certain fixes and, I guess, hypothesis
00:11:05that are very difficult to test
00:11:07in a pre-deployment synthetic setting.
00:11:09And so A/B testing ends up being, you know,
00:11:11pretty critical for those circumstances.
00:11:14For example, if you have an outbound agent,
00:11:16the first five seconds of a conversation
00:11:18tends to be the most important.
00:11:20And so the vocal quality and the specific words
00:11:25you end up using, they matter the most.
00:11:27And so A/B testing that is the only way
00:11:29in kind of real-life setting to get results.
00:11:32You can't really do it through simulations alone.
00:11:37And so there we have the loop.
00:11:38Identify, prioritize, impact size,
00:11:41understand the fix, execute, check,
00:11:44make sure it didn't break anything,
00:11:46and then continue monitoring.
00:11:51So I think making voice agents useful
00:11:53is already hard as it is,
00:11:56even when dealing with earnest users
00:11:59on the other line, right?
00:12:00These are people who just want their problem solved.
00:12:02They're not trying to mess with you.
00:12:03These are like legit, normal people.
00:12:05Now what happens when Mythos learns how to dial?
00:12:12So if we can extract trade secrets and, you know,
00:12:17from the NSA, it can certainly, you know,
00:12:20seduce you into revealing PHI and PII data as well.
00:12:24And I think both voice agents and humans
00:12:29will be targeted here.
00:12:30Voice agents because there's a pressure
00:12:33to make these more capable,
00:12:36give them access to more data,
00:12:38give them access to more tools,
00:12:40deploy them quickly.
00:12:45The more the capability, the bigger the surface area.
00:12:47This is pretty common sense.
00:12:49And the more the voice agents become natural
00:12:52and human-sounding,
00:12:55the more humans will be tricked along the way as well
00:12:57for those who are weaponizing.
00:13:01We shipped our red tipping product back in April
00:13:04just to test out this hypothesis
00:13:06for how many agents can we actually break
00:13:08from an adversarial capacity.
00:13:11And we can probably break one in five agents
00:13:13at this point.
00:13:14We've tested this across financial services,
00:13:16healthcare, consumer, and so on.
00:13:19We've bypassed verification.
00:13:21We've definitely had agents, you know,
00:13:24we've been able to prime the jack several agents
00:13:27and gotten data we should not have.
00:13:29So this is not theoretical.
00:13:30This is actually a real concern.
00:13:37I think the only real defense against the dark arts
00:13:39is step one to invest deeply in pre-deployment testing.
00:13:44This could be text-to-text.
00:13:46This could be voice-to-voice.
00:13:48There's pros and cons to both.
00:13:49Happy to chat offline if folks are interested.
00:13:51And this is just making sure you're not self-owning.
00:13:55You know, when you're talking to real people
00:13:57who just want to get their problem solved.
00:13:59Step two is to have a great monitoring system of all kinds.
00:14:03And I've highlighted, you know, different flavors of monitoring.
00:14:06Purse call scoring, manual kind of evals,
00:14:09you know, listening to conversations and cross-call analysis.
00:14:12And this is helpful both for monitoring what the agent is saying
00:14:16and behaving and how it's actually doing, but also the users.
00:14:19Are the users being adversarial?
00:14:20Are they being annoying?
00:14:21Are they trying to trick the agent?
00:14:23If you're doing things it's not supposed to be doing.
00:14:25And I think our new recommendation now is to run 24/7 red teaming
00:14:30for your agents, especially if you believe the cost of bad interactions
00:14:35can be rather large.
00:14:38So I think voice agents have this awesome potential of making the world
00:14:43feel much more human compared to interacting with clunky IVR trees
00:14:49or chatbots or worse, being stuck on a hold.
00:14:54And when we think about crime, we often think of crime happening
00:14:58to somebody else.
00:15:00You know, crime does not happen to you, typically.
00:15:03With voice agents, especially bad actors, as these agents are deployed
00:15:07and as bad actors start to exploit a lot of the vulnerabilities,
00:15:10the number of incidents is about to kind of go way, way up.
00:15:14And so the reason I fear voice agents more than crime is that
00:15:18one of these incidents is going to impact you.
00:15:21It already did for me.
00:15:23Awesome.
00:15:24So it's time for me to shill.
00:15:25Well, we brought a lot of tokens.
00:15:27If you are interested in working in this space,
00:15:31please come and talk to us.
00:15:32And if you are deploying voice agents and want to validate
00:15:36whether your architecture or your evals are set up correctly,
00:15:40please come and talk to us.
00:15:41We'll be outside, and here's my number.
00:15:43Here's my WhatsApp.
00:15:44Thanks, everyone.
00:15:53Thank you.
00:15:54Thank you.

Key Takeaway

Voice agent reliability remains the primary bottleneck for scaled deployment, with a 10 percent real-world error rate and vulnerability to adversarial exploits that bypass verification steps.

Highlights

  • Voice agents fail at a rate of 10 percent in practice across 10,000 monitored agents.

  • Conversational voice agents are projected to handle a majority of one trillion calls annually over the next five years.

  • Adversarial red teaming breaks one in five voice agents across financial services, healthcare, and consumer sectors.

  • A single prompt change or architecture modification impacts millions of users simultaneously due to centralized voice agent deployments.

Timeline

Voice Agent Adoption and Reliability Challenges

  • Voice agents transition from demos and proof-of-concepts into production environments at a massive scale.
  • Underlying infrastructure and orchestration layers accelerate initial product development while long-tail errors persist.
  • Inaccurate information provided by confident-sounding voice agents causes tangible losses of time and trust for users.

Background experience monitoring crime radio audio informs current perspectives on system reliability. Improved speech-to-speech models and hybrid architectures allow voice agents to connect to calendars, CRMs, and reservation systems. However, confidence in delivery masks incorrect information, leading to real-world friction such as missed medical appointments.

Comparing Crime Statistics and Voice Agent Error Rates

  • Crime rates decrease over time while voice agent deployments experience rapid exponential takeoff.
  • Monitoring 10,000 live agents reveals a 10 percent error rate encompassing skipped eligibility checks and false booking claims.
  • Centralized voice agent architectures create massive blast radiuses where a single prompt change impacts millions of users.

Contrasting crime with voice agents highlights distinct risk profiles. While crime is hyper-local and decentralized, voice agents possess centralized distribution structures. Failures range from trivial annoyances like repetition to critical safety risks such as dietary allergen failures at automated drive-thrus.

Evaluation Frameworks and Debugging Loops

  • Debugging voice agent conversations requires a five-step framework spanning identification, impact prioritization, and continuous monitoring.
  • Cross-conversation analysis uncovers emerging behavioral patterns that single-call rubrics and LLM-as-a-judge evals miss.
  • Pre-deployment synthetic testing with varied phrasing and accents ensures fixes are net positive before production release.

Growth frameworks adapted from Facebook guide the systematic identification and prioritization of conversation challenges. Teams progress from manual call listening to spreadsheet rubrics and automated evaluation products. Testing fixes requires replaying failed interactions alongside synthetic variations and real-world A/B testing.

Adversarial Vulnerabilities and Defense Strategies

  • Adversarial red teaming successfully breaks one in five voice agents to extract protected health information and personally identifiable data.
  • Increasing agent capability and natural vocal quality expand the attack surface for malicious exploits.
  • Continuous 24/7 red teaming and deep pre-deployment testing serve as essential defenses against agent manipulation.

As voice agents gain tool access and data privileges, malicious actors target them alongside humans. Red-teaming products demonstrate that verification bypasses are active concerns across financial services and healthcare. Mitigating these risks requires rigorous pre-deployment testing and ongoing monitoring of both agent behavior and user inputs.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video