Voice Agents Can Just Do Things — Charlie Guo, OpenAI

English
AAI Engineer
컴퓨터/소프트웨어

스크립트

00:00:00So my name is Charlie, and I work on the developer experience team at OpenAI.
00:00:18And part of my job is talking to developers to understand and see what and how they're
00:00:25building with our models, whether that's text, image, or audio.
00:00:31And lately I've been thinking about a misconception that I have seen, or maybe it's just a misunderstanding.
00:00:39And it's the idea that voice agents have to talk back.
00:00:48And to some of you that might sound, you know, absurd.
00:00:51It's a voice agent.
00:00:52What do you mean it's not supposed to talk?
00:00:54But I think if there's one thing that you take away from this presentation, I would like
00:01:00it to be the idea that speech is not the only way that a voice model has to respond.
00:01:12And I think models are getting intelligent enough and capable enough that they're starting to
00:01:17open up some new modes of design.
00:01:21There's actually three kind of modes that I kind of see emerging these days, right?
00:01:27Speech to speech, speech to action, and event to speech.
00:01:32And there's a couple of things I think worth pointing out about these three categories.
00:01:37The first is that they're not new, right?
00:01:41I think as we've seen from previous talks, even just today, there's a long history of building
00:01:47these types of systems in and around voice.
00:01:50If you squint, you could make the argument that the movie phone hotline where you called
00:01:54in to get showtimes was an example of a speech-to-speech system.
00:01:58And I think you could pretty reasonably make the argument that GPS navigation in your car,
00:02:04which has existed since I was a kid, is an example of an event-to-speech system.
00:02:09So it's not that they are brand new, but I think it is that we are able to do some much more
00:02:14interesting things with them now that we're in this era, right?
00:02:18And I think the other thing I would mention here is that they're remixable, right?
00:02:24They're not meant to be mutually exclusive, and I think as we'll see in a little bit, the
00:02:27best products exist in a way that combines all of these modes.
00:02:32So speech-to-speech, everybody knows it, hopefully everybody loves it.
00:02:36The user talks, and then the model talks back.
00:02:40And I think there are a few examples that I can give for this type of use case, right?
00:02:45You've got things like live practice and coaching, especially around language learning, right?
00:02:50I think the ability to hear a lot of emphasis or emotion and give people that feedback is
00:02:55really powerful.
00:02:57I think you can have, you know, what I am sort of cheekily calling concierge experiences,
00:03:01which I think is just another way to say customer support plus-plus.
00:03:06The first voice tutorial that most people, you know, try to build when they have access
00:03:10to this technology is some sort of customer support chat bot for very good reasons.
00:03:14But I think with, you know, when we add a richness and a depth to the voice models, as has been
00:03:19happening in recent months and recent years, we can build something that is, like, much more
00:03:24enjoyable to use than, like, talking your way through a phone tree.
00:03:28And so, you know, I have the hope that, like, soon if not, like, you know, now we are capable
00:03:33of building support experiences with agents that actually feel much more enjoyable to talk
00:03:38to than, like, arguably, like, the median human support agent.
00:03:44And as we just saw, if you were here for the last talk, live translation, right?
00:03:47The models have gotten good enough and fast enough that we can just dynamically translate
00:03:51content on the fly with, like, little to no latency.
00:03:55It wouldn't shock me if at next year's keynote, you know, they live streamed it from the main
00:03:58stage but also dubbed it in real time across multiple languages.
00:04:05The second category is speech to action.
00:04:07These are talks and the model uses tools.
00:04:10And I think this is one of the most under-explored areas that we have.
00:04:15I actually almost titled this talk "Voice is the Next Capability Overhang" because I think
00:04:20there is just a vast, vast amount of stuff that we could be doing in this category that we are
00:04:25not currently doing.
00:04:27For example, there's a broad spectrum.
00:04:30I don't have, you know, there's way too many examples even fit on this slide.
00:04:33But three categories that I find particularly interesting.
00:04:37First is form filling, right?
00:04:38So much of the internet is just filling out forms.
00:04:43And there is, you know, today no reason why you shouldn't be able to just talk.
00:04:46And I would love it if instead of spending an hour filling out a government document, I
00:04:50could just talk for five minutes and it would get 90% of it for me and I would do a quick
00:04:55check, you know, just to make sure that everything looked good, right?
00:04:57That is a vastly superior experience than, like, having to type in every single name and address
00:05:02that I've lived in the last five years and, you know, all of my previous identities.
00:05:05And so I think that one is, though it may seem boring, you know, affects a significant GDP
00:05:11of the internet, right?
00:05:13The next category is creative tools, where I am privileged enough that I can speak the
00:05:17language of software.
00:05:18And so I can tell codecs, you know, here's exactly what I want you to build and I can articulate
00:05:23it in a way that I get much more leverage than sort of just, like, cludgily trying to iterate
00:05:28one thing at a time.
00:05:29But I can't do that when it comes to, you know, using, making music or painting.
00:05:33And so if I don't have the ability to articulate the exact aesthetic that I'm looking for and
00:05:40if I don't know how to use Photoshop or Ableton, I'm left in this state where, you know, my taste
00:05:45exceeds my capability.
00:05:46And so I'm really looking forward to integrating voice into creative tools so that I can just
00:05:51sort of cludgily go along and, you know, vibe create, vibe compose, vibe paint, and make
00:05:56something that's really beautiful to me.
00:05:59And I think the generalizable category here, right, then just starts to become computer use.
00:06:05And we've already seen some companies start to do this.
00:06:07You know, it raises the question of, like, look, if the models are just getting good enough
00:06:11to do everything on a computer that a human can do, like, why am I talking to an app?
00:06:17Why am I talking to a terminal?
00:06:18Why am I not just talking to the entire computer?
00:06:22And so I think that's sort of a really interesting way to start exploring.
00:06:25But if you're a developer today, right, oops, if you're a developer today, what does that
00:06:30mean for building your own software, right?
00:06:32And I think it is, like, much easier than you think to start adding audio as an intelligence
00:06:36layer to the -- intelligence layer to the apps that you already have.
00:06:40If you're building a modern web application, you already expose so much of it as like actions,
00:06:45as nouns and verbs, right?
00:06:46If you think about all the verbs that you have, you have API endpoints, you have, you know,
00:06:51React hooks.
00:06:52Each of those things can, like, pretty relatively easily be converted into a tool that you expose
00:06:56to a model.
00:06:57And then you can give the user the ability to just drive your existing software with their
00:07:01voice, right?
00:07:02And yes, you still need guardrails.
00:07:03You still need safety checks.
00:07:04Like, many of the talks today are going to talk about securing and, you know, productizing
00:07:08this.
00:07:09And for this, I just want you to think about, you know, what would it mean to take your existing
00:07:12software and just talk to it?
00:07:17And to go back to that misconception, right, I think there are a lot of, you know, like,
00:07:20if you're talking to the software, maybe you can talk back, but we've been developing other
00:07:24ways of communicating with the user for decades, right?
00:07:26We know these things.
00:07:27We know we can show notifications and popups.
00:07:29We can change state, like the color of a button or a drop shadow.
00:07:32We can highlight text.
00:07:34If you've used computer use in the Codex app, you know, there's this amazing, like, little
00:07:37ghost cursor animation that goes around and clicks things for you.
00:07:40So we don't have to use words to actually tell the user what is happening on screen with
00:07:44their software.
00:07:48And then the last bucket here is event-to-speech, right?
00:07:51The model receives an event and talks to the user.
00:07:54And sort of the counterpoint from speech-to-action, I think this one is still very, very exploratory,
00:07:58right?
00:07:59You know, if you saw Quinn's talk, I think there's a lot of space here of, like, things
00:08:02we can do.
00:08:04And to me, we haven't quite seen what AI native really looks like in this vein yet.
00:08:09But of the things that I've seen, I think there's a couple of through lines that I tend
00:08:12to notice, right?
00:08:13The first is hands-free or screen-free experiences.
00:08:17There might be times where I need to interact with software, interact with objects, and I can't
00:08:21use my hands or, more importantly, my attention is diverted elsewhere.
00:08:24That might be something, you know, like a recipe app.
00:08:28Maybe I'm cooking and I need to just say, like, what's going on and have something else, have
00:08:32something happen.
00:08:34The other category is proactive outreach, right?
00:08:36Where you, the model needs to be able to tell you something or get your attention in a
00:08:39way that you might not be looking at, right?
00:08:41I think every developer has an endless amount of notifications and events happening in their
00:08:48software.
00:08:49But no developer in their right mind would sort of say, I should show all of these logs,
00:08:52nor, you know, would they say, I should speak all of these logs.
00:08:55But we can start to conceive of voice as this, like, upper level in this escalatory path of,
00:09:00like, okay, maybe you animate something and then maybe you pop something up and then if that
00:09:03doesn't work, maybe you talk to the user to get their attention.
00:09:08And underlying both of these categories and I think this, you know, this whole presentation
00:09:11is this broader theme of accessibility.
00:09:13On a personal note, I know, like, multiple developers who over the course of their careers
00:09:19lost mobility in their hands, lost dexterity in their fingers, and for many of them they
00:09:24thought their career as a programmer was more or less over.
00:09:28And then came large language models, right?
00:09:30Then came coding agents and voice agents and now they generate orders of magnitude more
00:09:34code than they, like, previously did, you know, on a given day or month.
00:09:39And so I think there's a lot that we can unlock here for the broader world as well.
00:09:45To go back to, like I said, you know, I think, like, when it comes to these three modalities,
00:09:48you can mix and match them and we already have some, you know, rudimentary ways that we're
00:09:52seeing this.
00:09:53I think there's things like, you know, all of these pieces for in-car assistance exist,
00:09:57though nothing has quite, like, combined them into this seamless way.
00:10:00You can talk to, like, the CarPlay dashboard, you can tell it, hey, go play some Spotify
00:10:05music for me.
00:10:06And then it can come back and tell you, Google Maps can come back and tell you, hey, like,
00:10:09you know, there's traffic on this route, we're going to reroute you.
00:10:12But, like, we can now start to think about what does it mean to combine that into, like,
00:10:15a single voice agent across multiple modes.
00:10:18Similarly, you know, there's a lot of experimentation in the game space with multimodality.
00:10:23You can think about a real-life character where you're talking to it to, you know, mine information
00:10:27about the game, about the world, you can talk to it to execute actions on your behalf.
00:10:32And then it can react to, like, world events, right, that are happening and then give that
00:10:35information to you rather than just, like, a simple notification.
00:10:39And I think the question, you know, behind the question here, right, is, like I mentioned,
00:10:44we've had all these things for a while, people have been prototyping them for a while.
00:10:47Why focus on them now?
00:10:48Why think about building with them now?
00:10:50And I think that brings me to a little bit of context here, right?
00:10:54As hopefully most of you know, traditionally, voice agents are built in this chained model,
00:10:59right?
00:11:00You talk, you transcribe, you send that to a language model, it calls tools, hopefully it
00:11:04doesn't take too long to respond.
00:11:05It then generates text output, you make that into audio, and then you play that back to the
00:11:11user.
00:11:12So a long time ago, OpenAI, you know, decided on a different approach, right?
00:11:15The real-time model family does not do any transcription behind the scenes.
00:11:20It is trained on native audio as tokens, so you send audio in and you get audio back out.
00:11:27And the industry, I think, in general, has been, you know, trending more in this direction,
00:11:31and not just making it native audio, but even just letting go of the turn-based abstraction
00:11:35that we've had, right?
00:11:36And so making it that it's just continuous streaming audio in and out.
00:11:42And the reason that OpenAI did this was, you know, it turns out there's a lot of stuff
00:11:46that you lose when you transcribe speech and when you transcribe audio, right?
00:11:50There's the old saying that when humans communicate face-to-face, 55% of the information is in
00:11:55body language, another 38% is in your tone of voice, and the last, like, 7% is the actual
00:12:00words you are saying.
00:12:01And so when you transcribe, you lose tone and cadence and emotional, you know, impact.
00:12:06You lose, like, whether they're trying to interrupt you, you lose background noise, all of this
00:12:09stuff which is really important context for the model to understand.
00:12:11A much more quantitative reason to do it is that, you know, the first two voice modes in
00:12:18ChatGPT were built with this chained approach, and as you can see, had, you know, significantly
00:12:23higher latency than using the native approach with advanced voice mode.
00:12:29And that brings me to GPT real-time 2.
00:12:32And this is going to be the one part of the talk where, you know, I make my shameless plug.
00:12:36Real-time 2 is the latest model in the real-time family.
00:12:39We released it a couple of months ago.
00:12:42And the really cool thing about this model is that it brings reasoning to the audio medium.
00:12:48And so much like our text models, it can now think before it speaks.
00:12:51I'm sure many of us have seen some demos of voice models saying things that are a little
00:12:55bit less than intelligent.
00:12:57And so you can now, you know, try to ensure that you give it more reasoning budget to come
00:13:02up with a good answer.
00:13:03Part of why that's also useful is that we introduced tool calling a little while ago.
00:13:07And so the model, in addition to thinking, it can also delegate parallel tool calls.
00:13:11You can start to bring these together, though, of course, that adds latency, right?
00:13:15And, you know, that's why we also added preambles.
00:13:18Preambles are a way that you can prompt the model to give the user a heads up if it's going
00:13:23to be thinking or if it's going to be calling tools.
00:13:25You know, if you think about the scenario of a travel agent, right, if I called a travel agent
00:13:30on the phone, you would want the travel agent to say, "Hey, like, I'm going to go check flight
00:13:35prices, right?
00:13:36Give me a couple seconds to do that."
00:13:38And now with an AI travel agent and preambles, you can actually have it communicate that to
00:13:43the user while it's performing actions in the background.
00:13:46There's a few other things here, right?
00:13:48It's got longer context, better domain understanding, more natural voices, and it's much more steerable.
00:13:53There's some really cool features that it can do when it comes to, like, wake words and just
00:13:58waiting for you to tell it.
00:14:00You know, you can give it a name.
00:14:01You can say, like, you know, "Hey, Marin, do you want to say hi to the room?"
00:14:05And if you've, like, prompted that into the model, then it'll, you know, go ahead and respond
00:14:10to you, right?
00:14:11And, of course, you know, obligatory benchmark slide.
00:14:14It does pretty well on the latest audio benchmarks, too.
00:14:18So TLDR is a pretty good model.
00:14:23But I think the kind of final thing that, you know, I want to leave you with here is when building voice
00:14:30agents, not to start with the question of, like, what kind of voice agent am I trying to build, right?
00:14:37I think the thing I want to leave you with is start with the question of, like, what is the role of voice and audio in this interaction?
00:14:45And then how do I move forward from there, right?
00:14:48And often when I ask that question, it leads to a bunch more questions after that.
00:14:52Things like, what can the model perceive?
00:14:54What context does it have, right?
00:14:56What tools are available to it?
00:14:58And which of those tools should it be, you know, executing safely and correctly?
00:15:03Should it communicate now?
00:15:05Should it wait?
00:15:06You know, how should it communicate?
00:15:08Should it be sending visual notifications or using audio?
00:15:12And so taken together, I hope everybody in here can start to build some much richer experiences with voice.
00:15:18Because, like others have said, I do believe that AGI will be spoken, not typed.
00:15:25Thank you very much.
00:15:26I'll be at the OpenAI booth for any Q&A after.
00:15:29Yeah, have a good event.
00:15:30Thank you.

핵심 요약

GPT real-time 2 brings reasoning, native audio processing, and parallel tool calling to voice agents, shifting interaction modes beyond speech-to-speech into speech-to-action and event-to-speech.

하이라이트

  • Speech is not the only response mode for voice models, allowing interaction through actions and events rather than just talking back.

  • Traditional voice agents rely on a chained model of transcription and text-to-speech, whereas native real-time audio models process audio as tokens directly without transcription loss.

  • GPT real-time 2 introduces reasoning to the audio medium, enabling models to think before speaking and delegate parallel tool calls with configured preambles.

  • Voice models can drive existing web applications by converting API endpoints and React hooks into tools controlled via voice.

  • Face-to-face human communication relies heavily on body language and tone, meaning transcription loses critical context like emotional impact, cadence, and interruptions.

타임라인

Misconceptions and Emerging Modalities of Voice Agents

  • Voice agents do not always have to talk back to users.
  • Three distinct interaction modes include speech-to-speech, speech-to-action, and event-to-speech.
  • These modes are remixable and combine to form advanced product experiences.

Voice models are reaching an intelligence level that opens up design patterns beyond traditional conversational loops. Historical precedents like movie phone hotlines and GPS navigation demonstrate that these categories are long-standing, but modern architectures allow significantly more complex execution. Speech-to-speech handles live practice, concierge customer support, and real-time translation.

Speech-to-Action and Event-to-Speech Paradigms

  • Speech-to-action enables form filling, creative tools, and complete computer use through voice commands.
  • Developers can easily convert existing API endpoints and React hooks into tools for voice models.
  • Event-to-speech handles hands-free, screen-free experiences and proactive user outreach.

Form filling eliminates manual data entry for complex government documents by letting users talk through fields. Creative tools allow users to compose music or paint using natural language descriptions. Developers can expose backend software actions directly to voice agents while retaining visual feedback mechanisms like cursor animations and state changes instead of mandatory verbal responses.

Native Audio Architecture and GPT Real-Time 2

  • Native real-time audio models process audio tokens directly, avoiding the information loss and latency of traditional transcription chains.
  • GPT real-time 2 introduces audio reasoning, allowing models to think before speaking and manage parallel tool calls.
  • Preambles let models communicate status updates to users while performing background actions.

Traditional voice architectures suffer from high latency and lose tone, emotion, and cadence during transcription. OpenAI's real-time family eliminates transcription by training directly on native audio streams. The addition of reasoning budgets, longer context windows, steerability, and preambles makes voice agents vastly more reliable for complex multi-step tasks.

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기