스크립트
00:00:00So my name is Charlie, and I work on the developer experience team at OpenAI.
00:00:18And part of my job is talking to developers to understand and see what and how they're
00:00:25building with our models, whether that's text, image, or audio.
00:00:31And lately I've been thinking about a misconception that I have seen, or maybe it's just a misunderstanding.
00:00:39And it's the idea that voice agents have to talk back.
00:00:48And to some of you that might sound, you know, absurd.
00:00:51It's a voice agent.
00:00:52What do you mean it's not supposed to talk?
00:00:54But I think if there's one thing that you take away from this presentation, I would like
00:01:00it to be the idea that speech is not the only way that a voice model has to respond.
00:01:12And I think models are getting intelligent enough and capable enough that they're starting to
00:01:17open up some new modes of design.
00:01:21There's actually three kind of modes that I kind of see emerging these days, right?
00:01:27Speech to speech, speech to action, and event to speech.
00:01:32And there's a couple of things I think worth pointing out about these three categories.
00:01:37The first is that they're not new, right?
00:01:41I think as we've seen from previous talks, even just today, there's a long history of building
00:01:47these types of systems in and around voice.
00:01:50If you squint, you could make the argument that the movie phone hotline where you called
00:01:54in to get showtimes was an example of a speech-to-speech system.
00:01:58And I think you could pretty reasonably make the argument that GPS navigation in your car,
00:02:04which has existed since I was a kid, is an example of an event-to-speech system.
00:02:09So it's not that they are brand new, but I think it is that we are able to do some much more
00:02:14interesting things with them now that we're in this era, right?
00:02:18And I think the other thing I would mention here is that they're remixable, right?
00:02:24They're not meant to be mutually exclusive, and I think as we'll see in a little bit, the
00:02:27best products exist in a way that combines all of these modes.
00:02:32So speech-to-speech, everybody knows it, hopefully everybody loves it.
00:02:36The user talks, and then the model talks back.
00:02:40And I think there are a few examples that I can give for this type of use case, right?
00:02:45You've got things like live practice and coaching, especially around language learning, right?
00:02:50I think the ability to hear a lot of emphasis or emotion and give people that feedback is
00:02:55really powerful.
00:02:57I think you can have, you know, what I am sort of cheekily calling concierge experiences,
00:03:01which I think is just another way to say customer support plus-plus.
00:03:06The first voice tutorial that most people, you know, try to build when they have access
00:03:10to this technology is some sort of customer support chat bot for very good reasons.
00:03:14But I think with, you know, when we add a richness and a depth to the voice models, as has been
00:03:19happening in recent months and recent years, we can build something that is, like, much more
00:03:24enjoyable to use than, like, talking your way through a phone tree.
00:03:28And so, you know, I have the hope that, like, soon if not, like, you know, now we are capable
00:03:33of building support experiences with agents that actually feel much more enjoyable to talk
00:03:38to than, like, arguably, like, the median human support agent.
00:03:44And as we just saw, if you were here for the last talk, live translation, right?
00:03:47The models have gotten good enough and fast enough that we can just dynamically translate
00:03:51content on the fly with, like, little to no latency.
00:03:55It wouldn't shock me if at next year's keynote, you know, they live streamed it from the main
00:03:58stage but also dubbed it in real time across multiple languages.
00:04:05The second category is speech to action.
00:04:07These are talks and the model uses tools.
00:04:10And I think this is one of the most under-explored areas that we have.
00:04:15I actually almost titled this talk "Voice is the Next Capability Overhang" because I think
00:04:20there is just a vast, vast amount of stuff that we could be doing in this category that we are
00:04:25not currently doing.
00:04:27For example, there's a broad spectrum.
00:04:30I don't have, you know, there's way too many examples even fit on this slide.
00:04:33But three categories that I find particularly interesting.
00:04:37First is form filling, right?
00:04:38So much of the internet is just filling out forms.
00:04:43And there is, you know, today no reason why you shouldn't be able to just talk.
00:04:46And I would love it if instead of spending an hour filling out a government document, I
00:04:50could just talk for five minutes and it would get 90% of it for me and I would do a quick
00:04:55check, you know, just to make sure that everything looked good, right?
00:04:57That is a vastly superior experience than, like, having to type in every single name and address
00:05:02that I've lived in the last five years and, you know, all of my previous identities.
00:05:05And so I think that one is, though it may seem boring, you know, affects a significant GDP
00:05:11of the internet, right?
00:05:13The next category is creative tools, where I am privileged enough that I can speak the
00:05:17language of software.
00:05:18And so I can tell codecs, you know, here's exactly what I want you to build and I can articulate
00:05:23it in a way that I get much more leverage than sort of just, like, cludgily trying to iterate
00:05:28one thing at a time.
00:05:29But I can't do that when it comes to, you know, using, making music or painting.
00:05:33And so if I don't have the ability to articulate the exact aesthetic that I'm looking for and
00:05:40if I don't know how to use Photoshop or Ableton, I'm left in this state where, you know, my taste
00:05:45exceeds my capability.
00:05:46And so I'm really looking forward to integrating voice into creative tools so that I can just
00:05:51sort of cludgily go along and, you know, vibe create, vibe compose, vibe paint, and make
00:05:56something that's really beautiful to me.
00:05:59And I think the generalizable category here, right, then just starts to become computer use.
00:06:05And we've already seen some companies start to do this.
00:06:07You know, it raises the question of, like, look, if the models are just getting good enough
00:06:11to do everything on a computer that a human can do, like, why am I talking to an app?
00:06:17Why am I talking to a terminal?
00:06:18Why am I not just talking to the entire computer?
00:06:22And so I think that's sort of a really interesting way to start exploring.
00:06:25But if you're a developer today, right, oops, if you're a developer today, what does that
00:06:30mean for building your own software, right?
00:06:32And I think it is, like, much easier than you think to start adding audio as an intelligence
00:06:36layer to the -- intelligence layer to the apps that you already have.
00:06:40If you're building a modern web application, you already expose so much of it as like actions,
00:06:45as nouns and verbs, right?
00:06:46If you think about all the verbs that you have, you have API endpoints, you have, you know,
00:06:51React hooks.
00:06:52Each of those things can, like, pretty relatively easily be converted into a tool that you expose
00:06:56to a model.
00:06:57And then you can give the user the ability to just drive your existing software with their
00:07:01voice, right?
00:07:02And yes, you still need guardrails.
00:07:03You still need safety checks.
00:07:04Like, many of the talks today are going to talk about securing and, you know, productizing
00:07:08this.
00:07:09And for this, I just want you to think about, you know, what would it mean to take your existing
00:07:12software and just talk to it?
00:07:17And to go back to that misconception, right, I think there are a lot of, you know, like,
00:07:20if you're talking to the software, maybe you can talk back, but we've been developing other
00:07:24ways of communicating with the user for decades, right?
00:07:26We know these things.
00:07:27We know we can show notifications and popups.
00:07:29We can change state, like the color of a button or a drop shadow.
00:07:32We can highlight text.
00:07:34If you've used computer use in the Codex app, you know, there's this amazing, like, little
00:07:37ghost cursor animation that goes around and clicks things for you.
00:07:40So we don't have to use words to actually tell the user what is happening on screen with
00:07:44their software.
00:07:48And then the last bucket here is event-to-speech, right?
00:07:51The model receives an event and talks to the user.
00:07:54And sort of the counterpoint from speech-to-action, I think this one is still very, very exploratory,
00:07:58right?
00:07:59You know, if you saw Quinn's talk, I think there's a lot of space here of, like, things
00:08:02we can do.
00:08:04And to me, we haven't quite seen what AI native really looks like in this vein yet.
00:08:09But of the things that I've seen, I think there's a couple of through lines that I tend
00:08:12to notice, right?
00:08:13The first is hands-free or screen-free experiences.
00:08:17There might be times where I need to interact with software, interact with objects, and I can't
00:08:21use my hands or, more importantly, my attention is diverted elsewhere.
00:08:24That might be something, you know, like a recipe app.
00:08:28Maybe I'm cooking and I need to just say, like, what's going on and have something else, have
00:08:32something happen.
00:08:34The other category is proactive outreach, right?
00:08:36Where you, the model needs to be able to tell you something or get your attention in a
00:08:39way that you might not be looking at, right?
00:08:41I think every developer has an endless amount of notifications and events happening in their
00:08:48software.
00:08:49But no developer in their right mind would sort of say, I should show all of these logs,
00:08:52nor, you know, would they say, I should speak all of these logs.
00:08:55But we can start to conceive of voice as this, like, upper level in this escalatory path of,
00:09:00like, okay, maybe you animate something and then maybe you pop something up and then if that
00:09:03doesn't work, maybe you talk to the user to get their attention.
00:09:08And underlying both of these categories and I think this, you know, this whole presentation
00:09:11is this broader theme of accessibility.
00:09:13On a personal note, I know, like, multiple developers who over the course of their careers
00:09:19lost mobility in their hands, lost dexterity in their fingers, and for many of them they
00:09:24thought their career as a programmer was more or less over.
00:09:28And then came large language models, right?
00:09:30Then came coding agents and voice agents and now they generate orders of magnitude more
00:09:34code than they, like, previously did, you know, on a given day or month.
00:09:39And so I think there's a lot that we can unlock here for the broader world as well.
00:09:45To go back to, like I said, you know, I think, like, when it comes to these three modalities,
00:09:48you can mix and match them and we already have some, you know, rudimentary ways that we're
00:09:52seeing this.
00:09:53I think there's things like, you know, all of these pieces for in-car assistance exist,
00:09:57though nothing has quite, like, combined them into this seamless way.
00:10:00You can talk to, like, the CarPlay dashboard, you can tell it, hey, go play some Spotify
00:10:05music for me.
00:10:06And then it can come back and tell you, Google Maps can come back and tell you, hey, like,
00:10:09you know, there's traffic on this route, we're going to reroute you.
00:10:12But, like, we can now start to think about what does it mean to combine that into, like,
00:10:15a single voice agent across multiple modes.
00:10:18Similarly, you know, there's a lot of experimentation in the game space with multimodality.
00:10:23You can think about a real-life character where you're talking to it to, you know, mine information
00:10:27about the game, about the world, you can talk to it to execute actions on your behalf.
00:10:32And then it can react to, like, world events, right, that are happening and then give that
00:10:35information to you rather than just, like, a simple notification.
00:10:39And I think the question, you know, behind the question here, right, is, like I mentioned,
00:10:44we've had all these things for a while, people have been prototyping them for a while.
00:10:47Why focus on them now?
00:10:48Why think about building with them now?
00:10:50And I think that brings me to a little bit of context here, right?
00:10:54As hopefully most of you know, traditionally, voice agents are built in this chained model,
00:10:59right?
00:11:00You talk, you transcribe, you send that to a language model, it calls tools, hopefully it
00:11:04doesn't take too long to respond.
00:11:05It then generates text output, you make that into audio, and then you play that back to the
00:11:11user.
00:11:12So a long time ago, OpenAI, you know, decided on a different approach, right?
00:11:15The real-time model family does not do any transcription behind the scenes.
00:11:20It is trained on native audio as tokens, so you send audio in and you get audio back out.
00:11:27And the industry, I think, in general, has been, you know, trending more in this direction,
00:11:31and not just making it native audio, but even just letting go of the turn-based abstraction
00:11:35that we've had, right?
00:11:36And so making it that it's just continuous streaming audio in and out.
00:11:42And the reason that OpenAI did this was, you know, it turns out there's a lot of stuff
00:11:46that you lose when you transcribe speech and when you transcribe audio, right?
00:11:50There's the old saying that when humans communicate face-to-face, 55% of the information is in
00:11:55body language, another 38% is in your tone of voice, and the last, like, 7% is the actual
00:12:00words you are saying.
00:12:01And so when you transcribe, you lose tone and cadence and emotional, you know, impact.
00:12:06You lose, like, whether they're trying to interrupt you, you lose background noise, all of this
00:12:09stuff which is really important context for the model to understand.
00:12:11A much more quantitative reason to do it is that, you know, the first two voice modes in
00:12:18ChatGPT were built with this chained approach, and as you can see, had, you know, significantly
00:12:23higher latency than using the native approach with advanced voice mode.
00:12:29And that brings me to GPT real-time 2.
00:12:32And this is going to be the one part of the talk where, you know, I make my shameless plug.
00:12:36Real-time 2 is the latest model in the real-time family.
00:12:39We released it a couple of months ago.
00:12:42And the really cool thing about this model is that it brings reasoning to the audio medium.
00:12:48And so much like our text models, it can now think before it speaks.
00:12:51I'm sure many of us have seen some demos of voice models saying things that are a little
00:12:55bit less than intelligent.
00:12:57And so you can now, you know, try to ensure that you give it more reasoning budget to come
00:13:02up with a good answer.
00:13:03Part of why that's also useful is that we introduced tool calling a little while ago.
00:13:07And so the model, in addition to thinking, it can also delegate parallel tool calls.
00:13:11You can start to bring these together, though, of course, that adds latency, right?
00:13:15And, you know, that's why we also added preambles.
00:13:18Preambles are a way that you can prompt the model to give the user a heads up if it's going
00:13:23to be thinking or if it's going to be calling tools.
00:13:25You know, if you think about the scenario of a travel agent, right, if I called a travel agent
00:13:30on the phone, you would want the travel agent to say, "Hey, like, I'm going to go check flight
00:13:35prices, right?
00:13:36Give me a couple seconds to do that."
00:13:38And now with an AI travel agent and preambles, you can actually have it communicate that to
00:13:43the user while it's performing actions in the background.
00:13:46There's a few other things here, right?
00:13:48It's got longer context, better domain understanding, more natural voices, and it's much more steerable.
00:13:53There's some really cool features that it can do when it comes to, like, wake words and just
00:13:58waiting for you to tell it.
00:14:00You know, you can give it a name.
00:14:01You can say, like, you know, "Hey, Marin, do you want to say hi to the room?"
00:14:05And if you've, like, prompted that into the model, then it'll, you know, go ahead and respond
00:14:10to you, right?
00:14:11And, of course, you know, obligatory benchmark slide.
00:14:14It does pretty well on the latest audio benchmarks, too.
00:14:18So TLDR is a pretty good model.
00:14:23But I think the kind of final thing that, you know, I want to leave you with here is when building voice
00:14:30agents, not to start with the question of, like, what kind of voice agent am I trying to build, right?
00:14:37I think the thing I want to leave you with is start with the question of, like, what is the role of voice and audio in this interaction?
00:14:45And then how do I move forward from there, right?
00:14:48And often when I ask that question, it leads to a bunch more questions after that.
00:14:52Things like, what can the model perceive?
00:14:54What context does it have, right?
00:14:56What tools are available to it?
00:14:58And which of those tools should it be, you know, executing safely and correctly?
00:15:03Should it communicate now?
00:15:05Should it wait?
00:15:06You know, how should it communicate?
00:15:08Should it be sending visual notifications or using audio?
00:15:12And so taken together, I hope everybody in here can start to build some much richer experiences with voice.
00:15:18Because, like others have said, I do believe that AGI will be spoken, not typed.
00:15:25Thank you very much.
00:15:26I'll be at the OpenAI booth for any Q&A after.
00:15:29Yeah, have a good event.
00:15:30Thank you.
커뮤니티 글
아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!
이 영상에 대해 글쓰기