Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
AAI Engineer
Computing/SoftwareConsumer Electronics
Transcript
00:00:00with like no noise you know in a sealed environment they're happening on the go on the train while
00:00:04you're on a walk and that's the type of experiences that we want to facilitate with the speech-to-speech
00:00:08model lastly another thing we're really excited about is not only voice out but also multimodal
00:00:15out we believe that you know the true AGI conversational frontier will also require
00:00:20visual presence so we were excited to launch our first kind of pilot demo with Citi in cloud next
00:00:26which supports customized real-time avatars and you can kind of like personalize anything from a
00:00:33hyper-realistic human to a cartoon appearance and everything in the middle and powered by the same
00:00:38speech-to-speech model we've been showing it allows to have you know low latency multilingual lip-syncing
00:00:43and a really kind of like uh ongoing flowy conversation that has visual presence so let me
00:00:48show you a demo that also brings our venn diagram together and what we're excited about all this is
00:00:54reminding me of my daughter's college fund how are we tracking on that
00:01:00you're tracking well jackson i've also identified a new opportunity that may get you there even sooner
00:01:06oh and i can see lisa just joined you hi lisa she must be so excited about her college acceptance
00:01:12congratulations it's wonderful to see your savings goals coming to life
00:01:16excellent that's good news i always joke that the user's audio in spanish is worse than the audio
00:01:34model speaking back but um this is kind of just to show how like our venn diagram of combining you know
00:01:40multi-modality in and out you know tool calling to like pull up the relevant examples from the user
00:01:45and also conversational fluidity with itnn are starting slowly to come together in these types
00:01:50of demos that we're excited to keep pushing the frontier of um so with this parting thought i guess last
00:01:57kind of thought that we have for you is that we believe that agi will not be typed that it will be
00:02:01spoken um and for it to be spoken there's a lot of things that need to work together in a single
00:02:06promptable versatile model that allows a user to switch between all the sorts of conversation modes
00:02:12that we're looking at right from translation to taking action to brainstorming to rambling and we
00:02:17truly believe in the power of these speech-to-speech models to achieve that like seamless switching um
00:02:22so we're excited to push the frontier on that so if you're excited or want to learn more
00:02:26please come talk to us and thank you so much for coming
00:02:46you
00:02:56Thank you.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video