Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

AAI Engineer
컴퓨터/소프트웨어가전제품/카메라

스크립트

00:00:00with like no noise you know in a sealed environment they're happening on the go on the train while
00:00:04you're on a walk and that's the type of experiences that we want to facilitate with the speech-to-speech
00:00:08model lastly another thing we're really excited about is not only voice out but also multimodal
00:00:15out we believe that you know the true AGI conversational frontier will also require
00:00:20visual presence so we were excited to launch our first kind of pilot demo with Citi in cloud next
00:00:26which supports customized real-time avatars and you can kind of like personalize anything from a
00:00:33hyper-realistic human to a cartoon appearance and everything in the middle and powered by the same
00:00:38speech-to-speech model we've been showing it allows to have you know low latency multilingual lip-syncing
00:00:43and a really kind of like uh ongoing flowy conversation that has visual presence so let me
00:00:48show you a demo that also brings our venn diagram together and what we're excited about all this is
00:00:54reminding me of my daughter's college fund how are we tracking on that
00:01:00you're tracking well jackson i've also identified a new opportunity that may get you there even sooner
00:01:06oh and i can see lisa just joined you hi lisa she must be so excited about her college acceptance
00:01:12congratulations it's wonderful to see your savings goals coming to life
00:01:16excellent that's good news i always joke that the user's audio in spanish is worse than the audio
00:01:34model speaking back but um this is kind of just to show how like our venn diagram of combining you know
00:01:40multi-modality in and out you know tool calling to like pull up the relevant examples from the user
00:01:45and also conversational fluidity with itnn are starting slowly to come together in these types
00:01:50of demos that we're excited to keep pushing the frontier of um so with this parting thought i guess last
00:01:57kind of thought that we have for you is that we believe that agi will not be typed that it will be
00:02:01spoken um and for it to be spoken there's a lot of things that need to work together in a single
00:02:06promptable versatile model that allows a user to switch between all the sorts of conversation modes
00:02:12that we're looking at right from translation to taking action to brainstorming to rambling and we
00:02:17truly believe in the power of these speech-to-speech models to achieve that like seamless switching um
00:02:22so we're excited to push the frontier on that so if you're excited or want to learn more
00:02:26please come talk to us and thank you so much for coming
00:02:46you
00:02:56Thank you.

핵심 요약

Achieving advanced conversational AGI requires a single promptable speech-to-speech model supporting low-latency multilingual lip-syncing, multimodal visual output, and real-time tool calling.

하이라이트

  • Speech-to-speech models facilitate on-the-go conversational experiences in noisy real-world environments like trains or walks.

  • A pilot demo with Citi in cloud next supports customized real-time avatars ranging from hyper-realistic humans to cartoon appearances.

  • The speech-to-speech model enables low-latency multilingual lip-syncing alongside ongoing conversational fluidity.

  • The conversational frontier requires multimodal output capabilities, incorporating visual presence alongside voice generation.

  • AGI will be spoken rather than typed, requiring a single promptable model that handles translation, action-taking, brainstorming, and rambling.

타임라인

Multimodal Speech-to-Speech Integration and Avatars

  • Real-world interactions happen on the go, requiring robust speech-to-speech models.
  • Conversational AI frontiers demand multimodal output, including custom real-time visual avatars.
  • A pilot demonstration with Citi in cloud next features low-latency multilingual lip-syncing and fluid conversational dynamics.

User interactions occur in diverse, noisy environments like public transit or walks, forming the primary use case for speech-to-speech systems. Expanding beyond audio output, the system incorporates visual presence through customizable real-time avatars that range from cartoon appearances to hyper-realistic humans. A demonstration combining input-output multimodality, tool calling, and conversational fluidity illustrates these capabilities in a live dialogue involving savings tracking and visual participant recognition.

Future Frontier of Spoken AGI

  • Tool calling allows models to retrieve relevant user examples dynamically.
  • Advanced conversational systems combine multimodality, tool integration, and conversational fluidity.
  • Spoken AGI relies on a versatile model capable of switching between translation, action-taking, brainstorming, and rambling.

The integration of multi-modality, tool calling, and text normalization frameworks enables cohesive interaction workflows. The overarching trajectory of artificial general intelligence centers on speech rather than text. A versatile, promptable model must seamlessly handle rapid shifts across various conversational modes without friction.

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기