Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data
AAI Engineer
컴퓨터/소프트웨어경제 뉴스AI/미래기술
스크립트
00:00:00Hi everybody, welcome. For those that came to actually see me, thank you. For those that are
00:00:21just hanging out here, also thank you. I'm gonna try to entertain you guys, teach you something new.
00:00:26Again, my name is Rafael, I work at Bright Data, I've been with Bright Data for over eight years and in Bright Data we are collecting data and we're trying to innovate.
00:00:38So today I want to talk about video discovery for agentic world models training and we're going to explain what that is in a little bit.
00:00:48So what does it take to today to build a system that can see, understand and act?
00:00:56Right? I mean, we figured out what an AI is, we're constantly improving it, like, you know, that's kind of handled,
00:01:04but now the hard part is what data do you provide to it because everybody knows AI without data is just a box.
00:01:13Right? So let me take you a little bit back in history, right? In 2022, Google taught a robot by images,
00:01:25real world actions, right? Then going forward, there was a leap. In 2023, we had an AI actually control multiple robots.
00:01:37In 2024, the open source came out. Now this is where the game changed. Once we had an open source, all the labs had access to AI and they started already putting AI into robots, right?
00:01:51And so now these models are already driving humanoid robots. I'm seeing some, of course, most here are remote controlled.
00:02:00But I mean, in San Francisco, you got Wimeo driving without any drivers. How cool is that?
00:02:07First time I got in the car, I thought I was going to die, but it was actually very nice.
00:02:12Right? Now, how do you think it learned how to drive? By actually teaching itself, by learning on dashboard cameras, right?
00:02:22Every car has a dashboard. So if you provide that to an AI, it can figure out how to actually operate a machine.
00:02:31So as I said, the AI is no longer the hard part, the data is, and this is what I want to talk about, right?
00:02:39So for chat LLMs, there's trillions of words and texts, right? So we have so much text data to train the LLMs.
00:02:49For image generation, we got billions of images, labeled images, so that's also a huge database.
00:02:56But for robotics, there's only about a million videos. It's a very small, limited data set of robots doing things, right?
00:03:09And so the idea here is that, of course, a lot of companies out there, what they do is they pay people to record actions, right?
00:03:20So for example, hey, record me how you open a door, or record to me how you're sitting on the chair.
00:03:28But then the problem becomes is that when somebody is told to do something, they do not do it naturally.
00:03:34So I call it instructed, right? If you're told to record how you open the door, and you're doing it for somebody,
00:03:40it's not going to be the same as if you just walk into the house. The movement is totally different.
00:03:45So is that data good for robotic training? In my assumption, no. It's biased data, and it doesn't deliver the same results.
00:03:53So obviously, the robot is not going to be as effective as if it's intuitive, right?
00:03:59So I just threw in some examples on the slide, right?
00:04:06The build-by-hand data is not the same as the data that is kind of on the side, right?
00:04:12And also, the fact is that people behave very differently when they're on the camera.
00:04:16Some of you know, some people are camera shy, so as soon as you put a camera on them, they change.
00:04:23But in the natural world, it's a totally different thing.
00:04:27So this is what we are trying to solve in bright data, right?
00:04:35Again, so why is the usual data sources are not enough, right?
00:04:37Simulations. Virtual, it's cheap, but it's, you know, everybody's playing video games.
00:04:44The physics in video games are almost there, but they're not good enough to train a robot on it, right?
00:04:50A hand control, the person controlling a robot, it's doable, but how many hours a day can you record a robot?
00:04:57How many people you need to actually get huge-scale data, right?
00:05:01I mean, you can record maybe eight hours a day, every day.
00:05:04How many hours are you going to get in a year?
00:05:07It's not really scalable.
00:05:09And of course, there's already pre-built data sets, but as I said before, they're small.
00:05:13There's only about a million videos at this point.
00:05:15So what is the alternative?
00:05:17The alternative is the web.
00:05:20Let's just think about YouTube.
00:05:22How many videos are there on YouTube?
00:05:25I don't know, five billion videos?
00:05:28How many actions are in those videos that a robot can replicate?
00:05:34I mean, let's think about it.
00:05:35How many do you think videos are there of a person opening the door, first person view?
00:05:42Millions of hours.
00:05:43You'll be surprised.
00:05:45So every day, the web video shows gravity, motions, right?
00:05:53So cause and effect, obviously accidents, the biggest cause and effect, right?
00:06:00How people handle objects, and billions of hours.
00:06:04Great training material for robots, but what is the problem?
00:06:08The problem is that there's a lot of noise, right?
00:06:12Oh, and for example, right?
00:06:14So one of the AI models that Meta trained, they gave it about a million hours of real world videos.
00:06:21And then all it took is 62 hours of real robotics data to actually control a real robot.
00:06:29So if you think about it, the data was collected publicly, right?
00:06:33So it trained the robot on just random things, and then in 62 robotic hours, it was already moving.
00:06:41So no simulations needed, autonomous robotics delivered fast.
00:06:52But the videos don't show how the robots move, right?
00:06:54So you might think, well, listen, how useful is it a person pouring water into a cup for a robot?
00:07:01So there is methods actually out there how an AI can distinguish and actually learn from the actions that are in the video.
00:07:09If you take two frames, frame by frame, and the AI measures the movement difference, right?
00:07:15So you have an image A and image 2, and there's a small movement that is changing between the two images.
00:07:21An AI can actually measure that change.
00:07:24So if you keep doing that through the whole video, an AI can actually figure out angles, distance,
00:07:30and everything that it needs to train and robot, right?
00:07:34So we don't necessarily need the sensor's data, but of course, the video itself that you download
00:07:41from YouTube or the video that you extract the data from doesn't necessarily get you 100% there.
00:07:48You do need to have some processing on top of that, but it's available.
00:07:53It's out there, and all you need to do is process it.
00:07:58So a few more examples, right?
00:07:59So for example, NVIDIA, when they train in the Cosmos robot, they're throwing out about 96% of the video.
00:08:07So what does that mean?
00:08:09They download a million hours of video,
00:08:12and the only part is 4% is actually useful.
00:08:15Everything else gets thrown out.
00:08:16That's wasted compute.
00:08:18That's wasted bandwidth.
00:08:19That's wasted storage.
00:08:21That's just a lot of wasted money.
00:08:23Okay, stable video diffusion is throwing out 74% of the videos that they're downloading.
00:08:30So in Bright Data, we are trying to solve that problem, and we're trying to take it a step forward.
00:08:37So what we propose is search first, collect second.
00:08:41What we do is we are actually indexing videos, and we are allowing you to search for specific actions.
00:08:50Since we're talking about a person opening a door, let's keep using that example.
00:08:55What you can do on our platform is actually input detailed information, right?
00:09:01A query, person washing dishes, person folding a t-shirt, and so on.
00:09:06And we will provide the snippets from the videos of these actions.
00:09:13So what does that mean?
00:09:16That means that there's less waste of data collection.
00:09:22There's less waste of storage and bandwidth, right?
00:09:25So a little more on how it works.
00:09:27Obviously, you define what you're looking for.
00:09:32We search through billions and billions of pre-indexed videos.
00:09:36Not by keywords, but by actions.
00:09:39And we provide you with ready-to-use clips.
00:09:43And obviously, they're already prepared for your training.
00:09:47So all you gotta do is process them a little bit for maybe motion, distance sensors, and so on.
00:09:52And it's ready to be ingested into a robot.
00:09:57I have a small example of a video of how it actually works on our platform.
00:10:02So let me just hit play here, if I can find it.
00:10:07So here we see a person washing dishes by hand in sink, removing grease from the plate,
00:10:12using sponge and soap, including rinse and placing dishes into drying rack, close-up hand interactions.
00:10:19So right now, what is happening is it's searching through billions of videos.
00:10:24Now, on this video, it's a bit old.
00:10:26We only got about 100 million index there, but now we're up to 1.1 billion videos.
00:10:31And it will literally, like, this is a little sped up, so I didn't want to waste you guys' time.
00:10:35But it will provide to you clips from videos of people washing dishes.
00:10:41And you can see right here all the videos that we find.
00:10:47Now, some of these videos might not have to do anything with washing dishes.
00:10:50They might have a cleaning kitchen.
00:10:52They might have to do something else.
00:10:54And here's another example.
00:10:55For example, human folding different types of clothes.
00:10:59Right?
00:10:59So, a little more description.
00:11:01The more description you give, the more you actually get back.
00:11:06So again, a little sped up.
00:11:08Let's see.
00:11:14I want you to understand, it could be anything.
00:11:16A person putting on makeup.
00:11:17Let's say you have a brand and you want to find videos where your brand is being found.
00:11:21All that can be also provided.
00:11:23So, you see people folding clothes.
00:11:26So, now you can train a robot on how to fold clothes.
00:11:34Of course, everything is available via API.
00:11:36We don't expect people to actually do anything manually.
00:11:39Right?
00:11:39So, you can trigger it by API.
00:11:41You get a return snippet of the video, the URL.
00:11:46We give you the link to the original video if you want to watch it.
00:11:49But the main thing here is that we can give you the snippets of the video to train your robotics.
00:11:57Now, I'm not sure if any of you are training any robots.
00:12:01So, I'm going to give you another example of how useful it could be.
00:12:04Let's say you have a brand.
00:12:06And you want to know where your brand is being demoed on videos.
00:12:13The video titles doesn't matter because it doesn't actually, it's not about your product.
00:12:19It's just a person doing some podcast, recording something.
00:12:22But, you see a woman putting on makeup.
00:12:25And you see on the table that this is your brand of a makeup.
00:12:30So, now all of a sudden, you can find all videos for your brand where it's being used.
00:12:34Whatever it is.
00:12:35Right?
00:12:35By searching the name of the video, you won't find it.
00:12:38By video indexing, you can actually find specific things that you are interested in.
00:12:45Apple falling from the tree.
00:12:47And so on.
00:12:48So, what comes back?
00:12:50We provide to you this time stamp.
00:12:52We provide to you the score matching.
00:12:53How close it is to your query.
00:12:55And of course, the frame count.
00:12:57How many frames are about what you're looking for.
00:13:03So, what does this change?
00:13:06This changes a lot.
00:13:07Okay, less waste.
00:13:09You don't need to download millions of hours of videos.
00:13:12You can actually get exactly the specific actions
00:13:16that you are interested in.
00:13:20Again, it's very crucial for robotic training.
00:13:24Because noise creates problems, hallucinations.
00:13:28Right?
00:13:29This removes the noise.
00:13:34What is it useful for?
00:13:35Not only robots, but obviously self-driving.
00:13:38For example, Wimeo.
00:13:39Right?
00:13:40Trained on dash video cameras.
00:13:43And there's millions, hundreds of millions of hours on YouTube.
00:13:46Of dash cams.
00:13:47How many of you like to watch accidents?
00:13:49Dash cam cam accidents.
00:13:52I'm sure everybody watched some.
00:13:53Crazy driving in Russia.
00:13:55Right?
00:13:56So, this is what we're talking about.
00:13:58Right?
00:13:58The videos are out there.
00:14:01Okay?
00:14:01A person running a red light.
00:14:02A person stopping on a red light.
00:14:04Making a left turn.
00:14:05Making a right turn.
00:14:06And so on.
00:14:06All of that can be a very useful data for training.
00:14:11And any other world models.
00:14:13Physics.
00:14:14Right?
00:14:14Robots need to understand physics.
00:14:16What is the gravity?
00:14:18How to sit?
00:14:18How to walk?
00:14:19How to move?
00:14:21Again, of course, you can pre-record it.
00:14:25A lot of times right now, I speak to some of the people.
00:14:27And they have millions of people recording videos for them.
00:14:30Hey, can you guys record me how you open doors?
00:14:33That's crazy.
00:14:34Why do you need a million people doing that when everything is available online?
00:14:39Online is one of the hugest databases out there.
00:14:43It's just people don't really, like, it's not that easy to use it, right?
00:14:46I mean, right now the only search available to us is by the keyword of the name of the video.
00:14:53And of course, one source, the whole public video web in one place.
00:14:59I mean, we got YouTube, we got Vimeo, we got so many different video providers out there.
00:15:06Billions and billions of hours of training data just there waiting for you to grab it,
00:15:13collect it, and process it.
00:15:14So basically, this is a new product that we're doing at Bright Data, right?
00:15:18Indexing videos so that you guys can find specific actions.
00:15:23Anything you can think about, you know, again, it could be useful for brands.
00:15:28As well, it's not only for training data.
00:15:31Anything you can think of could be useful.
00:15:35Maybe you are a gamer and you want to know how do you beat this level?
00:15:40But there's no exactly video like that.
00:15:42So you can search, okay, people beating the level on the video game, right?
00:15:47Anything, the world is yours.
00:15:52So I'm going to wrap it up.
00:15:54I don't know if you have any questions, but if you want to connect with me, talk about it more,
00:15:58feel free to add me on LinkedIn.
00:16:01And yeah, I'm going to thank you guys for your attention.
00:16:06And I'm at a Bright Data booth if you want to talk more about it.
00:16:10And thanks for coming.
00:16:17I'll see you next time.
커뮤니티 글
아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!
이 영상에 대해 글쓰기