Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data

AAI Engineer
Computing/SoftwareBusiness NewsInternet Technology

Transcript

00:00:00Hi everybody, welcome. For those that came to actually see me, thank you. For those that are
00:00:21just hanging out here, also thank you. I'm gonna try to entertain you guys, teach you something new.
00:00:26Again, my name is Rafael, I work at Bright Data, I've been with Bright Data for over eight years and in Bright Data we are collecting data and we're trying to innovate.
00:00:38So today I want to talk about video discovery for agentic world models training and we're going to explain what that is in a little bit.
00:00:48So what does it take to today to build a system that can see, understand and act?
00:00:56Right? I mean, we figured out what an AI is, we're constantly improving it, like, you know, that's kind of handled,
00:01:04but now the hard part is what data do you provide to it because everybody knows AI without data is just a box.
00:01:13Right? So let me take you a little bit back in history, right? In 2022, Google taught a robot by images,
00:01:25real world actions, right? Then going forward, there was a leap. In 2023, we had an AI actually control multiple robots.
00:01:37In 2024, the open source came out. Now this is where the game changed. Once we had an open source, all the labs had access to AI and they started already putting AI into robots, right?
00:01:51And so now these models are already driving humanoid robots. I'm seeing some, of course, most here are remote controlled.
00:02:00But I mean, in San Francisco, you got Wimeo driving without any drivers. How cool is that?
00:02:07First time I got in the car, I thought I was going to die, but it was actually very nice.
00:02:12Right? Now, how do you think it learned how to drive? By actually teaching itself, by learning on dashboard cameras, right?
00:02:22Every car has a dashboard. So if you provide that to an AI, it can figure out how to actually operate a machine.
00:02:31So as I said, the AI is no longer the hard part, the data is, and this is what I want to talk about, right?
00:02:39So for chat LLMs, there's trillions of words and texts, right? So we have so much text data to train the LLMs.
00:02:49For image generation, we got billions of images, labeled images, so that's also a huge database.
00:02:56But for robotics, there's only about a million videos. It's a very small, limited data set of robots doing things, right?
00:03:09And so the idea here is that, of course, a lot of companies out there, what they do is they pay people to record actions, right?
00:03:20So for example, hey, record me how you open a door, or record to me how you're sitting on the chair.
00:03:28But then the problem becomes is that when somebody is told to do something, they do not do it naturally.
00:03:34So I call it instructed, right? If you're told to record how you open the door, and you're doing it for somebody,
00:03:40it's not going to be the same as if you just walk into the house. The movement is totally different.
00:03:45So is that data good for robotic training? In my assumption, no. It's biased data, and it doesn't deliver the same results.
00:03:53So obviously, the robot is not going to be as effective as if it's intuitive, right?
00:03:59So I just threw in some examples on the slide, right?
00:04:06The build-by-hand data is not the same as the data that is kind of on the side, right?
00:04:12And also, the fact is that people behave very differently when they're on the camera.
00:04:16Some of you know, some people are camera shy, so as soon as you put a camera on them, they change.
00:04:23But in the natural world, it's a totally different thing.
00:04:27So this is what we are trying to solve in bright data, right?
00:04:35Again, so why is the usual data sources are not enough, right?
00:04:37Simulations. Virtual, it's cheap, but it's, you know, everybody's playing video games.
00:04:44The physics in video games are almost there, but they're not good enough to train a robot on it, right?
00:04:50A hand control, the person controlling a robot, it's doable, but how many hours a day can you record a robot?
00:04:57How many people you need to actually get huge-scale data, right?
00:05:01I mean, you can record maybe eight hours a day, every day.
00:05:04How many hours are you going to get in a year?
00:05:07It's not really scalable.
00:05:09And of course, there's already pre-built data sets, but as I said before, they're small.
00:05:13There's only about a million videos at this point.
00:05:15So what is the alternative?
00:05:17The alternative is the web.
00:05:20Let's just think about YouTube.
00:05:22How many videos are there on YouTube?
00:05:25I don't know, five billion videos?
00:05:28How many actions are in those videos that a robot can replicate?
00:05:34I mean, let's think about it.
00:05:35How many do you think videos are there of a person opening the door, first person view?
00:05:42Millions of hours.
00:05:43You'll be surprised.
00:05:45So every day, the web video shows gravity, motions, right?
00:05:53So cause and effect, obviously accidents, the biggest cause and effect, right?
00:06:00How people handle objects, and billions of hours.
00:06:04Great training material for robots, but what is the problem?
00:06:08The problem is that there's a lot of noise, right?
00:06:12Oh, and for example, right?
00:06:14So one of the AI models that Meta trained, they gave it about a million hours of real world videos.
00:06:21And then all it took is 62 hours of real robotics data to actually control a real robot.
00:06:29So if you think about it, the data was collected publicly, right?
00:06:33So it trained the robot on just random things, and then in 62 robotic hours, it was already moving.
00:06:41So no simulations needed, autonomous robotics delivered fast.
00:06:52But the videos don't show how the robots move, right?
00:06:54So you might think, well, listen, how useful is it a person pouring water into a cup for a robot?
00:07:01So there is methods actually out there how an AI can distinguish and actually learn from the actions that are in the video.
00:07:09If you take two frames, frame by frame, and the AI measures the movement difference, right?
00:07:15So you have an image A and image 2, and there's a small movement that is changing between the two images.
00:07:21An AI can actually measure that change.
00:07:24So if you keep doing that through the whole video, an AI can actually figure out angles, distance,
00:07:30and everything that it needs to train and robot, right?
00:07:34So we don't necessarily need the sensor's data, but of course, the video itself that you download
00:07:41from YouTube or the video that you extract the data from doesn't necessarily get you 100% there.
00:07:48You do need to have some processing on top of that, but it's available.
00:07:53It's out there, and all you need to do is process it.
00:07:58So a few more examples, right?
00:07:59So for example, NVIDIA, when they train in the Cosmos robot, they're throwing out about 96% of the video.
00:08:07So what does that mean?
00:08:09They download a million hours of video,
00:08:12and the only part is 4% is actually useful.
00:08:15Everything else gets thrown out.
00:08:16That's wasted compute.
00:08:18That's wasted bandwidth.
00:08:19That's wasted storage.
00:08:21That's just a lot of wasted money.
00:08:23Okay, stable video diffusion is throwing out 74% of the videos that they're downloading.
00:08:30So in Bright Data, we are trying to solve that problem, and we're trying to take it a step forward.
00:08:37So what we propose is search first, collect second.
00:08:41What we do is we are actually indexing videos, and we are allowing you to search for specific actions.
00:08:50Since we're talking about a person opening a door, let's keep using that example.
00:08:55What you can do on our platform is actually input detailed information, right?
00:09:01A query, person washing dishes, person folding a t-shirt, and so on.
00:09:06And we will provide the snippets from the videos of these actions.
00:09:13So what does that mean?
00:09:16That means that there's less waste of data collection.
00:09:22There's less waste of storage and bandwidth, right?
00:09:25So a little more on how it works.
00:09:27Obviously, you define what you're looking for.
00:09:32We search through billions and billions of pre-indexed videos.
00:09:36Not by keywords, but by actions.
00:09:39And we provide you with ready-to-use clips.
00:09:43And obviously, they're already prepared for your training.
00:09:47So all you gotta do is process them a little bit for maybe motion, distance sensors, and so on.
00:09:52And it's ready to be ingested into a robot.
00:09:57I have a small example of a video of how it actually works on our platform.
00:10:02So let me just hit play here, if I can find it.
00:10:07So here we see a person washing dishes by hand in sink, removing grease from the plate,
00:10:12using sponge and soap, including rinse and placing dishes into drying rack, close-up hand interactions.
00:10:19So right now, what is happening is it's searching through billions of videos.
00:10:24Now, on this video, it's a bit old.
00:10:26We only got about 100 million index there, but now we're up to 1.1 billion videos.
00:10:31And it will literally, like, this is a little sped up, so I didn't want to waste you guys' time.
00:10:35But it will provide to you clips from videos of people washing dishes.
00:10:41And you can see right here all the videos that we find.
00:10:47Now, some of these videos might not have to do anything with washing dishes.
00:10:50They might have a cleaning kitchen.
00:10:52They might have to do something else.
00:10:54And here's another example.
00:10:55For example, human folding different types of clothes.
00:10:59Right?
00:10:59So, a little more description.
00:11:01The more description you give, the more you actually get back.
00:11:06So again, a little sped up.
00:11:08Let's see.
00:11:14I want you to understand, it could be anything.
00:11:16A person putting on makeup.
00:11:17Let's say you have a brand and you want to find videos where your brand is being found.
00:11:21All that can be also provided.
00:11:23So, you see people folding clothes.
00:11:26So, now you can train a robot on how to fold clothes.
00:11:34Of course, everything is available via API.
00:11:36We don't expect people to actually do anything manually.
00:11:39Right?
00:11:39So, you can trigger it by API.
00:11:41You get a return snippet of the video, the URL.
00:11:46We give you the link to the original video if you want to watch it.
00:11:49But the main thing here is that we can give you the snippets of the video to train your robotics.
00:11:57Now, I'm not sure if any of you are training any robots.
00:12:01So, I'm going to give you another example of how useful it could be.
00:12:04Let's say you have a brand.
00:12:06And you want to know where your brand is being demoed on videos.
00:12:13The video titles doesn't matter because it doesn't actually, it's not about your product.
00:12:19It's just a person doing some podcast, recording something.
00:12:22But, you see a woman putting on makeup.
00:12:25And you see on the table that this is your brand of a makeup.
00:12:30So, now all of a sudden, you can find all videos for your brand where it's being used.
00:12:34Whatever it is.
00:12:35Right?
00:12:35By searching the name of the video, you won't find it.
00:12:38By video indexing, you can actually find specific things that you are interested in.
00:12:45Apple falling from the tree.
00:12:47And so on.
00:12:48So, what comes back?
00:12:50We provide to you this time stamp.
00:12:52We provide to you the score matching.
00:12:53How close it is to your query.
00:12:55And of course, the frame count.
00:12:57How many frames are about what you're looking for.
00:13:03So, what does this change?
00:13:06This changes a lot.
00:13:07Okay, less waste.
00:13:09You don't need to download millions of hours of videos.
00:13:12You can actually get exactly the specific actions
00:13:16that you are interested in.
00:13:20Again, it's very crucial for robotic training.
00:13:24Because noise creates problems, hallucinations.
00:13:28Right?
00:13:29This removes the noise.
00:13:34What is it useful for?
00:13:35Not only robots, but obviously self-driving.
00:13:38For example, Wimeo.
00:13:39Right?
00:13:40Trained on dash video cameras.
00:13:43And there's millions, hundreds of millions of hours on YouTube.
00:13:46Of dash cams.
00:13:47How many of you like to watch accidents?
00:13:49Dash cam cam accidents.
00:13:52I'm sure everybody watched some.
00:13:53Crazy driving in Russia.
00:13:55Right?
00:13:56So, this is what we're talking about.
00:13:58Right?
00:13:58The videos are out there.
00:14:01Okay?
00:14:01A person running a red light.
00:14:02A person stopping on a red light.
00:14:04Making a left turn.
00:14:05Making a right turn.
00:14:06And so on.
00:14:06All of that can be a very useful data for training.
00:14:11And any other world models.
00:14:13Physics.
00:14:14Right?
00:14:14Robots need to understand physics.
00:14:16What is the gravity?
00:14:18How to sit?
00:14:18How to walk?
00:14:19How to move?
00:14:21Again, of course, you can pre-record it.
00:14:25A lot of times right now, I speak to some of the people.
00:14:27And they have millions of people recording videos for them.
00:14:30Hey, can you guys record me how you open doors?
00:14:33That's crazy.
00:14:34Why do you need a million people doing that when everything is available online?
00:14:39Online is one of the hugest databases out there.
00:14:43It's just people don't really, like, it's not that easy to use it, right?
00:14:46I mean, right now the only search available to us is by the keyword of the name of the video.
00:14:53And of course, one source, the whole public video web in one place.
00:14:59I mean, we got YouTube, we got Vimeo, we got so many different video providers out there.
00:15:06Billions and billions of hours of training data just there waiting for you to grab it,
00:15:13collect it, and process it.
00:15:14So basically, this is a new product that we're doing at Bright Data, right?
00:15:18Indexing videos so that you guys can find specific actions.
00:15:23Anything you can think about, you know, again, it could be useful for brands.
00:15:28As well, it's not only for training data.
00:15:31Anything you can think of could be useful.
00:15:35Maybe you are a gamer and you want to know how do you beat this level?
00:15:40But there's no exactly video like that.
00:15:42So you can search, okay, people beating the level on the video game, right?
00:15:47Anything, the world is yours.
00:15:52So I'm going to wrap it up.
00:15:54I don't know if you have any questions, but if you want to connect with me, talk about it more,
00:15:58feel free to add me on LinkedIn.
00:16:01And yeah, I'm going to thank you guys for your attention.
00:16:06And I'm at a Bright Data booth if you want to talk more about it.
00:16:10And thanks for coming.
00:16:17I'll see you next time.

Key Takeaway

Training autonomous physical AI models requires shifting from downloading raw web video—where up to 96% of data is discarded—to action-based video indexing that extracts precise, uninstructed human motion snippets directly via API.

Highlights

  • Robotics AI faces a massive data bottleneck, with only about one million dedicated robotics training videos available compared to trillions of words for LLMs and billions of labeled images.

  • Data collected by paying people to perform instructed actions produces unnatural, biased movements that fail to deliver intuitive robotic performance.

  • Meta trained a real-world robot using roughly one million hours of public video paired with just 62 hours of actual robotics telemetry.

  • AI models waste massive compute, bandwidth, and storage when processing unindexed public video, with NVIDIA's Cosmos filtering out 96% of downloaded footage and Stable Video Diffusion discarding 74%.

  • Pre-indexing public web videos by action rather than keyword allows developers to query exact movement snippets, timestamp data, match scores, and frame counts via API.

Timeline

The Data Disparity in Physical AI Development

  • Robotics AI development progressed from single-robot image training in 2022 to multi-robot control in 2023 and open-source humanoid models by 2024.
  • AI model architectures are no longer the primary development bottleneck; training data availability is the main constraint.
  • Robotics models suffer from a extreme data shortage, possessing roughly one million training videos compared to trillions of text words and billions of labeled images for software AI.

While autonomous systems like Waymo demonstrate real-world deployment by learning from dashboard cameras, the broader robotics industry lacks foundational training data. While natural language processing and computer vision benefit from massive web-scale datasets, physical AI development remains constrained by the limited volume of available robotics recordings.

Flaws in Instructed Data and Traditional Data Collection

  • Paying humans to record specific physical tasks yields biased, unnatural motion data because subjects alter their movements when self-conscious or instructed.
  • Physics engines in video game simulations lack the necessary real-world fidelity to train physical robots effectively.
  • Manual teleoperation is unscalable, yielding at most eight hours of daily recorded data per human operator.

Standard data collection methods fail to scale or capture authentic movement. Staged demonstrations create unnatural biomechanical patterns, game engine physics fail to replicate real-world dynamics accurately, and direct human teleoperation requires too many labor hours to supply model needs.

Extracting Physical Laws and Movements from Unstructured Web Video

  • Public web platforms like YouTube contain billions of videos capturing real-world physics, cause-and-effect scenarios, and uninstructed object handling.
  • Meta successfully trained a robot by pairing one million hours of unstructured real-world video with 62 hours of physical robot data.
  • Computer vision models calculate angles, distances, and trajectory changes by analyzing frame-by-frame movement deltas across video sequences.

Unscripted video footage on the web provides a massive repository of implicit physical training data, including gravity, motion, and object interactions. By evaluating positional differences frame by frame, AI systems extract accurate geometric and spatial vectors without requiring specialized physical sensor logs during pre-training.

Eliminating Compute Waste Through Action-Based Video Indexing

  • Major AI training pipelines waste substantial compute, storage, and bandwidth, with NVIDIA discarding 96% of downloaded video for Cosmos and Stable Video Diffusion discarding 74%.
  • Action-based indexing allows developers to search pre-indexed video repositories using natural language prompts for specific physical behaviors rather than video titles.
  • An API-driven action search returns exact video URLs, timestamps, match confidence scores, frame counts, and cropped activity snippets.

Traditional video ingestion forces training pipelines to download massive media files only to discard the vast majority of non-relevant frames. Pre-indexing video datasets by specific human actions—such as folding clothes, washing dishes, or unboxing products—allows developers to retrieve only the precise frame sequences needed, eliminating redundant data transfer and processing overhead.

Applications for Autonomous Driving, Physics Models, and Brand Monitoring

  • Action indexing retrieves long-tail autonomous driving scenarios, such as traffic violations or dashcam accident footage, across public video repositories.
  • Video action indexing isolates visual product placements and brand interactions inside unlabelled media like podcasts and vlogs.
  • Replacing manual recording campaigns with public web video indexing cuts training costs while supplying foundational world models with real-world physics data.

Searching media by visual content rather than metadata expands dataset acquisition beyond robotics into autonomous driving edge cases, physical world simulation, and visual brand intelligence. Aggregating action snippets across global video platforms provides an immediate, scalable alternative to paying human contractors for staged video recordings.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video