Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI

스크립트

00:00:00I think all of you know here that how much progress has AI
00:00:15made in the last five years.
00:00:17Seems like it's more than the last 100 years of progress.
00:00:22And what is happening, especially in AI,
00:00:24is AI came for language, then speech, audio, video.
00:00:28And everybody's excitement is next is robotics.
00:00:32And this excitement is going on peak this year for some reason,
00:00:37which is still beyond my understanding.
00:00:40And you can also see leaders like Jensen
00:00:43talking about physical AI as the next frontier of robotics.
00:00:46So it seems like if you open Twitter or LinkedIn,
00:00:49it seems like robotics is already here.
00:00:51And we'll have home robots in our homes sometime by December,
00:00:56as it's being said by multiple people.
00:00:58But I'm here to highlight that robotics has been almost here
00:01:03for the last 70 years.
00:01:06Now, those of you who got into robotics in the last four or five
00:01:10years, go and take any course in robotics.
00:01:14And if you can distinguish the videos in robotics
00:01:17from 30 years ago versus today, I love your dinner.
00:01:23So to give you an example, let's look at this robot.
00:01:26So this robot here is an image.
00:01:27And the goal of this robot is to look at this image of blocks
00:01:31and arrange these blocks in front of it
00:01:34so that it looks the same pattern as in the image.
00:01:37OK?
00:01:38This is an extremely simple task.
00:01:40But if you think about this, it's a 3D block.
00:01:42You have to look at understand 3D from each side.
00:01:44And the robot can do this pretty precisely.
00:01:47Any guesses how old this result is?
00:01:51Anyone?
00:01:51It's a guess.
00:01:52You can make any guess.
00:01:53Huh?
00:01:5440 years.
00:01:55OK.
00:01:55That's the highest?
00:01:59This is from 1960s.
00:02:01This is way before there were computers.
00:02:03OK?
00:02:04So this is called MIT Copy Demo.
00:02:06This was the start of--
00:02:09this predates AI, as you know today.
00:02:12Look at this one.
00:02:17One exhibit at the nuclear congress in Philadelphia
00:02:19has the perfect formula for popularity.
00:02:21What's happening here, the guy behind the scene
00:02:24is controlling these robots, a leader-follower system.
00:02:27And if you look at any big lab, any big company,
00:02:31any big academic lab, how they get data,
00:02:33they use tele-operation system, which
00:02:35follows the exact same principles?
00:02:38Any guesses for the year for this one?
00:02:40The skilled operator of the electronic device--
00:02:42Now, everybody will correct aggressively on the year.
00:02:45But this is from 1957, 68 years ago.
00:02:49OK?
00:02:51Ask anybody in your family who is older than--
00:02:53actually, they may not be around.
00:02:55So because this is--
00:02:57you have to ask somebody who is 80 years old,
00:02:59what was the state of technology at that time?
00:03:02Like RAM, for a few kb of RAM, you would have a size of a room,
00:03:06like this big of a setup.
00:03:07And this is from that time that people could make it work.
00:03:11On not just this, you keep coming back every decade,
00:03:14the videos look as cool.
00:03:15Like here is a robot, a humanoid, juggling,
00:03:18playing foosball with a human.
00:03:21And this is also from, they were all pre-date deep learning.
00:03:25This is from eight years ago.
00:03:26So this robot is a result from Berkeley,
00:03:29can clean up the whole table, can arrange items in a box.
00:03:33This was done by one grad student with a single GPU machine.
00:03:36Nothing more than that.
00:03:37OK?
00:03:38Now, this picture of-- and if I colorize all these videos
00:03:41from the past and apply Gen AI filter for the modern voice,
00:03:45you cannot tell that this video is from 19657
00:03:49or is it from today.
00:03:50And these are not even the oldest results.
00:03:52The oldest ones go to 1940s.
00:03:54OK?
00:03:55So what is this that if you look at last 70 years,
00:04:00every other technology, every other technology
00:04:02has come way far--
00:04:04language understanding, computer vision, mobile phone, chips.
00:04:08Why is robotics stuck in this primitive age for this long?
00:04:11And the reason is a general brain.
00:04:14Robotics has always been approached as a hardware
00:04:17problem from ground up.
00:04:18And this is the reason that everything around robotics
00:04:21has progressed, but robotics is still stuck in the same land
00:04:24from the last 70 years.
00:04:25So there's a very famous paradox called Moravec's paradox.
00:04:27I don't know if you know about Moravec.
00:04:29He was also a CME professor, and I'm also a CME professor.
00:04:32So I have to quote him for sure.
00:04:34He was one of the founding figures in AI.
00:04:36After 30 years of working in robotics,
00:04:39he arrived at a very simple conclusion.
00:04:41Hard is easy.
00:04:42Easy is hard.
00:04:43OK?
00:04:44Whatever a human believed to be hard
00:04:45is extremely easy for computers and vice versa.
00:04:49And you can see it happening right in front of your eyes.
00:04:51You would say doing math is hard.
00:04:53Math Olympiad gold medal is really, really hard.
00:04:56Climbing a stair?
00:04:57It was super easy.
00:04:59Now look around you in technology.
00:05:01Where we are in terms of what has been solved
00:05:02and what's not being solved.
00:05:04So this is robotics for you, OK?
00:05:05It's very-- it is not yet another application of AI.
00:05:09It is what AI was founded for in the very beginning
00:05:12and has made very little progress towards.
00:05:14This is not yet another application
00:05:15where deep networks can come in and attack the field.
00:05:18This has to be thought of with fundamentally first principles
00:05:22from the ground up.
00:05:23This is a very famous example of where
00:05:25you can have a computer beat Garry Casparov on chess
00:05:29in '90s, but you still don't have a computer, a robot,
00:05:34which can pick up the chess pieces and arrange them
00:05:36on the chess board.
00:05:37Now, so how do we go about solving this, OK?
00:05:43So on a more positive note, we have a magic sauce
00:05:47for AI success, right?
00:05:49You get big data set.
00:05:51You train big models.
00:05:52And magic happens, OK?
00:05:54Now, we know this template is working very well
00:05:57in many topics, OK?
00:05:59Now, can we apply the same recipe to robotics?
00:06:02Well, it's not directly applicable because we have no data.
00:06:06There is no internet of robotics data.
00:06:09And as I said earlier, you can go and collect data manually
00:06:12on the robot called teleoperation.
00:06:14And as you notice, teleoperation is not a five-year-old thing.
00:06:18This is a 68 or 70-year-old thing.
00:06:21So the funny part is, if you go back
00:06:24and you look at the whole evolution of GPT-3, GPT-4 models,
00:06:28go back to GPT-3 three years ago.
00:06:31Seems like a lifetime ago.
00:06:33GPT-3 only began working and caught people's attention
00:06:36when you could train those models on trillions of tokens.
00:06:40So GPT-3 was already 30 trillion-plus token.
00:06:42And today, hundreds of trillions of tokens.
00:06:44And if you collect data manually by teleoperation,
00:06:47it takes you about one minute to get one example.
00:06:49You can do the math.
00:06:50If you had all of US population, it
00:06:53will take you more than a century to reach the same scale
00:06:55as of GPT-3, which is like a--
00:06:59and nobody uses GPT-3 today.
00:07:00OK?
00:07:01So this is very slow and expensive.
00:07:02So in robotics, I would argue nobody
00:07:05has really scaled robotics yet.
00:07:07And we are very far from talking about scaling robotics
00:07:10with the way other areas have seen scale.
00:07:14OK?
00:07:14So this is where we have been focusing on.
00:07:17Scaled is about three years old.
00:07:19But I have been working on the problem for more than a decade.
00:07:22That's all I've done in my career.
00:07:24Nothing is.
00:07:26So at scale, what our thesis is, is
00:07:27to build what we call an omnibodied intelligence.
00:07:32Any robot, any task, one brain.
00:07:35OK?
00:07:36Any robot-- it can be a humanoid.
00:07:37It can be a quadruped.
00:07:38It can be a robotic arm on a conveyor belt or a dextrous hand.
00:07:41It doesn't really matter.
00:07:42And even this hypothesis is way more
00:07:45general than one would argue humans are,
00:07:48because we control our own body.
00:07:50And why do we have to go so general?
00:07:51Well, the argument is, in robotics,
00:07:54there is no data anyway.
00:07:56So I can't pick and choose which hardware do I use data from.
00:07:59And we should be able to use data from any kind of hardware,
00:08:02any kind of task, any kind of scenario.
00:08:04And that's the only way to truly achieve
00:08:06the scale of what language models achieved three years ago.
00:08:10And the goal here is this whole idea of any robot,
00:08:13any task, one brain.
00:08:15And through this talk, I'll hopefully
00:08:16convince you why this is the way to go towards robotics.
00:08:20OK?
00:08:20So this is the rough intro.
00:08:23But before I go into the more details,
00:08:24let me show you just a teaser result.
00:08:28So in this result, every single robot
00:08:31is from a different company, different hardware.
00:08:34And they're all controlled by a skilled brain,
00:08:37whether it's humanoids going up and down stairs,
00:08:39any kind of scenarios.
00:08:42And these are not new results.
00:08:43They're like a couple of year-old results in here.
00:08:47Robust to disturbances.
00:08:51You can put them zero shot in new scenarios.
00:09:03So the idea here, every single robot in this video takes
00:09:08any action anywhere in the world.
00:09:11The brain behind the scene improves,
00:09:13because it's an omnibodied brain.
00:09:15Yeah, this is really hard, right?
00:09:16Because eggs are OK.
00:09:17So that's a teaser of what we work on.
00:09:33So I want to make this talk more scientific and more informative
00:09:38than a company ad.
00:09:39So I'll talk about how do we scale data in robotics.
00:09:42Let's take a tool back as to how people address this.
00:09:45I'll take a look back at my own career.
00:09:48And my hypothesis for data has been changing over the years.
00:09:51When I began working and scaling things,
00:09:55the idea was, you can collect data for robots manually,
00:09:59but it is too slow.
00:10:00And robots should be allowed to collect data by themselves.
00:10:03So we had this whole idea of curiosity-driven exploration,
00:10:06allow robots to explore in a curious way in the environment.
00:10:10And it has a lot of good parts, like no human required.
00:10:13Robots can go and keep playing, collect more data.
00:10:15We called it play data at the time.
00:10:17It's very rich, because it has force, sensors, and joint angles.
00:10:21But the difficulty is, it is only in the physical world.
00:10:24So it's very difficult to scale.
00:10:27Google was at it at the time, having hundreds of robots.
00:10:31Even for Google, it's too expensive and too slow to scale.
00:10:34The other idea is, again, tele-operation.
00:10:36As I said, it's an old idea.
00:10:38But again, the same issues.
00:10:39Impossible to scale, because now you don't need robot.
00:10:42You also need humans.
00:10:43So it's even more expensive.
00:10:45And diversity is very limited.
00:10:46Because even if you put a robot in one setup with a human,
00:10:50you can get 1,000 examples.
00:10:51But they'll all be in the same setup.
00:10:53So what you ideally want, carry the robot to a new home
00:10:57every day, a new scenario.
00:10:58And that's just impossibly hard to scale.
00:11:01Then we had a major breakthrough in learning
00:11:03from simulation.
00:11:04This was one of the first results where one could
00:11:08show deep reinforcement learning, trained in simulation,
00:11:12transferred to a real robot.
00:11:13This was an award-winning paper at the time.
00:11:16And now it's used in every humanoid, every company out there.
00:11:20This was from a lab at Berkeley and CMU.
00:11:23But again, there are pros and cons.
00:11:25It's very easy to scale.
00:11:27But diversity is hard to get, because you have to engineer
00:11:29every scene in simulation manually.
00:11:31So if you're listening to this, there
00:11:34is no single answer I'm coming to.
00:11:35There is no single solution.
00:11:39This is another work we did earlier.
00:11:41This is all before scale.
00:11:43Learning from human videos.
00:11:44You watch the human do things, and then robot copies this.
00:11:47Again, high diversity.
00:11:48You can use videos from YouTube, et cetera.
00:11:50Highly scalable.
00:11:51But it's a very poor form of data,
00:11:54because it's very far from robot.
00:11:57So what is the solution here?
00:11:59The answer is there is no golden path.
00:12:01You have to think about data in the context
00:12:05of different features.
00:12:06And in my opinion, there are only three features
00:12:09that matter in robotics--
00:12:12scalability, diversity, and how close you are
00:12:15to your robot joint angles.
00:12:18Scalability means, can I quickly scale it across scenarios?
00:12:22Diversity means, can I get diverse data?
00:12:25Because just having 100 trillion tokens
00:12:27is completely useless if they're all
00:12:29in the same environment and same scenarios.
00:12:32And closeness to robot means, how far are you
00:12:35from the ground truth of robot-owned joint angles?
00:12:38So if you look at simulation, very scalable, low diverse,
00:12:42but moderately close to robot.
00:12:44Human videos are very scalable and diverse,
00:12:46but very far from robot data.
00:12:47You have to learn to map the human to robot.
00:12:49This other two, tele-op and the manipulation interfaces,
00:12:52they are kind of in between.
00:12:54And you can see here, right now, around the world,
00:12:57if you look at different companies,
00:12:58they're all focusing on one of these approaches.
00:13:01Most of them are on tele-op or Yumi setups.
00:13:05Very few on everything else.
00:13:07But there is no golden answer.
00:13:09Each one of them have a downside.
00:13:11But the golden light here is that they're all complemented
00:13:15to each other.
00:13:16They are not-- their cons are not exactly matching from each
00:13:19other.
00:13:20This is where the recipe that we have converged to over the years
00:13:24is to realize-- separate the training into two parts,
00:13:27pre-training and post-training.
00:13:29But that's a surprise, right?
00:13:30But how do we pre-train?
00:13:31You want to pre-train using data, which
00:13:32is highly scalable and diverse, but maybe low quality.
00:13:35So simulation, human videos, et cetera.
00:13:37So that's how you pre-train.
00:13:39And then for post-training, you use data from tele-operation.
00:13:42Now, what is tele-operation data?
00:13:43It's very high quality, but low in amount, OK?
00:13:47High quality, low in amount.
00:13:48It exactly reminds us of the recipe in language models.
00:13:51You pre-train data on the internet.
00:13:53Then you post-train for coding, for your own company, et cetera.
00:13:59But in robotics, there is one more bucket,
00:14:00which is deployment data.
00:14:02And deployment data is highly scalable
00:14:05once deployment's scaled.
00:14:07So over time, this data will take over everything else.
00:14:11And what we are trying to build is what
00:14:14we call this data flywheel, which
00:14:16takes this data from post-training time
00:14:18and puts back in pre-training.
00:14:20And now you can see why this idea of omnibodied brain
00:14:23is extremely important, because this
00:14:25is very unlikely that you have only one robot, only one
00:14:28version deployed forever in every task around the world.
00:14:32This never happens in any area of hardware, except chips,
00:14:36because they are very hard to manufacture.
00:14:38And you can still see, even in the chips,
00:14:40inference chips are coming left and right
00:14:42for many companies these days.
00:14:43So in hardware, this is never the case.
00:14:45You have only one version of the robot, one shape,
00:14:48which is why omnibodied intelligence
00:14:50is the enabler of what we call a deployment data flywheel.
00:14:56So let's now look at a few results very quick.
00:14:58Now, this brain is extremely general,
00:15:00so we can do a variety of tasks very quickly.
00:15:03You may have seen many tasks like laundry folding,
00:15:06et cetera.
00:15:07It's a very popular task in Silicon Valley for some reason.
00:15:10And the argument here is, when have you ever
00:15:14thought, while folding a t-shirt, that if I miss my hand
00:15:18by one centimeter, my t-shirt fold will be a blender?
00:15:22You don't think like this.
00:15:23People don't even think while folding.
00:15:25They just hold anywhere.
00:15:26You just do something.
00:15:27So the tolerance for error is extremely, extremely high.
00:15:31Then why is this task hard?
00:15:35Anybody-- why do people get fooled
00:15:38into believing this task is hard?
00:15:39Let me put it that way.
00:15:42Fabric, right?
00:15:43Why is fabric hard?
00:15:46Simulation, but nobody's using simulation anyway.
00:15:48This is all from teleoperation.
00:15:49Why is this hard?
00:15:51You know why is it hard?
00:15:52Because it is hard for classical version of robotics.
00:15:56Classically in robotics, people would model the whole physics,
00:15:59create models by hand, and then do this.
00:16:01It's very hard for that.
00:16:02But for deep learning-based robotics,
00:16:04this is the easiest task possible,
00:16:06because it has high tolerance for error.
00:16:08So it is completely--
00:16:09now, so many companies focusing on this task.
00:16:12Now, what is hard is, I would say,
00:16:13somewhere like, let's say, this task.
00:16:17If I ask you, before seeing this video,
00:16:19is this task doable without having hands?
00:16:22Very likely, half the people will say no,
00:16:24because it requires putting--
00:16:25like, how many of you have lost airports?
00:16:29And in our company, there's a channel
00:16:30called right airpod, because people keep
00:16:32losing their right airpod.
00:16:34Now, in this case, the robot does not have hand.
00:16:36It has gripper.
00:16:37So here, the task is very hard for the gripper.
00:16:41So it requires a higher level of intelligence,
00:16:43because the arm has to go and orient itself
00:16:46to pick up the airpod in the right manner,
00:16:49such that it can be inserted, because the grippers can only
00:16:52close up like this.
00:16:54They're parallel to your grippers.
00:16:55So you cannot turn the airpod at the very end.
00:16:59Now, what you are seeing here, these are not real airpods.
00:17:02They are fake ones from Temu, like $5, $10 each.
00:17:05So they don't have a magnet inside.
00:17:07So the robot has to really work hard to put this inside properly,
00:17:11because there is no magnet to pull it easily.
00:17:13So these are all fake ones.
00:17:15This is even harder than it appears in the video.
00:17:17You can also-- once you build the general brain behind the scene,
00:17:23you can also go and learn it from a variety of just human videos
00:17:26without actually having any fine-tuning data
00:17:28at teleoperation time.
00:17:29So in this scenario, what we did, we
00:17:31trained on human videos like this, like egocentric videos.
00:17:34We are now trying to transfer it to more third-person videos.
00:17:38And if third-person works, you can learn from YouTube,
00:17:40any kind of open source data.
00:17:43So in this case, we see the video,
00:17:45and then we add less than one hour of robot data.
00:17:47So very quick transformation.
00:17:50And then it can transfer to humanoids.
00:17:53Now, these ones have hands.
00:17:54Hands or no hands is a big debate.
00:17:57People often use hands as an excuse as to why robotics is not
00:18:00here, but that's not the case.
00:18:02It's always the intelligence behind the scene.
00:18:07We don't deploy hands right now, because there are none
00:18:09available which can be deployed in factories.
00:18:12They all break within 100 hours.
00:18:14Pick anyone.
00:18:17If you have a better hand, I would love
00:18:19to buy 10 units ASAP to test.
00:18:22So it's robust scenarios.
00:18:25Now, the idea is you can do many tasks by watching humans.
00:18:28And the reason we can work with very less data, which
00:18:31is less than one hour of robot data,
00:18:33is because robot imagines things in its own head
00:18:36and tries to multiply the learning for many scenarios.
00:18:39What you see here is a completely fake video.
00:18:42You can call it robot's dream.
00:18:43So it's inside the robot's own model,
00:18:45where it can imagine scenarios.
00:18:48So this is not real data.
00:18:49This is all fake from robot's own model.
00:18:52You can transfer it to more complex tasks
00:18:56and even lower cost hardware.
00:18:57So this is an egg making, like omelet making task.
00:19:02Now here, this whole setup costs about $4,000.
00:19:06So extremely cheap arms compared to what you see out there.
00:19:11And if you notice here, this setup is too cheap
00:19:15to even have any sensors.
00:19:16So the only sensor here is just a camera, no force, nothing else.
00:19:20And the robot can do tasks which require forces from vision.
00:19:25Now, I'm not saying that's the future.
00:19:26Like you shouldn't be doing this.
00:19:27I'm sure the sensors will improve.
00:19:29But from the existing sensors, we
00:19:33are way behind in where we can be from just intelligence
00:19:36perspective.
00:19:38So this can keep on going.
00:19:39This is my advisor from Berkeley.
00:19:41He did not believe that robots can do it.
00:19:43He just kept standing for like one hour.
00:19:45And the robot kept making omelet.
00:19:48And the interesting part here is this
00:19:49is fully end-to-end system.
00:19:51There is no state machine, nothing.
00:19:53So the reason robot is making egg,
00:19:55because there is an empty plate in the front.
00:19:56I don't know why it's fluctuating.
00:19:58There is no cut in the video.
00:20:00I assure you of that.
00:20:01It's HDMI problem.
00:20:02So as soon as you remove the plate--
00:20:06sorry, I don't know why this happened.
00:20:08Yeah, so there is no state machine.
00:20:10When the robot starts, when the robot ends,
00:20:11it's all automated.
00:20:12And it's very robust to even disturbances.
00:20:14You can change objects around.
00:20:16You can add-- these are all unseen objects.
00:20:20Only the pan and the gas stove is the same.
00:20:22Everything else is different.
00:20:23And the robot can keep on going.
00:20:25So you can get basic robustness.
00:20:27So this is all trained with less than 10 hours of data.
00:20:30So it's very low data to be learning this robustness.
00:20:33And it's coming from the base model behind the scene.
00:20:37I'm skipping this in the interest of time.
00:20:39So unlike traditional robotics pipelines,
00:20:42where you have planning, mapping, et cetera,
00:20:44this is an end-to-end brain.
00:20:46Now, end-to-end is heavily used term in many areas of AI.
00:20:50But when I mean end-to-end, I really mean end-to-end.
00:20:53It reads directly from the cameras.
00:20:55It applies power directly to motors.
00:20:59So we use nothing in between, except just a PID
00:21:02at the very end, to go from 100 to 500 Hertz.
00:21:04But PID is not robotics contribution.
00:21:06This is before.
00:21:07It predates to World War II or something.
00:21:10Now, the interesting part is, vision for us
00:21:13is yet another input.
00:21:15So nothing that crazy about it.
00:21:17So if you have seen-- you may have seen a lot of results
00:21:20of robots dancing, doing karate, kung fu, backflip.
00:21:23Go back and think, how many results
00:21:27have you seen of robots going up and down stairs?
00:21:30Very few.
00:21:31Like, even the top companies out there in humanoid,
00:21:34they have not shown beyond one sample stair
00:21:37just to check the tick box.
00:21:39And people talk about humanoids are coming, China versus US.
00:21:42All of this is sort of BS.
00:21:44Like, if humanoids cannot climb stairs,
00:21:47what's the point of having legs?
00:21:50That's the only reason why you have legs.
00:21:51Now, this is actually a Moravex paradox at action.
00:21:55Left one actually is much easier for a robot.
00:21:58A backflip, dancing, it's much easier for the robot.
00:22:01And you may think, OK, why does this paradox exist?
00:22:04Think one level deeper.
00:22:05When a robot is doing a backflip, it
00:22:09has to only know about its own body, nothing else.
00:22:13Right?
00:22:13So it's-- and when everything is known or fully observed,
00:22:17that's what computers are good at.
00:22:20Because they can simulate every possible setup.
00:22:25But the right one, even simply climbing on stairs,
00:22:27requires seeing the stairs at the first place.
00:22:30Or what's the height, what's the width.
00:22:32You don't measure height and width exactly,
00:22:34but you adjust to all the disturbances.
00:22:35And that requires vision.
00:22:36So right one requires understanding.
00:22:38And that's why it is hard, and you don't see any of this.
00:22:40But for us, it's just get another result.
00:22:42It doesn't really matter.
00:22:43So we can-- we put this result out one and a half year ago.
00:22:46We shot this two and a half years ago.
00:22:48But this robot can go any scenario.
00:22:52It is not stumbling on things.
00:22:53It intentionally steps over things.
00:22:55And there is no mapping or planning.
00:22:56It doesn't make any 3D map of the surroundings.
00:22:59It's all operating from camera on the torso.
00:23:02And you can put this in any kind of setup, any kind of stairs.
00:23:05I'm going faster here.
00:23:07These are, you know, fire-scape stairs.
00:23:11It's a bit weird.
00:23:12Fire-scape stairs are to be used when there is fire in urgency.
00:23:15And they are the hardest in every building.
00:23:17So we are testing in all those setups.
00:23:20But you can put this in any way.
00:23:21You can disturb it in stairs.
00:23:22It doesn't really matter.
00:23:23This is superhuman capability.
00:23:25Because the robot cannot see it's being pulled.
00:23:27So it's a surprise factor for the robot
00:23:29that you're pulling it while climbing.
00:23:31And again, it's the same model everywhere
00:23:33in all these scenarios.
00:23:35So very easy.
00:23:36Although easy, parkour does look fun.
00:23:40And people often find, like, this is, you know,
00:23:45the robots can do all this.
00:23:47But this is so much easier than what I showed earlier.
00:23:51And even if you see these videos right now,
00:23:53many of you will find this more impressive,
00:23:55even though I'm telling you this is easy.
00:23:57It takes half an hour to train this,
00:23:59while the previous one takes much longer and much more
00:24:01difficult to train.
00:24:03So you can take this base model, and you can transfer it
00:24:06to a variety of tasks very quickly.
00:24:08So here, this is a very old result
00:24:09from like one and a half year ago.
00:24:11For inserting land cables, et cetera,
00:24:13we have very advanced systems now.
00:24:14But we are deploying these models already
00:24:17across variety of application.
00:24:18So for instance, because to build a data flywheel,
00:24:21the hardest part is the deployment itself.
00:24:24There are so many other issues besides intelligence
00:24:26and hardware that come up when you deploy robotics.
00:24:30This is not the same as deploying chat GPT over an app.
00:24:34Because you have to work around many of the--
00:24:36and I think Skydio gave a talk before this,
00:24:38and they can talk to you about all day
00:24:40about the hardness of deployment.
00:24:42So we have to start deployment now,
00:24:44so that we can have this data flywheel in a forcible future.
00:24:47And I'll give you a few examples of deployment we are doing.
00:24:49So one of the examples with NVIDIA--
00:24:51like NVIDIA is opening their first factory in Houston.
00:24:54And their GPUs, what you use right now,
00:24:56they're all built outside US and in Taiwan mainly.
00:24:59And it's all done manually over there.
00:25:01And humans are really efficient and really, really good
00:25:04at making these things.
00:25:05But sustaining cost and scalability and throughput
00:25:08here in US, it's very hard to maintain
00:25:10without this labor force.
00:25:11So here, we are automating this GPU assembly for them.
00:25:15This was-- we did a live demo in NVIDIA GTC.
00:25:19This is deployed in factory already last week,
00:25:21so it's already live.
00:25:22But the interesting part I want to highlight--
00:25:25this is a very traditional factory task.
00:25:29But if you look at the whole setup,
00:25:32this is extremely randomized, extremely noisy.
00:25:36And if you go to any factory, traditional setup,
00:25:38they're extremely clean.
00:25:39So here, the robot, the same brain,
00:25:42can not only do very precise tasks,
00:25:44it can be very robust to any kind of disturbance, which
00:25:46means now these robots do not need to be in a cage.
00:25:50And they can be just deployed as is on these factory lines
00:25:54without any change to any factory line.
00:25:56Alongside humans, and wherever there is humans, there is mess.
00:26:00OK?
00:26:00You can put the same brain to delivery applications.
00:26:02So here, the robot delivers from the truck to the door.
00:26:06It has to find where the front door is.
00:26:08So it's a common sense problem.
00:26:09It's not a mobility problem.
00:26:10And we already have been delivering packages
00:26:12to working behind the scene with many partners
00:26:15to delivering packages to front door of the houses
00:26:17in these areas.
00:26:19Same setup, works across warehouses and all this.
00:26:23Now, just to close the talk, close the topic,
00:26:26I mentioned this is an omnibodied brain.
00:26:29And I hope it's very clear why it's extremely essential
00:26:31to have an omnibodied brain to build a data flywheel.
00:26:34But there's one extra benefit.
00:26:35And the benefit is your robots may break over time
00:26:38and on the safety scenarios.
00:26:41And safety is a byproduct of this omnibodied brain.
00:26:44Because if you have a--
00:26:45we can put the-- I'm just skipping very fast.
00:26:47You can see the online.
00:26:48This is all online.
00:26:49But we can put the same brain across each of these systems,
00:26:52each of these robots.
00:26:55And in this case, we did not even train on these robots.
00:26:58This is completely zero-shot transfer
00:27:00to all these humanoids and all these things.
00:27:02And even if the robot breaks, and here the leg breaks,
00:27:05it can recover within milliseconds.
00:27:07Because when your leg breaks, your three-legged robot
00:27:10is a new robot.
00:27:12So it doesn't really matter what the shape of the robot is.
00:27:14Anything changing, here we disable the legs of the robot.
00:27:17It learns to walk on two legs in just three trials.
00:27:21So what you are seeing on screen is all the training
00:27:24for the robot that is happening in these systems.
00:27:27So it adapts in few milliseconds to 30 seconds or so.
00:27:32And like one example here, it's going on wheels.
00:27:34You jam the wheel.
00:27:36It starts walking.
00:27:37This is another byproduct of omnibodied brain.
00:27:40Safety takes a very different meaning with these kind of models
00:27:45when the robot can fly--
00:27:46or sorry, when the robots can walk or operate
00:27:49with only half the body.
00:27:50I removed some of the Gori scenes from this,
00:27:52but we also cut the robot in half.
00:27:54It can still work with both the two halves separately.
00:27:57But for that, you can go to YouTube.
00:27:59And that's all I have.
00:28:00Thank you.

설명

Show a robotics demo from 1957 next to one from today, and most people can't tell which is which. Deepak Pathak, co-founder and CEO of Skild AI and a professor at Carnegie Mellon, argues robotics stalled because it was treated as a hardware problem rather than a problem of building a general brain. Collecting robot data by teleoperation at one example a minute, he calculates, would take the entire US population more than a century to reach GPT-3 scale. Skild's answer is an "omni-bodied" brain: one model for any robot and any task. It's pre-trained on scalable data like simulation and human video, post-trained on teleoperation, and improved by a deployment flywheel. Pathak shows it inserting AirPods with a simple gripper, learning from human video with under an hour of robot data, and cooking omelets on $4,000 arms with only a camera. He explains why climbing stairs is harder than a backflip. He also shows GPU assembly for NVIDIA's Houston factory, and a robot learning to walk on two legs in three tries after its other legs are disabled. Speaker info: X/Twitter: @pathak2206 (https://x.com/pathak2206) LinkedIn: https://www.linkedin.com/in/pathak22 Website: https://www.cs.cmu.edu/~dpathak Related links: Skild AI: https://www.skild.ai Timestamps: 0:00 AI's progress and the robotics hype 1:02 Robotics has been "almost here" for 70 years 1:27 A 1960s block-copying robot 2:17 Teleoperation in 1957 3:57 Why robotics is stuck: the missing general brain 4:27 Moravec's paradox 5:42 There's no internet of robot data 7:22 Skild's thesis: one brain, any robot, any task 8:22 Teaser: many robots, one brain 9:31 How to scale robot data: a career retrospective 12:01 The three properties of robot data that matter 13:21 Pre-training, post-training and the deployment flywheel 14:56 Why laundry folding is actually easy 16:17 The AirPods task 17:21 Learning from human videos 18:36 The robot's "dream": imagined training data 18:56 Cooking omelets on $4,000 arms 20:41 Truly end-to-end: camera in, motor power out 21:21 Why stairs are harder than backflips 24:50 Deployment: GPU assembly for NVIDIA 26:00 Package delivery to the front door 26:25 Safety: adapting when the robot breaks

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기