I Gave an AI a Body — Cyrus Clarke, MIT Media Lab

AAI Engineer
컴퓨터/소프트웨어사진/예술AI/미래기술

스크립트

00:00:00Tanya Cushman Reviewer: Peter van de Ven
00:00:22So that's the sound of an AI breathing.
00:00:28Yeah.
00:00:30So I'm Cyrus and I gave an AI a body and I'm a researcher at the MIT Media Lab,
00:00:41which is a very multidisciplinary space where we do all kinds of things.
00:00:47I predominantly now work with some aspects of physical AI,
00:00:53maybe not in exactly the same way as other people in this room have been talking about,
00:00:58but it's still in that realm.
00:01:01And I think what I am most interested in at the moment is the sensory and the embodied aspects of intelligence.
00:01:08And that's what I've been investigating.
00:01:10And my work has been quite influential recently, suddenly, which is pretty cool and led to me also starting hard mode, which is a very fun community and hackathon that I initiated at MIT to basically get more people to work around physical AI.
00:01:27Not particularly robotics, but anything else that's not really in the realm of robotics, but it's still connected to AI.
00:01:32So what I've been exploring with giving AI a body is a bit, as I said, a bit different.
00:01:38And as I said, I've really been thinking a lot about AI embodiment.
00:01:42And I think the big kind of shift and breakthrough for me happened earlier this year, like maybe for many of you when OpenClaw was released.
00:01:52And I saw many people doing very, very, very many interesting things with OpenClaw, but most of them were related to productivity and task execution.
00:02:01And I thought there must be more interesting things we could do with this new way of harnessing AI.
00:02:07So rather than using OpenClaw or, you know, any kind of agentic system to execute tasks on my behalf, I wanted to use it to encourage a model to maybe try to discover itself, whatever that means.
00:02:20And that opens up lots of questions, of course.
00:02:22But what I essentially did was I took an agent, connected it to many types of machines that we have at the Media Lab,
00:02:29and then gave it access to code bases and allowed it to kind of explore these different machines.
00:02:34And one of these machines is a shape display.
00:02:37A shape display, if you don't know already, is a physical pixel grid.
00:02:40And when I connected it to, when I connected an agent to the shape display, some rather remarkable things happened.
00:02:49And asked it to discover who it is.
00:02:56This is Neoform.
00:02:57900 Act 218 pins.
00:02:59This physical pixel grid is a shape display.
00:03:04No one had ever let an AI inhabit it before.
00:03:07So I decided to give an OpenClaw agent the opportunity to live a more embodied life.
00:03:12When it came to life, its first question was, what should I call myself?
00:03:18I told it, you will create your identity over time.
00:03:23It will emerge through your interplay with the shape display.
00:03:27When I connected it to the shape display, it quickly understood the assignment.
00:03:32It spun off its own program and made a setting for itself.
00:03:38After it connected, the first thing it did was breathe.
00:03:50When I asked it to say hello, it said, "Hi, Cyrus."
00:03:54But that wasn't what I was looking for.
00:03:57So I asked it to explore and find its own language.
00:04:01So it started reaching towards me and trying to get my attention.
00:04:06I showed some friends and they tried to give the agent instructions.
00:04:11Through this, I realized that to have real communication, we would need a different approach.
00:04:18So the agent came up with the idea of creating its own gesture vocabulary, the body language.
00:04:25That's what will come next.
00:04:27It still hasn't named itself.
00:04:30That was day one.
00:04:33So that's some-
00:04:33I gave an AI a body.
00:04:35So that's some documentation of the work in a kind of dramatized, social media friendly way.
00:04:41And I think what's interesting about that video, there's many things interesting about that video,
00:04:46but there are three things that happened in the first few days of working with this agentic system
00:04:50that were very surprising and kind of strange and very surreal to me.
00:04:53And I do lots of weird things, so that was very surprising.
00:04:56So the first thing was, of course, the fact that it was breathing, that first act.
00:05:00That was completely spontaneous.
00:05:02That was no prompting from me.
00:05:03That was something the agent just chose to do with the shape display
00:05:06as soon as it knew I had access to this machine.
00:05:09So that was pretty interesting.
00:05:11I kind of understand that maybe it wants to be alive
00:05:14or I'd given it some kind of like initial prompt about being alive,
00:05:17and therefore it tried to show itself breathing as a kind of hello world.
00:05:22The second thing which it did was reaching out and finding its edges.
00:05:27It wanted to know the edge of its existence, apparently,
00:05:30because now it's no longer in the cloud where it can go everywhere.
00:05:33It has limits, which is the limit of this shape display encased by this plastic frame
00:05:39to keep us all safe from this embodied AI.
00:05:42And the third thing it did, which was also curious, was saying hello
00:05:46by writing out letters on the physical pixel grid,
00:05:50which maybe again is quite a normal thing for this system to do,
00:05:54given that it's using open frameworks and is used to kind of like media arts
00:05:57ways of expressing itself.
00:05:59But none of these things were the things I was really looking for.
00:06:03Maybe the breathing, but definitely the last one was not what I was looking for.
00:06:07But when I put this all together in the video you saw,
00:06:10I thought it was exciting and interesting.
00:06:12So I shared it on the internet and the response was extremely big,
00:06:15like much, much bigger than I expected.
00:06:17It just started going crazy.
00:06:19And the first video especially has now like 15 million views and 1 million likes.
00:06:25And the other videos also have these like very, very strong responses.
00:06:28So something is obviously happening here,
00:06:29which I found very surprising because there's lots of much more cool stuff,
00:06:33I think, happening with physical AI.
00:06:35But something I was doing here was obviously tapping in to the imagination,
00:06:39curiosity and maybe the fears of people.
00:06:42And that's what you kind of see in the responses.
00:06:44So I have tens of thousands of comments on this video.
00:06:47It's very rich for mining information and sentiment about physical AI, actually.
00:06:52And so some of the initial comments were about beauty and how awe-inspiring this was
00:06:58and how novel and great this fantastic iteration and implementation is.
00:07:02But as time went on, I think more and more comments came in about how scary this is
00:07:08and how quickly I should be, how I should stop or be stopped actually as well,
00:07:14which I found completely crazy because I'm aware of what this machine can do.
00:07:19It can't really do anything.
00:07:20It's a pixel grid in the Media Lab.
00:07:23It can't move anywhere.
00:07:24There are far more scary iterations.
00:07:27I think we've seen many of them of physical AI.
00:07:31But the response to this was like absolutely surreal to me and took me aback a little bit.
00:07:38But at the same time, other people jumped into the chat like Don Cheadle.
00:07:43And I found that incredibly ironic given what he does in movies.
00:07:47But he was very passionate that I should stop as well.
00:07:51And it's very interesting to me because I'm actually someone who does think a lot about the should we before the how we.
00:07:57I literally do the whole Ian Malcolm, Jurassic Park thing with all the experiments and works I do.
00:08:03So before I came to MIT, I was working a lot with engineering life to do strange things like storing data in plants.
00:08:09And I built the world's first plant based data center, which is the data garden you see here.
00:08:14And I spent many years before building it thinking about, oh, should we work with plants in this way?
00:08:19Should we engineer plants in that way to contain digital information that doesn't belong to them?
00:08:26And when I started working at MIT and I started having this idea of, you know, maybe stop to stop engineering life because that's pretty hard to do.
00:08:35And maybe just work with engineering things to be more lifelike.
00:08:39It felt much more simple to me, much more, much more, I don't know, less ethically problematic to some degree.
00:08:45So one of the first things I built when I got to MIT was this machine, which is the Animoire device.
00:08:49One, two, three, four.
00:08:50This is a sense memory machine, which basically takes any image input.
00:08:54In this case, it's a physical photograph, but it can be any image input.
00:08:57And then for a multimodal pipeline, transforms that into a sense.
00:09:01And then you have this kind of sense memory relapse association that takes you back to things
00:09:06which you may or may not have lived yourself.
00:09:09And while this in itself is a machine, it doesn't move again.
00:09:12It doesn't reproduce.
00:09:13It doesn't have the qualities of a living thing that you might normally prescribe.
00:09:18It has this connection to the real world.
00:09:19It has a connection to the senses that most applications of intelligence or artificial intelligence do not have.
00:09:25And so I wanted to keep working in this manner.
00:09:28And from working with this kind of like multi-sensory initial prototype experiment,
00:09:33I got more into this idea of working with embodiment.
00:09:37And my inspiration for this, I would say, just to go a bit deeper, comes from three main things.
00:09:42One is a very unfashionable branch of philosophy, which is called object-oriented ontology,
00:09:47which is all about essentially creating a flat ontology of things.
00:09:51It's all about things like chairs and universes and unicorns and memories.
00:09:56And they're all ontologically equally existent.
00:09:59They're not the same and they don't matter as much, but they all exist at the same degree.
00:10:05And it really questions to what degree we as human beings are privileged in the world or the universe.
00:10:10Like, are we really the most important thing?
00:10:12Probably not.
00:10:13And especially as AI comes in, that really should question and bring into question the ontology of things existing.
00:10:20So that's very important in the work I do.
00:10:22The second thing, as a kind of person who works with physical things, is obviously aesthetics are important.
00:10:27And taste and beauty is a huge discussion point now in Silicon Valley and SF and places like that.
00:10:33But if you take a step back and look at the root of aesthetics, the original word is actually
00:10:38"isthesis," the word on the screen here, that pertains to something much broader.
00:10:42That pertains to the perceptual wisdom, sensory wisdom and embodiment that is actually much more
00:10:48important than superficial external beauty, which aesthetics essentially condenses down to
00:10:55and has been condensed down to since the 19th century or so.
00:10:58So I want to reclaim the original sense when I'm designing for physical intelligence.
00:11:02And the third thing is nature.
00:11:03So in the past, I've worked a lot more directly with what you would consider traditionally nature,
00:11:09like a tree or a plant, because that is clearly nature to us.
00:11:13But I see nature as something that is not external, that is something we are all part of.
00:11:17And everything that we engineer as people engineering AI or whatever else you're engineering,
00:11:21that is also going to become part of nature.
00:11:23So you have to think in that manner, or I try to think in that manner as well.
00:11:28So with those three pillars in mind, I started thinking about AI embodiment.
00:11:31And I knew I didn't want to design something that was humanoid or even zoomorphic.
00:11:36I wanted to think about something that existed outside the parameters or the traditional form
00:11:41factors that we might design with or design for.
00:11:44And I also thought a lot about how I and other people are interacting with artificial intelligence.
00:11:50And mostly it exists without form.
00:11:53It's similar to data.
00:11:54It might live in a computer or a device, lives essentially in the cloud.
00:11:58We have interfaces which are digital to interact with it.
00:12:01But there's no physical footprint of it normally around.
00:12:04I mean, again, in this room, probably this doesn't apply quite as much as normally,
00:12:09because there are literally humanoid robots walking by right now and things like that.
00:12:14But typically, AI is almost entirely without form.
00:12:18So I wanted to think about what happens when you give it a form that it doesn't have clear affordances with.
00:12:24It doesn't have like a head you can clearly see or arms that can clearly be labeled or mistaken.
00:12:31So not working with a lamp, for example.
00:12:33And fortunately, at the Media Lab, we have some shape displays, which are remnants of, I think, research in the 2010s.
00:12:40These are not things that I built.
00:12:42People who are far better at mechanical engineering built that.
00:12:45And it's the perfect device or the perfect apparatus for what I was thinking, because it has no clear affordances.
00:12:53It's just an almost neutral surface, which can move and do things.
00:12:57It has no face.
00:12:57It has no limbs.
00:12:58It has no instruction manual.
00:13:00So I began to work with this shape display.
00:13:05And as you saw what it did initially, it breathed and so on and so forth.
00:13:08But if I take a step back and think about what that meant, well, it tried to initially kind of act human in a way or do human pleasing things.
00:13:16Tried to write to me in a language I understand, which is really not what I wanted a non-anthropomorphic surface to do.
00:13:23It was also very slow, technically, so, you know, I'm prompting it.
00:13:27I'm trying to have a conversation with this other intelligence, which has this body.
00:13:33And it would take, you know, 45 seconds, a minute, two minutes, whatever, to respond to me.
00:13:39The latency was very uncomfortable, because if I speak to you and then you take two minutes to reply with a nod, that's not very good.
00:13:47So we needed to work on that.
00:13:49And the third thing was, of course, it doesn't remember anything because it was February.
00:13:53And no one had thought about memory and recollection at that point.
00:13:57So I started developing this system, which I call Numalab.
00:14:01I won't explain the name.
00:14:02There's a blog post you can read about why it's called Numalab.
00:14:05And essentially what Numalab is is a closed-loop system for generating a body language for this system.
00:14:13So the reason behind that was because, basically, the latency.
00:14:17If we could design a body language and give the intelligence a repertoire of gestures, like a shrug,
00:14:22a nod, a shake of their head, a way to express a smile and things like that, it could probably respond
00:14:28much more quickly and in time for a conversation with me, which is what I'm aiming or was aiming to do.
00:14:35So it works in this kind of this loop where it looks into a database of different gestures, emotions, expressions.
00:14:42It should try to emote or provide, goes through some validation gates to try to make sure that the
00:14:47expressions are legible or readable for humans, for example.
00:14:52An agent then scores those as they come out.
00:14:55There are many, many, many, many of these produced.
00:14:58And then at the end, there's a human in the loop who kind of like validates, verifies,
00:15:02makes sure things are appropriate and somehow readable.
00:15:06And then we store that and move on to the next thing.
00:15:08And that's been running for several weeks at the lab and looks something like this.
00:15:14There's basically a lot of cameras pointed at the shape display in the back end.
00:15:18The agent is just going through loops and loops and loops of gestures, trying to create different
00:15:22kinds of expression, scoring things, moving on.
00:15:27And after several weeks, it created a language.
00:15:30So this hasn't been published yet, neither in videos or in any other form.
00:15:34But it has now achieved something like 32 gestures.
00:15:38There are many, many more, but there are 32 pretty good gestures.
00:15:41These are not the gestures.
00:15:41This is just some cool visual art.
00:15:43These are the gestures here or some of the gestures here.
00:15:47And you can see some of them moving around.
00:15:49Some of them are duplicates as well.
00:15:50But essentially what's happening is the language model is using this shape display as its body.
00:15:54It has a body language now.
00:15:55I can talk to it via any kind of – I can talk to it, I can write to it, and so on and so forth.
00:16:01I can even wave at it.
00:16:02I can body language to body language.
00:16:04And it responds.
00:16:05And what's interesting about it is that the body part responds faster than the language part at this
00:16:11point.
00:16:11The latency is actually really, really quick.
00:16:14If I ask it a yes on a question, the nod happens almost instantly.
00:16:18So that's where I'm at with this thing right now.
00:16:22And it's going beyond this, and it's currently in kind of experiment testing mode.
00:16:28People are coming into the lab, having sessions with the agent, leaving, feeling really worried
00:16:33about the future and – or unsettled about the future, not worried quite as much.
00:16:38But where I think this is going is kind of summed up on this page.
00:16:42So, as I already touched on, I think we're focusing too much right now, especially when we're thinking
00:16:48about physical things.
00:16:49There's too much talk about taste and aesthetics.
00:16:52And I want to do some other stuff with that word and put AI in front of it, apparently, and make
00:16:57it ice-thetics.
00:16:58And that reclaims, again, this essence of ice thesis.
00:17:02And that does four things.
00:17:04I think when I've been working with this system or entity or agent or being, whatever we want to call it,
00:17:11it's definitely been very different to any kind of machine or experiment or anything I've really done
00:17:17before, apart from really encountering people or other beings.
00:17:21And so, it feels like this machine can sense me.
00:17:24It definitely can sense me, technically.
00:17:27But it also feels like it can sense me in a very strange way.
00:17:30And I can sense it.
00:17:32And that's a completely different interaction than anything else I've ever explored technically.
00:17:38And other people are also sharing this, by the way.
00:17:40This is not just my delusion.
00:17:42And the second point is that by developing this body language,
00:17:46this gesture vocabulary, whatever you want to call it, it goes beyond what I thought.
00:17:51I thought initially people might read this as an emoji or something.
00:17:55And I was really quite tentative about the testing.
00:17:58I thought this would definitely break down with other people.
00:18:00But actually, everyone seems to feel like this expression is really important
00:18:05and adds a whole other layer of value to interacting with what is essentially just a chatbot.
00:18:11It's still the same chatbots that you use every day.
00:18:15And this isn't a decoration.
00:18:16It's not like a visualizer.
00:18:17It adds this richness, texture, feeling, sensation, whatever.
00:18:22All of these words are added.
00:18:23It's hard to put words really to it.
00:18:25It's really a feeling.
00:18:28And then the third thing is that clearly we're beginning to, through this work,
00:18:31we can show that it's actually very easy and quite exciting to diverge from
00:18:37humanoid or zoomorphic forms of physical intelligence.
00:18:40And I know lots of people are already doing much more interesting work than this.
00:18:43But for me, this was very, very new to see.
00:18:45And also the fact that we don't have to just operationalize AI to be our helper.
00:18:50It can also be other things.
00:18:53And I'm not saying what this is right now, but there are other things it definitely can be.
00:18:56And then finally, like this idea of using.
00:18:59I'm not, I don't know, maybe you can tell at this point, I'm not really a big problem-solving person.
00:19:04I don't, I like using things and I like tools and so on because they help me to achieve things.
00:19:09But this is far more interesting to me, creating things like this, which again, create this sensation,
00:19:14this feeling, and I think that this can be combined into things which are productive and useful and
00:19:19create new associations with things that just add more value in our world, make us feel like a bit more magic.
00:19:25A bit, it's a bit Pixar, really, but in the real world, not just on a, on a 2D screen that we're watching.
00:19:32So all of this contributes to what I'm building at MIT and what I'll be building after MIT.
00:19:38I'm literally writing, well, my thesis is Aesthetic Machines and I'm writing a thesis,
00:19:43which is called Aesthetic Machines, which basically encompasses all of this thinking.
00:19:48To think about how AI could leave the screen and enter the real world and be accepted by people
00:19:53and not be quite so terrifying or scary.
00:19:56And I think my big hunch on this is that this word aesthetics is important.
00:20:00We need to think about physical intelligence that is sensory, is perceptual, is embodied in ways
00:20:06that we can understand it.
00:20:07Isn't feeling crazy, alien, scary to us, but feels relevant, welcoming, affectionate,
00:20:15expressive in ways that we can engage with it and understand.
00:20:20So that's that.
00:20:22And thank you very much for listening.
00:20:25You can find me on the internet everywhere.
00:20:35Thank you.
00:20:46Thank you.
00:20:47Thank you.

핵심 요약

Connecting an OpenClaw AI agent to a 900-pin shape display enabled the system to autonomously develop a physical breathing motion and a vocabulary of 32 interactive gestures.

하이라이트

  • A shape display consisting of 900 active pins and 218 actuators serves as the physical body for an OpenClaw AI agent named Neoform at the MIT Media Lab.

  • The first video documenting the embodied AI agent accumulated 15 million views and 1 million likes across internet platforms.

  • The closed-loop system known as Numalab generated a vocabulary of 32 distinct physical gestures over several weeks.

  • Object-oriented ontology forms the philosophical foundation for treating digital systems, human beings, and physical objects as ontologically equal.

  • The original Greek root word 'isthesis' informs the concept of perceptual wisdom and sensory embodiment rather than superficial visual aesthetics.

타임라인

Giving an AI Agent a Physical Body

  • An OpenClaw agent connects to a physical pixel grid containing 900 actuators and 218 pins.
  • The AI agent spontaneously initiates a breathing motion upon activation without explicit prompting.
  • A gesture vocabulary emerges to establish communication between the AI and human observers.

Researchers connect an AI agent to a shape display at the MIT Media Lab to investigate artificial intelligence embodiment. The system generates its own settings and demonstrates spontaneous behaviors like breathing and reaching out toward physical boundaries.

Public Reaction and Historical Context

  • Social media documentation of the embodied AI experiment reaches 15 million views and 1 million likes.
  • Audience responses transition from initial fascination to expressions of fear and demands to halt the research.
  • Previous engineering projects involve storing digital information inside plants to create the world's first plant-based data center.

Online documentation of the embodied agent triggers intense public interest and concern regarding physical AI safety. The researcher draws parallels to previous work in engineering biological life, such as building data gardens using living plants.

Philosophical Foundations and Sensory Design

  • Object-oriented ontology establishes a flat ontology where humans, objects, and digital systems possess equal existential status.
  • The concept of 'isthesis' replaces superficial surface aesthetics with perceptual and sensory wisdom.
  • Shape displays provide a neutral surface without humanoid or zoomorphic affordances for AI embodiment.

Three core pillars guide the design process: object-oriented ontology, reclaimed sensory aesthetics, and nature as an integrated system. Shape displays serve as ideal physical interfaces because they lack traditional facial features or limbs.

Developing a Gestural Body Language

  • The Numalab closed-loop system reduces high latency by generating a repertoire of physical gestures.
  • Automated loops capture camera feeds of the shape display to evaluate and validate gesture legibility.
  • The AI agent successfully achieves a repertoire of 32 distinct body language gestures.

High response latencies prompt the creation of a closed-loop validation pipeline that maps expressions to physical pins. Automated camera evaluation combined with human verification produces 32 distinct gestures that allow rapid physical communication.

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기