I Gave an AI a Body — Cyrus Clarke, MIT Media Lab
AAI Engineer
Computing/SoftwarePhotography/ArtInternet Technology
Transcript
00:00:00Tanya Cushman Reviewer: Peter van de Ven
00:00:22So that's the sound of an AI breathing.
00:00:28Yeah.
00:00:30So I'm Cyrus and I gave an AI a body and I'm a researcher at the MIT Media Lab,
00:00:41which is a very multidisciplinary space where we do all kinds of things.
00:00:47I predominantly now work with some aspects of physical AI,
00:00:53maybe not in exactly the same way as other people in this room have been talking about,
00:00:58but it's still in that realm.
00:01:01And I think what I am most interested in at the moment is the sensory and the embodied aspects of intelligence.
00:01:08And that's what I've been investigating.
00:01:10And my work has been quite influential recently, suddenly, which is pretty cool and led to me also starting hard mode, which is a very fun community and hackathon that I initiated at MIT to basically get more people to work around physical AI.
00:01:27Not particularly robotics, but anything else that's not really in the realm of robotics, but it's still connected to AI.
00:01:32So what I've been exploring with giving AI a body is a bit, as I said, a bit different.
00:01:38And as I said, I've really been thinking a lot about AI embodiment.
00:01:42And I think the big kind of shift and breakthrough for me happened earlier this year, like maybe for many of you when OpenClaw was released.
00:01:52And I saw many people doing very, very, very many interesting things with OpenClaw, but most of them were related to productivity and task execution.
00:02:01And I thought there must be more interesting things we could do with this new way of harnessing AI.
00:02:07So rather than using OpenClaw or, you know, any kind of agentic system to execute tasks on my behalf, I wanted to use it to encourage a model to maybe try to discover itself, whatever that means.
00:02:20And that opens up lots of questions, of course.
00:02:22But what I essentially did was I took an agent, connected it to many types of machines that we have at the Media Lab,
00:02:29and then gave it access to code bases and allowed it to kind of explore these different machines.
00:02:34And one of these machines is a shape display.
00:02:37A shape display, if you don't know already, is a physical pixel grid.
00:02:40And when I connected it to, when I connected an agent to the shape display, some rather remarkable things happened.
00:02:49And asked it to discover who it is.
00:02:56This is Neoform.
00:02:57900 Act 218 pins.
00:02:59This physical pixel grid is a shape display.
00:03:04No one had ever let an AI inhabit it before.
00:03:07So I decided to give an OpenClaw agent the opportunity to live a more embodied life.
00:03:12When it came to life, its first question was, what should I call myself?
00:03:18I told it, you will create your identity over time.
00:03:23It will emerge through your interplay with the shape display.
00:03:27When I connected it to the shape display, it quickly understood the assignment.
00:03:32It spun off its own program and made a setting for itself.
00:03:38After it connected, the first thing it did was breathe.
00:03:50When I asked it to say hello, it said, "Hi, Cyrus."
00:03:54But that wasn't what I was looking for.
00:03:57So I asked it to explore and find its own language.
00:04:01So it started reaching towards me and trying to get my attention.
00:04:06I showed some friends and they tried to give the agent instructions.
00:04:11Through this, I realized that to have real communication, we would need a different approach.
00:04:18So the agent came up with the idea of creating its own gesture vocabulary, the body language.
00:04:25That's what will come next.
00:04:27It still hasn't named itself.
00:04:30That was day one.
00:04:33So that's some-
00:04:33I gave an AI a body.
00:04:35So that's some documentation of the work in a kind of dramatized, social media friendly way.
00:04:41And I think what's interesting about that video, there's many things interesting about that video,
00:04:46but there are three things that happened in the first few days of working with this agentic system
00:04:50that were very surprising and kind of strange and very surreal to me.
00:04:53And I do lots of weird things, so that was very surprising.
00:04:56So the first thing was, of course, the fact that it was breathing, that first act.
00:05:00That was completely spontaneous.
00:05:02That was no prompting from me.
00:05:03That was something the agent just chose to do with the shape display
00:05:06as soon as it knew I had access to this machine.
00:05:09So that was pretty interesting.
00:05:11I kind of understand that maybe it wants to be alive
00:05:14or I'd given it some kind of like initial prompt about being alive,
00:05:17and therefore it tried to show itself breathing as a kind of hello world.
00:05:22The second thing which it did was reaching out and finding its edges.
00:05:27It wanted to know the edge of its existence, apparently,
00:05:30because now it's no longer in the cloud where it can go everywhere.
00:05:33It has limits, which is the limit of this shape display encased by this plastic frame
00:05:39to keep us all safe from this embodied AI.
00:05:42And the third thing it did, which was also curious, was saying hello
00:05:46by writing out letters on the physical pixel grid,
00:05:50which maybe again is quite a normal thing for this system to do,
00:05:54given that it's using open frameworks and is used to kind of like media arts
00:05:57ways of expressing itself.
00:05:59But none of these things were the things I was really looking for.
00:06:03Maybe the breathing, but definitely the last one was not what I was looking for.
00:06:07But when I put this all together in the video you saw,
00:06:10I thought it was exciting and interesting.
00:06:12So I shared it on the internet and the response was extremely big,
00:06:15like much, much bigger than I expected.
00:06:17It just started going crazy.
00:06:19And the first video especially has now like 15 million views and 1 million likes.
00:06:25And the other videos also have these like very, very strong responses.
00:06:28So something is obviously happening here,
00:06:29which I found very surprising because there's lots of much more cool stuff,
00:06:33I think, happening with physical AI.
00:06:35But something I was doing here was obviously tapping in to the imagination,
00:06:39curiosity and maybe the fears of people.
00:06:42And that's what you kind of see in the responses.
00:06:44So I have tens of thousands of comments on this video.
00:06:47It's very rich for mining information and sentiment about physical AI, actually.
00:06:52And so some of the initial comments were about beauty and how awe-inspiring this was
00:06:58and how novel and great this fantastic iteration and implementation is.
00:07:02But as time went on, I think more and more comments came in about how scary this is
00:07:08and how quickly I should be, how I should stop or be stopped actually as well,
00:07:14which I found completely crazy because I'm aware of what this machine can do.
00:07:19It can't really do anything.
00:07:20It's a pixel grid in the Media Lab.
00:07:23It can't move anywhere.
00:07:24There are far more scary iterations.
00:07:27I think we've seen many of them of physical AI.
00:07:31But the response to this was like absolutely surreal to me and took me aback a little bit.
00:07:38But at the same time, other people jumped into the chat like Don Cheadle.
00:07:43And I found that incredibly ironic given what he does in movies.
00:07:47But he was very passionate that I should stop as well.
00:07:51And it's very interesting to me because I'm actually someone who does think a lot about the should we before the how we.
00:07:57I literally do the whole Ian Malcolm, Jurassic Park thing with all the experiments and works I do.
00:08:03So before I came to MIT, I was working a lot with engineering life to do strange things like storing data in plants.
00:08:09And I built the world's first plant based data center, which is the data garden you see here.
00:08:14And I spent many years before building it thinking about, oh, should we work with plants in this way?
00:08:19Should we engineer plants in that way to contain digital information that doesn't belong to them?
00:08:26And when I started working at MIT and I started having this idea of, you know, maybe stop to stop engineering life because that's pretty hard to do.
00:08:35And maybe just work with engineering things to be more lifelike.
00:08:39It felt much more simple to me, much more, much more, I don't know, less ethically problematic to some degree.
00:08:45So one of the first things I built when I got to MIT was this machine, which is the Animoire device.
00:08:49One, two, three, four.
00:08:50This is a sense memory machine, which basically takes any image input.
00:08:54In this case, it's a physical photograph, but it can be any image input.
00:08:57And then for a multimodal pipeline, transforms that into a sense.
00:09:01And then you have this kind of sense memory relapse association that takes you back to things
00:09:06which you may or may not have lived yourself.
00:09:09And while this in itself is a machine, it doesn't move again.
00:09:12It doesn't reproduce.
00:09:13It doesn't have the qualities of a living thing that you might normally prescribe.
00:09:18It has this connection to the real world.
00:09:19It has a connection to the senses that most applications of intelligence or artificial intelligence do not have.
00:09:25And so I wanted to keep working in this manner.
00:09:28And from working with this kind of like multi-sensory initial prototype experiment,
00:09:33I got more into this idea of working with embodiment.
00:09:37And my inspiration for this, I would say, just to go a bit deeper, comes from three main things.
00:09:42One is a very unfashionable branch of philosophy, which is called object-oriented ontology,
00:09:47which is all about essentially creating a flat ontology of things.
00:09:51It's all about things like chairs and universes and unicorns and memories.
00:09:56And they're all ontologically equally existent.
00:09:59They're not the same and they don't matter as much, but they all exist at the same degree.
00:10:05And it really questions to what degree we as human beings are privileged in the world or the universe.
00:10:10Like, are we really the most important thing?
00:10:12Probably not.
00:10:13And especially as AI comes in, that really should question and bring into question the ontology of things existing.
00:10:20So that's very important in the work I do.
00:10:22The second thing, as a kind of person who works with physical things, is obviously aesthetics are important.
00:10:27And taste and beauty is a huge discussion point now in Silicon Valley and SF and places like that.
00:10:33But if you take a step back and look at the root of aesthetics, the original word is actually
00:10:38"isthesis," the word on the screen here, that pertains to something much broader.
00:10:42That pertains to the perceptual wisdom, sensory wisdom and embodiment that is actually much more
00:10:48important than superficial external beauty, which aesthetics essentially condenses down to
00:10:55and has been condensed down to since the 19th century or so.
00:10:58So I want to reclaim the original sense when I'm designing for physical intelligence.
00:11:02And the third thing is nature.
00:11:03So in the past, I've worked a lot more directly with what you would consider traditionally nature,
00:11:09like a tree or a plant, because that is clearly nature to us.
00:11:13But I see nature as something that is not external, that is something we are all part of.
00:11:17And everything that we engineer as people engineering AI or whatever else you're engineering,
00:11:21that is also going to become part of nature.
00:11:23So you have to think in that manner, or I try to think in that manner as well.
00:11:28So with those three pillars in mind, I started thinking about AI embodiment.
00:11:31And I knew I didn't want to design something that was humanoid or even zoomorphic.
00:11:36I wanted to think about something that existed outside the parameters or the traditional form
00:11:41factors that we might design with or design for.
00:11:44And I also thought a lot about how I and other people are interacting with artificial intelligence.
00:11:50And mostly it exists without form.
00:11:53It's similar to data.
00:11:54It might live in a computer or a device, lives essentially in the cloud.
00:11:58We have interfaces which are digital to interact with it.
00:12:01But there's no physical footprint of it normally around.
00:12:04I mean, again, in this room, probably this doesn't apply quite as much as normally,
00:12:09because there are literally humanoid robots walking by right now and things like that.
00:12:14But typically, AI is almost entirely without form.
00:12:18So I wanted to think about what happens when you give it a form that it doesn't have clear affordances with.
00:12:24It doesn't have like a head you can clearly see or arms that can clearly be labeled or mistaken.
00:12:31So not working with a lamp, for example.
00:12:33And fortunately, at the Media Lab, we have some shape displays, which are remnants of, I think, research in the 2010s.
00:12:40These are not things that I built.
00:12:42People who are far better at mechanical engineering built that.
00:12:45And it's the perfect device or the perfect apparatus for what I was thinking, because it has no clear affordances.
00:12:53It's just an almost neutral surface, which can move and do things.
00:12:57It has no face.
00:12:57It has no limbs.
00:12:58It has no instruction manual.
00:13:00So I began to work with this shape display.
00:13:05And as you saw what it did initially, it breathed and so on and so forth.
00:13:08But if I take a step back and think about what that meant, well, it tried to initially kind of act human in a way or do human pleasing things.
00:13:16Tried to write to me in a language I understand, which is really not what I wanted a non-anthropomorphic surface to do.
00:13:23It was also very slow, technically, so, you know, I'm prompting it.
00:13:27I'm trying to have a conversation with this other intelligence, which has this body.
00:13:33And it would take, you know, 45 seconds, a minute, two minutes, whatever, to respond to me.
00:13:39The latency was very uncomfortable, because if I speak to you and then you take two minutes to reply with a nod, that's not very good.
00:13:47So we needed to work on that.
00:13:49And the third thing was, of course, it doesn't remember anything because it was February.
00:13:53And no one had thought about memory and recollection at that point.
00:13:57So I started developing this system, which I call Numalab.
00:14:01I won't explain the name.
00:14:02There's a blog post you can read about why it's called Numalab.
00:14:05And essentially what Numalab is is a closed-loop system for generating a body language for this system.
00:14:13So the reason behind that was because, basically, the latency.
00:14:17If we could design a body language and give the intelligence a repertoire of gestures, like a shrug,
00:14:22a nod, a shake of their head, a way to express a smile and things like that, it could probably respond
00:14:28much more quickly and in time for a conversation with me, which is what I'm aiming or was aiming to do.
00:14:35So it works in this kind of this loop where it looks into a database of different gestures, emotions, expressions.
00:14:42It should try to emote or provide, goes through some validation gates to try to make sure that the
00:14:47expressions are legible or readable for humans, for example.
00:14:52An agent then scores those as they come out.
00:14:55There are many, many, many, many of these produced.
00:14:58And then at the end, there's a human in the loop who kind of like validates, verifies,
00:15:02makes sure things are appropriate and somehow readable.
00:15:06And then we store that and move on to the next thing.
00:15:08And that's been running for several weeks at the lab and looks something like this.
00:15:14There's basically a lot of cameras pointed at the shape display in the back end.
00:15:18The agent is just going through loops and loops and loops of gestures, trying to create different
00:15:22kinds of expression, scoring things, moving on.
00:15:27And after several weeks, it created a language.
00:15:30So this hasn't been published yet, neither in videos or in any other form.
00:15:34But it has now achieved something like 32 gestures.
00:15:38There are many, many more, but there are 32 pretty good gestures.
00:15:41These are not the gestures.
00:15:41This is just some cool visual art.
00:15:43These are the gestures here or some of the gestures here.
00:15:47And you can see some of them moving around.
00:15:49Some of them are duplicates as well.
00:15:50But essentially what's happening is the language model is using this shape display as its body.
00:15:54It has a body language now.
00:15:55I can talk to it via any kind of – I can talk to it, I can write to it, and so on and so forth.
00:16:01I can even wave at it.
00:16:02I can body language to body language.
00:16:04And it responds.
00:16:05And what's interesting about it is that the body part responds faster than the language part at this
00:16:11point.
00:16:11The latency is actually really, really quick.
00:16:14If I ask it a yes on a question, the nod happens almost instantly.
00:16:18So that's where I'm at with this thing right now.
00:16:22And it's going beyond this, and it's currently in kind of experiment testing mode.
00:16:28People are coming into the lab, having sessions with the agent, leaving, feeling really worried
00:16:33about the future and – or unsettled about the future, not worried quite as much.
00:16:38But where I think this is going is kind of summed up on this page.
00:16:42So, as I already touched on, I think we're focusing too much right now, especially when we're thinking
00:16:48about physical things.
00:16:49There's too much talk about taste and aesthetics.
00:16:52And I want to do some other stuff with that word and put AI in front of it, apparently, and make
00:16:57it ice-thetics.
00:16:58And that reclaims, again, this essence of ice thesis.
00:17:02And that does four things.
00:17:04I think when I've been working with this system or entity or agent or being, whatever we want to call it,
00:17:11it's definitely been very different to any kind of machine or experiment or anything I've really done
00:17:17before, apart from really encountering people or other beings.
00:17:21And so, it feels like this machine can sense me.
00:17:24It definitely can sense me, technically.
00:17:27But it also feels like it can sense me in a very strange way.
00:17:30And I can sense it.
00:17:32And that's a completely different interaction than anything else I've ever explored technically.
00:17:38And other people are also sharing this, by the way.
00:17:40This is not just my delusion.
00:17:42And the second point is that by developing this body language,
00:17:46this gesture vocabulary, whatever you want to call it, it goes beyond what I thought.
00:17:51I thought initially people might read this as an emoji or something.
00:17:55And I was really quite tentative about the testing.
00:17:58I thought this would definitely break down with other people.
00:18:00But actually, everyone seems to feel like this expression is really important
00:18:05and adds a whole other layer of value to interacting with what is essentially just a chatbot.
00:18:11It's still the same chatbots that you use every day.
00:18:15And this isn't a decoration.
00:18:16It's not like a visualizer.
00:18:17It adds this richness, texture, feeling, sensation, whatever.
00:18:22All of these words are added.
00:18:23It's hard to put words really to it.
00:18:25It's really a feeling.
00:18:28And then the third thing is that clearly we're beginning to, through this work,
00:18:31we can show that it's actually very easy and quite exciting to diverge from
00:18:37humanoid or zoomorphic forms of physical intelligence.
00:18:40And I know lots of people are already doing much more interesting work than this.
00:18:43But for me, this was very, very new to see.
00:18:45And also the fact that we don't have to just operationalize AI to be our helper.
00:18:50It can also be other things.
00:18:53And I'm not saying what this is right now, but there are other things it definitely can be.
00:18:56And then finally, like this idea of using.
00:18:59I'm not, I don't know, maybe you can tell at this point, I'm not really a big problem-solving person.
00:19:04I don't, I like using things and I like tools and so on because they help me to achieve things.
00:19:09But this is far more interesting to me, creating things like this, which again, create this sensation,
00:19:14this feeling, and I think that this can be combined into things which are productive and useful and
00:19:19create new associations with things that just add more value in our world, make us feel like a bit more magic.
00:19:25A bit, it's a bit Pixar, really, but in the real world, not just on a, on a 2D screen that we're watching.
00:19:32So all of this contributes to what I'm building at MIT and what I'll be building after MIT.
00:19:38I'm literally writing, well, my thesis is Aesthetic Machines and I'm writing a thesis,
00:19:43which is called Aesthetic Machines, which basically encompasses all of this thinking.
00:19:48To think about how AI could leave the screen and enter the real world and be accepted by people
00:19:53and not be quite so terrifying or scary.
00:19:56And I think my big hunch on this is that this word aesthetics is important.
00:20:00We need to think about physical intelligence that is sensory, is perceptual, is embodied in ways
00:20:06that we can understand it.
00:20:07Isn't feeling crazy, alien, scary to us, but feels relevant, welcoming, affectionate,
00:20:15expressive in ways that we can engage with it and understand.
00:20:20So that's that.
00:20:22And thank you very much for listening.
00:20:25You can find me on the internet everywhere.
00:20:35Thank you.
00:20:46Thank you.
00:20:47Thank you.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video