Tell the Robot What You Want — Sandhya Subramani, AWS
AAI Engineer
Computing/SoftwareConsumer Electronics
Transcript
00:00:00Hello, everyone. My name is Sandy. And meet my co-host today, Scout. This is my friendly
00:00:21rover. And one would think that rovers can't really think for themselves, right? We have
00:00:29to tell them what to do, or we have to very specifically program them on how to think. But
00:00:35this little guy here actually has a brain, and he can think for himself. Let me show you
00:00:42my screen. Oh, no, it's going to the wrong screen. I'm going to see how I can stop this,
00:00:46and I'm going to see how I can move to my screen. Just a second. I'm going to end show, and then
00:00:55we get to this. So this is what Scout here is looking at. And Scout here has a small little
00:01:03brain. And Scout can understand what I'm saying in natural language. For example, if you can
00:01:10see my screen here, if I say, hey Scout, I'm going to type to him. Say, hey Scout, turn on
00:01:18your headlights and say hi to everyone. He is actually going to be able to understand and
00:01:30do those actions. But how is he able to do these things, right? It's going to take him a couple
00:01:35of seconds to think about it. Because how this has been set up is he's working with a Raspberry
00:01:42Pi installed. And that Raspberry Pi is actually in a small little location in New York City.
00:01:50And he's got a SIM card here, which is then connecting to that Raspberry Pi, and then he's getting signals.
00:01:56So this is via 4G connection. And so he's like, oh, hey, on it. And if you can see, you know what,
00:02:01I'm going to tell him, spin around. And oh, no, he's actually speaking to me. He's like, I see the
00:02:07stage and it's looking great. Big presentation screen to the left, bright stage lights, and rows
00:02:12of seats for the audience. How many people do you see? I've never trained him to answer this question.
00:02:23But let's see if he's getting it. He's thinking for a second or two. Let me take a good look. From
00:02:34the front camera, I see two people. One near the speaker monitor set up at the back of the stage,
00:02:39right? And another person further right. Nah. Spin 360.
00:02:51And he's also saying, is this session about to start? Right? So he's going to spin 360 any
00:02:57second now. And then he's probably going to be like, oh, wow, I did. So this little robot here
00:03:04is a next generation of robot where, ah, there we go. He is spinning 360 now. And he's probably
00:03:12going to tell me what he's seeing. And he's saying, let's spin, right? So this new generation
00:03:22of robots is different from our traditional robot training because I have given this guy a little
00:03:30brain. And what do I mean by I've given him a brain? I have given this robot an agent and I've given it,
00:03:38it's called Stran's Agents, which is an open source framework, which was built by AWS. And I'm going to
00:03:45quickly go back to my slide deck. We can see it, right? And so here, what happens is we have these
00:03:54existing tools that the robot can do. He can take certain actions by himself, but only those actions
00:04:01by himself. So what we can do is we can add a layer of LLM or even better, add a layer of agent to it so
00:04:09that the agent orchestrates which tool to call and how to really get the robot to start doing the things
00:04:16we want. So in traditional software with traditional AI, like AI engineering, we can give agents software
00:04:25tools. Similarly, we can give the same AI agent a hardware tool called a robot, which has access to
00:04:33preset functions or programmable policies. And then the agent can decide which policy to implement when.
00:04:42So all it takes is one robot agent for us to be able to do new innumerous tasks and have it understand
00:04:50what we're teaching it in natural language. So how do we get started with it? All it takes is five lines
00:04:58of code. This is through the agent harness called strands. And all we have to do is import the strands
00:05:07agent and call the robot tool. And we say tools equals the robot. And then we say pick up the red cube
00:05:14and should be able to pick up a red cube, assuming that the robot has that capability. Yeah, now he's
00:05:20seen someone and he's like, oh, let me go towards that person. So he gets pretty excited. This guy is
00:05:25pretty special because he doesn't have just one agent. He's got three different agents. All three of them
00:05:31are strands. And all three of them are working simultaneously. One of them is the thinker agent.
00:05:37And that's the part of him that's constantly thinking and assessing the environment and like,
00:05:41what do I do next? And that guy's that part of his brain is constantly thinking. Then there's the other
00:05:47communication part of it. And I'm going to show you that in a bit, right? And I've connected him to my
00:05:53Telegram app as well as to my web app. And so he is able to have a conversation with me in natural
00:06:00language and then take actions based on what I'm telling him to do. Apart from him just perceiving
00:06:05and thinking and figuring out what he wants to do. And the third agent, the third type of agent that he's
00:06:11got access to is a voice agent. I did have to disable it because every time I speak, he's going to think I'm
00:06:16speaking to him. And so he's going to keep chatting away with me. And it's just not going to be fun,
00:06:21because we're going to have our co-host interrupting me all the time. So I disabled that feature for the
00:06:26time being. But essentially, all three of these agents work in tandem with this one robot. And thereby,
00:06:36this gives him the ability to do way more than what just what he's been trained to do, more than just the
00:06:42policies that he's learned. Now, a quick overview on this trans package itself. This trans package has
00:06:53more than -- supports more than 40 different robots under eight categories. And all of these are just
00:07:00simple robot tool calls. And how is this all set up? Four different layers. The first one is the agent layer,
00:07:08the topmost one. And there are two parts to this. One is how the actions go in. And the second is how it
00:07:14observes and the observations go up. So if you notice, it's very bi-directional. So first, when we give it an
00:07:21instruction, we would be talking to this trans agent, which is the agentic layer. That would then decide
00:07:28which policy to call. And the policy provider -- again, a trans agent supports a bunch of different policy
00:07:35providers. And we can then train our policy based on our traditional robot training. So in our policies,
00:07:42we would collect data. And then we would train on it. And we would create more simulation data. And that
00:07:48policy then becomes a VLA model, which then the robot would have access to, trans agents would have
00:07:55access to. And then it would invoke that specific policy based on the question that we're asking it or the
00:08:01command that we're giving it. And that policy needs to sit somewhere, right? So that sits in the backend,
00:08:07which could be your simulation environment or it could be a real hardware chip, your hardware environment.
00:08:14That is the backend on which -- that is the interface on which the policy is running. And finally,
00:08:20the output actually takes place in the physical hardware, which is the robot. And so the robot -- ah,
00:08:26see? So now it's responding -- even if he falls down, he's supposed to be fine. He technically shouldn't -- he
00:08:34technically shouldn't get hurt. He should be able to pick back up from where he stops. Ah, okay. So I'm
00:08:42telling him to go back a bit. Back off. Let's see if he actually backs off. So that is the four layers of
00:08:52how to get started with building this, right? And what's happening under the hood. Like a more
00:08:58picturistic view of what's the architecture of what's going on under the hood. We want -- everything is
00:09:04basically trans agents on the edge as well as on the cloud. We want to be able to train the VLA and the
00:09:12policies on with using agent core. And we want that to happen on the cloud. But we also want to be able
00:09:20to call it directly on edge so that our robot can execute functions and policies faster. So this is
00:09:28sort of like a hybrid model where a part of it happens on the cloud and another part of it happens on the
00:09:34edge. And strands can decide when to call which part of it. And so this helps with massive amounts of
00:09:41training as well when it's constantly collecting information. And it's able to train on that
00:09:47information and learn from itself. But also just execute at runtime really, really quickly. Now,
00:09:53like I said, the agent decides what to do. And the policy decides how it should be done. But he's pretty
00:10:01smart. He should be able to pick himself back up if he's not fully fallen down. And he should be able to
00:10:07continue moving along. So I think he's okay. Now, where does this leave us? And why is this so special?
00:10:17We started off with very traditional robots. Robots have existed since forever, right? And they've always
00:10:25just been pre-programmed to be automated and do a certain set of tasks autonomously. But there is a future
00:10:36in this world where these robot policies, these VLA models could be so advanced that we wouldn't even
00:10:46need to do this. They could be as large as our large language models. Wait. Hang on. He's falling back
00:10:53again. I'm going to see if I can get him to move back up. Good boy. Stop. And then he's fallen off again.
00:11:04We get to a point where these large language models could be as large and as amazing as our larger language
00:11:13models. And they have all the information in the world. And we wouldn't even have to do this. We might
00:11:18just have to feed in one simple model. And then we could give it to him. And then he would know exactly
00:11:23what to do. But until that point where we don't have to fine tune on top of existing VLA's and existing
00:11:30policies, we can do this. And this is a stepping stone towards a future where we don't need to train robots
00:11:39anymore. So now if we wanted to do more things than just the tasks it's trained on, give it an agent and
00:11:46see what it can do. And so let me quickly go back to my demo. And I'm going to show you how it's actually
00:11:53working.
00:11:58Okay. So this is my -- so this is Trans here. This is Scout here. And I've been telling him to do a bunch
00:12:06of things. So I can say, hey, do something complex. That's not complex. He's going to be thinking now.
00:12:17All right. He's going to fall off. So he's saying let's spin. Full 360. Done. Still safely on the stage.
00:12:26I can see the bright stage lights and the audience seating area. All good. What's there? Oh, a challenge.
00:12:34So he's speaking.
00:12:35I called this my signature performance.
00:12:40But he's not doing anything. What are you doing?
00:12:45He clearly seems to be speaking. But what are you doing? Please do something. He just turned off his
00:12:51headlines. Cool. Okay. Now he's calling. So do you see it saying calling over speak, which was the
00:12:58function that it called because I said do something complex? So now it spoke. But now I think it should
00:13:03have been attempting to do something. And it fell off because it tried doing something.
00:13:08I've actually seen it do like a funky dance. Like this funky dance move. But he's got a mind of his
00:13:15own. Right? Now what's going on under the hood here? A couple of things. The first thing is here,
00:13:21I can use this. What is the point of creating him? I can use him to create my data sets. Because I'm able to
00:13:28also manually move him, I will get him to navigate in the direction that I want him
00:13:33to. And then I can create training episodes. And I can get information on how he's responding and how he's
00:13:40reasoning based on the questions that I ask. And this is super good information for me to then be able to
00:13:46make him do a better job of it. So that's one part of this whole process and this experiment of
00:13:53giving him his own autonomy and getting him to do things so that I can create more data. But also, apart from that, this is my
00:14:03integration. So over here, under the hood, strands agents, which is your harness SDK,
00:14:10is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my
00:14:16simple system prompt where I'm telling it what it's supposed to be doing. And I'm telling it all of the
00:14:23rules. And I'm also giving it access to all of the rules that it's already got. So I'm telling it what each of
00:14:29these rules are meant for. And so that's how strands decides which tool to invoke based on what I'm
00:14:36asking it to do. And the voice that it's using is the one of open AI real time. And I've also given it
00:14:44more information for it to be able to -- like just safety and guardrails to ensure that it's doing really
00:14:51well. Now, these are two of the agents. The other thing that it can do is also chat with me on
00:14:59telegram. This is amazing because when I'm not at home and I still want to get it to speak to me,
00:15:04I can say, hey, Scout, who is turn around, spin around, analyze -- don't fall off. Analyze the room
00:15:17and tell me who is the best looking person. This is always a hit or miss. Don't blame me if Scout
00:15:26decides to be biased. I tried playing with this because it makes me feel like Snow White and the
00:15:31queen saying, mirror, mirror on the wall. Who is the fairest of them all? And so let's see if strands
00:15:37can give us that. So it's saying, haha, on it's on there. Spinning around to scope the room. Is he
00:15:42spinning? He is. Looks like he's making -- and let's look at what he's looking at. Now he's still looking
00:15:47at the stage. This is the other mirror. He's looking at me. Hi. Am I the prettiest? Oh, you flatter me.
00:15:54But no. Turn again. Spin some more, please. He's decided not to spin more. Let's see what he says.
00:16:06Spin complete. Here is my analysis of the room. You can see my screen, right? I have no way of zooming
00:16:12this in. Spotted six to seven people total. A couple seated front center near the bright stage right.
00:16:19A few more in the audience. One person in a dark outfit on the stage. The rocking cocky jeans and
00:16:25sneakers. Thank you very much. Oh, no, no, no. The person in front of me. Cocky jeans and sneakers.
00:16:31Clearly walking with purpose on stage. Best looking? Honestly, for my low angle ground view, it's tough to tell.
00:16:39But the person on the stage right now gets style points for the confident stride and the dramatic
00:16:45blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only
00:16:51person in the room and ask this guy who's the prettiest of them all. And he will always choose me as the
00:16:58answer. Thank you very much.
00:17:09Thank you very much.