Tell the Robot What You Want — Sandhya Subramani, AWS

AAI Engineer
Computing/SoftwareConsumer Electronics

Transcript

00:00:00Hello, everyone. My name is Sandy. And meet my co-host today, Scout. This is my friendly
00:00:21rover. And one would think that rovers can't really think for themselves, right? We have
00:00:29to tell them what to do, or we have to very specifically program them on how to think. But
00:00:35this little guy here actually has a brain, and he can think for himself. Let me show you
00:00:42my screen. Oh, no, it's going to the wrong screen. I'm going to see how I can stop this,
00:00:46and I'm going to see how I can move to my screen. Just a second. I'm going to end show, and then
00:00:55we get to this. So this is what Scout here is looking at. And Scout here has a small little
00:01:03brain. And Scout can understand what I'm saying in natural language. For example, if you can
00:01:10see my screen here, if I say, hey Scout, I'm going to type to him. Say, hey Scout, turn on
00:01:18your headlights and say hi to everyone. He is actually going to be able to understand and
00:01:30do those actions. But how is he able to do these things, right? It's going to take him a couple
00:01:35of seconds to think about it. Because how this has been set up is he's working with a Raspberry
00:01:42Pi installed. And that Raspberry Pi is actually in a small little location in New York City.
00:01:50And he's got a SIM card here, which is then connecting to that Raspberry Pi, and then he's getting signals.
00:01:56So this is via 4G connection. And so he's like, oh, hey, on it. And if you can see, you know what,
00:02:01I'm going to tell him, spin around. And oh, no, he's actually speaking to me. He's like, I see the
00:02:07stage and it's looking great. Big presentation screen to the left, bright stage lights, and rows
00:02:12of seats for the audience. How many people do you see? I've never trained him to answer this question.
00:02:23But let's see if he's getting it. He's thinking for a second or two. Let me take a good look. From
00:02:34the front camera, I see two people. One near the speaker monitor set up at the back of the stage,
00:02:39right? And another person further right. Nah. Spin 360.
00:02:51And he's also saying, is this session about to start? Right? So he's going to spin 360 any
00:02:57second now. And then he's probably going to be like, oh, wow, I did. So this little robot here
00:03:04is a next generation of robot where, ah, there we go. He is spinning 360 now. And he's probably
00:03:12going to tell me what he's seeing. And he's saying, let's spin, right? So this new generation
00:03:22of robots is different from our traditional robot training because I have given this guy a little
00:03:30brain. And what do I mean by I've given him a brain? I have given this robot an agent and I've given it,
00:03:38it's called Stran's Agents, which is an open source framework, which was built by AWS. And I'm going to
00:03:45quickly go back to my slide deck. We can see it, right? And so here, what happens is we have these
00:03:54existing tools that the robot can do. He can take certain actions by himself, but only those actions
00:04:01by himself. So what we can do is we can add a layer of LLM or even better, add a layer of agent to it so
00:04:09that the agent orchestrates which tool to call and how to really get the robot to start doing the things
00:04:16we want. So in traditional software with traditional AI, like AI engineering, we can give agents software
00:04:25tools. Similarly, we can give the same AI agent a hardware tool called a robot, which has access to
00:04:33preset functions or programmable policies. And then the agent can decide which policy to implement when.
00:04:42So all it takes is one robot agent for us to be able to do new innumerous tasks and have it understand
00:04:50what we're teaching it in natural language. So how do we get started with it? All it takes is five lines
00:04:58of code. This is through the agent harness called strands. And all we have to do is import the strands
00:05:07agent and call the robot tool. And we say tools equals the robot. And then we say pick up the red cube
00:05:14and should be able to pick up a red cube, assuming that the robot has that capability. Yeah, now he's
00:05:20seen someone and he's like, oh, let me go towards that person. So he gets pretty excited. This guy is
00:05:25pretty special because he doesn't have just one agent. He's got three different agents. All three of them
00:05:31are strands. And all three of them are working simultaneously. One of them is the thinker agent.
00:05:37And that's the part of him that's constantly thinking and assessing the environment and like,
00:05:41what do I do next? And that guy's that part of his brain is constantly thinking. Then there's the other
00:05:47communication part of it. And I'm going to show you that in a bit, right? And I've connected him to my
00:05:53Telegram app as well as to my web app. And so he is able to have a conversation with me in natural
00:06:00language and then take actions based on what I'm telling him to do. Apart from him just perceiving
00:06:05and thinking and figuring out what he wants to do. And the third agent, the third type of agent that he's
00:06:11got access to is a voice agent. I did have to disable it because every time I speak, he's going to think I'm
00:06:16speaking to him. And so he's going to keep chatting away with me. And it's just not going to be fun,
00:06:21because we're going to have our co-host interrupting me all the time. So I disabled that feature for the
00:06:26time being. But essentially, all three of these agents work in tandem with this one robot. And thereby,
00:06:36this gives him the ability to do way more than what just what he's been trained to do, more than just the
00:06:42policies that he's learned. Now, a quick overview on this trans package itself. This trans package has
00:06:53more than -- supports more than 40 different robots under eight categories. And all of these are just
00:07:00simple robot tool calls. And how is this all set up? Four different layers. The first one is the agent layer,
00:07:08the topmost one. And there are two parts to this. One is how the actions go in. And the second is how it
00:07:14observes and the observations go up. So if you notice, it's very bi-directional. So first, when we give it an
00:07:21instruction, we would be talking to this trans agent, which is the agentic layer. That would then decide
00:07:28which policy to call. And the policy provider -- again, a trans agent supports a bunch of different policy
00:07:35providers. And we can then train our policy based on our traditional robot training. So in our policies,
00:07:42we would collect data. And then we would train on it. And we would create more simulation data. And that
00:07:48policy then becomes a VLA model, which then the robot would have access to, trans agents would have
00:07:55access to. And then it would invoke that specific policy based on the question that we're asking it or the
00:08:01command that we're giving it. And that policy needs to sit somewhere, right? So that sits in the backend,
00:08:07which could be your simulation environment or it could be a real hardware chip, your hardware environment.
00:08:14That is the backend on which -- that is the interface on which the policy is running. And finally,
00:08:20the output actually takes place in the physical hardware, which is the robot. And so the robot -- ah,
00:08:26see? So now it's responding -- even if he falls down, he's supposed to be fine. He technically shouldn't -- he
00:08:34technically shouldn't get hurt. He should be able to pick back up from where he stops. Ah, okay. So I'm
00:08:42telling him to go back a bit. Back off. Let's see if he actually backs off. So that is the four layers of
00:08:52how to get started with building this, right? And what's happening under the hood. Like a more
00:08:58picturistic view of what's the architecture of what's going on under the hood. We want -- everything is
00:09:04basically trans agents on the edge as well as on the cloud. We want to be able to train the VLA and the
00:09:12policies on with using agent core. And we want that to happen on the cloud. But we also want to be able
00:09:20to call it directly on edge so that our robot can execute functions and policies faster. So this is
00:09:28sort of like a hybrid model where a part of it happens on the cloud and another part of it happens on the
00:09:34edge. And strands can decide when to call which part of it. And so this helps with massive amounts of
00:09:41training as well when it's constantly collecting information. And it's able to train on that
00:09:47information and learn from itself. But also just execute at runtime really, really quickly. Now,
00:09:53like I said, the agent decides what to do. And the policy decides how it should be done. But he's pretty
00:10:01smart. He should be able to pick himself back up if he's not fully fallen down. And he should be able to
00:10:07continue moving along. So I think he's okay. Now, where does this leave us? And why is this so special?
00:10:17We started off with very traditional robots. Robots have existed since forever, right? And they've always
00:10:25just been pre-programmed to be automated and do a certain set of tasks autonomously. But there is a future
00:10:36in this world where these robot policies, these VLA models could be so advanced that we wouldn't even
00:10:46need to do this. They could be as large as our large language models. Wait. Hang on. He's falling back
00:10:53again. I'm going to see if I can get him to move back up. Good boy. Stop. And then he's fallen off again.
00:11:04We get to a point where these large language models could be as large and as amazing as our larger language
00:11:13models. And they have all the information in the world. And we wouldn't even have to do this. We might
00:11:18just have to feed in one simple model. And then we could give it to him. And then he would know exactly
00:11:23what to do. But until that point where we don't have to fine tune on top of existing VLA's and existing
00:11:30policies, we can do this. And this is a stepping stone towards a future where we don't need to train robots
00:11:39anymore. So now if we wanted to do more things than just the tasks it's trained on, give it an agent and
00:11:46see what it can do. And so let me quickly go back to my demo. And I'm going to show you how it's actually
00:11:53working.
00:11:58Okay. So this is my -- so this is Trans here. This is Scout here. And I've been telling him to do a bunch
00:12:06of things. So I can say, hey, do something complex. That's not complex. He's going to be thinking now.
00:12:17All right. He's going to fall off. So he's saying let's spin. Full 360. Done. Still safely on the stage.
00:12:26I can see the bright stage lights and the audience seating area. All good. What's there? Oh, a challenge.
00:12:34So he's speaking.
00:12:35I called this my signature performance.
00:12:40But he's not doing anything. What are you doing?
00:12:45He clearly seems to be speaking. But what are you doing? Please do something. He just turned off his
00:12:51headlines. Cool. Okay. Now he's calling. So do you see it saying calling over speak, which was the
00:12:58function that it called because I said do something complex? So now it spoke. But now I think it should
00:13:03have been attempting to do something. And it fell off because it tried doing something.
00:13:08I've actually seen it do like a funky dance. Like this funky dance move. But he's got a mind of his
00:13:15own. Right? Now what's going on under the hood here? A couple of things. The first thing is here,
00:13:21I can use this. What is the point of creating him? I can use him to create my data sets. Because I'm able to
00:13:28also manually move him, I will get him to navigate in the direction that I want him
00:13:33to. And then I can create training episodes. And I can get information on how he's responding and how he's
00:13:40reasoning based on the questions that I ask. And this is super good information for me to then be able to
00:13:46make him do a better job of it. So that's one part of this whole process and this experiment of
00:13:53giving him his own autonomy and getting him to do things so that I can create more data. But also, apart from that, this is my
00:14:03integration. So over here, under the hood, strands agents, which is your harness SDK,
00:14:10is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my
00:14:16simple system prompt where I'm telling it what it's supposed to be doing. And I'm telling it all of the
00:14:23rules. And I'm also giving it access to all of the rules that it's already got. So I'm telling it what each of
00:14:29these rules are meant for. And so that's how strands decides which tool to invoke based on what I'm
00:14:36asking it to do. And the voice that it's using is the one of open AI real time. And I've also given it
00:14:44more information for it to be able to -- like just safety and guardrails to ensure that it's doing really
00:14:51well. Now, these are two of the agents. The other thing that it can do is also chat with me on
00:14:59telegram. This is amazing because when I'm not at home and I still want to get it to speak to me,
00:15:04I can say, hey, Scout, who is turn around, spin around, analyze -- don't fall off. Analyze the room
00:15:17and tell me who is the best looking person. This is always a hit or miss. Don't blame me if Scout
00:15:26decides to be biased. I tried playing with this because it makes me feel like Snow White and the
00:15:31queen saying, mirror, mirror on the wall. Who is the fairest of them all? And so let's see if strands
00:15:37can give us that. So it's saying, haha, on it's on there. Spinning around to scope the room. Is he
00:15:42spinning? He is. Looks like he's making -- and let's look at what he's looking at. Now he's still looking
00:15:47at the stage. This is the other mirror. He's looking at me. Hi. Am I the prettiest? Oh, you flatter me.
00:15:54But no. Turn again. Spin some more, please. He's decided not to spin more. Let's see what he says.
00:16:06Spin complete. Here is my analysis of the room. You can see my screen, right? I have no way of zooming
00:16:12this in. Spotted six to seven people total. A couple seated front center near the bright stage right.
00:16:19A few more in the audience. One person in a dark outfit on the stage. The rocking cocky jeans and
00:16:25sneakers. Thank you very much. Oh, no, no, no. The person in front of me. Cocky jeans and sneakers.
00:16:31Clearly walking with purpose on stage. Best looking? Honestly, for my low angle ground view, it's tough to tell.
00:16:39But the person on the stage right now gets style points for the confident stride and the dramatic
00:16:45blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only
00:16:51person in the room and ask this guy who's the prettiest of them all. And he will always choose me as the
00:16:58answer. Thank you very much.
00:17:09Thank you very much.

Key Takeaway

Integrating open-source Strands agents with large language models enables robots to execute complex, natural language commands and coordinate multiple hardware tools beyond traditional pre-programmed policies.

Highlights

  • A Raspberry Pi connected via a 4G SIM card powers the rover named Scout from a remote location in New York City.

  • The Strands Agents open-source framework built by AWS provides the multi-agent architecture and orchestrates robot tool calls.

  • Three distinct Strands agents run simultaneously on the robot, handling thinking, communication, and voice interactions.

  • Anthropic Claude Opus 4.8 serves as the underlying language model acting as the brain for the Strands agent harness.

  • The Strands package supports more than 40 different robots across eight distinct hardware categories.

Timeline

Natural Language Robot Control and Hardware Setup

  • A rover named Scout processes natural language commands and responds with audio and physical actions.
  • A Raspberry Pi equipped with a 4G SIM card connects the robot to a remote server in New York City.
  • Traditional robot training relies on rigid pre-programmed policies rather than dynamic agent orchestration.

The demonstration shows the rover executing tasks like turning on headlights, spinning 360 degrees, and describing its visual surroundings using a front-facing camera. The hardware setup uses cellular connectivity to bridge the physical robot with backend processing systems, marking a shift away from standard automated robotics.

Multi-Agent Architecture and Strands Framework

  • The Strands open-source framework from AWS allows software agents to interface directly with hardware tools.
  • Three separate Strands agents operate concurrently to handle environment assessment, natural language conversation, and voice output.
  • A four-layer architecture links the agentic layer, policy provider, backend simulation or hardware interface, and physical robot hardware.

The system utilizes an agent harness where an LLM orchestrates preset robot functions and programmable policies. By running a thinker agent, communication agent, and voice agent simultaneously, the robot processes observations and executes runtime functions dynamically without requiring extensive retraining for every new task.

Hybrid Cloud-Edge Integration and Future Robotics

  • Vision-Language-Action models and policies are trained on the cloud using agent core while execution happens on the edge.
  • Anthropic Claude Opus 4.8 operates as the core intelligence under the hood of the Strands agent harness.
  • Integration with messaging applications like Telegram allows remote natural language interaction and data set collection.

The hybrid model balances heavy cloud-based training with rapid edge execution. Manual navigation and agent responses generate continuous training episodes and datasets, paving the way for future robotic systems powered directly by large language models.

Community Posts

View all posts