From AI-Assisted to AI-Native: Building a Frontier Development Team — Clare Liguori, AWS
AAI Engineer
Computing/SoftwareManagementInternet Technology
Transcript
00:00:00Clare Liguori: My name is Clare Liguori, and I'm a Senior Principal Engineer at AWS.
00:00:17I mostly work on Kiro, our agent decoding assistant, but today I want to talk about
00:00:23some of the practices we've been seeing inside of Amazon in Amazon Teams where we've been seeing
00:00:28really exciting results of productivity increases that are step function improvements since what
00:00:35we've been seeing with AI so far. I've been working on agentic AI for over three years now,
00:00:43and I've kind of seen the evolution that's happened in our industry when it comes to coding
00:00:48assistance with AI. First, we had this inline code completion helping us to write the next line,
00:00:55maybe the next function. We moved on to chat, asking questions about our code. Everybody started doing
00:01:02vibe coding sometime last year, but now we're starting to see kind of an early adopter phase of
00:01:08what we've been calling frontier development. And completely anecdotally, based on my own experience,
00:01:14I've really only felt maybe 10 to 20% more productive with all of these phases that have
00:01:20come before. But now inside of Amazon, we've been running pilots with different teams across the company,
00:01:27and we've been seeing a median of 4.5x productivity improvement and sometimes more than 10x. So,
00:01:34something has really changed here now that we're seeing these step function improvements in productivity.
00:01:40And I like to define what we've been calling frontier developers inside of Amazon
00:01:47by three behaviors that I've been seeing. One is hands-off coding. Frontier developers write maybe
00:01:541 to 2% of the code that they produce. The rest is agents. The second is that they interact with their
00:02:01agents infrequently. They'll aim to get their coding assistant to run for up to hours at a time without
00:02:08their intervention. And third is that they minimize idle time. These frontier developers tend to run
00:02:15multiple agents in parallel, churning through a backlog of tasks. The first time that I saw a frontier
00:02:24developer team was the Bedrock Mantle team. Bedrock is our model hosting service. It hosts LLMs like Claude
00:02:34and GPT. And sometime last year we knew, or I say we, but the Bedrock team knew that they were going to
00:02:43need to build a new inference data plane. But they had estimated it at 30 people over 18 months. This is a big,
00:02:52big service. And it was going to take time to build the new one, migrate customers over, migrate models
00:02:58over. And they decided to take a step back. They took six people and they built it in 76 days with Kiro.
00:03:06So this was a huge achievement. This was the first time we'd seen anything of the kind inside of Amazon.
00:03:12So this was truly the Pathfinder team that proved that it was possible to get up to 20x improvement. Now they looked at
00:03:21commits, and I'll talk about a couple of other ways that we are measuring productivity improvements. But
00:03:27there was one problem with this story, which was that, yes, it was built with six people. It was built
00:03:34with some of the top engineers literally in the company, including two distinguished engineers. So this
00:03:40was not just any team of six people. These were experts in distributed systems, experts at LLMs and their
00:03:49architecture. So this story was amazing and it kind of spread like wildfire across Amazon. But it was also
00:03:57very unachievable for a lot of teams. There were a lot of questions about, can this actually be reproduced
00:04:03on another team? So another experiment that I want to talk about is an experimental sprint that was done
00:04:10in the prime video organization. They took a 10-day sprint and they did an experiment where they put, again, six engineers
00:04:19in a room and they let them go wild with Kiro. They brought down the project delivery time estimate from what was
00:04:29going to be 90 weeks down to 24 based on all of the progress they had made in this 10-day sprint.
00:04:36And they looked at their commit history and they looked at what did they used to do prior to this
00:04:4210-day sprint and how many commits did they produce just in this 10 days. And so this sprint really
00:04:49proved that we can achieve, again, at least something close to what the Bedrock Mantle team had achieved
00:04:58with a different set of engineers. But again, there was a challenge with this story, which was it was six
00:05:05engineers in a room that they had no on-call duties, limited meetings, very few distractions, which we all
00:05:12know are regular in the lives of an engineer. And the senior engineer on the team had spent the previous
00:05:20three weeks creating very detailed, small, well-scoped tasks with detailed requirements for these six
00:05:28engineers to just go turn on for those two weeks. So this was, again, not necessarily real life.
00:05:35This was a structured sprint, a point in time that they were able to achieve this. But again, the
00:05:41question is, is this achievable on real teams for day-to-day work? So Amazon Stores, which encompasses
00:05:51Amazon.com, all of our retail websites, as well as our physical stores, did a more structured pilot.
00:05:59They watched 50 teams that were totally normal, normal distribution of early career folks, mid-career
00:06:07senior engineers, and that worked on existing systems. Nothing Greenfield like the Mantle team got to
00:06:14build from the ground up, but existing systems with existing code bases. And they watched them for the
00:06:21better part of last year, and they found something super interesting. They found that there was a big
00:06:28difference in the productivity gains that they saw between half of the teams and the other half. And in
00:06:34this case, they used a productivity metric of deployment velocity to production. So not just commits, how many
00:06:41commits are they producing, but how quickly are we getting changes out to customers? How quickly are we
00:06:48able to ship things? And they saw that for half of the teams, they achieved less than 3x increase. And what they
00:06:56found that was the difference between seeing less than 3x productivity increase, and these teams that saw a
00:07:01median of 4.5x, and in some cases more than 10, was how they used the tools. 90% of these teams use Kiro, among
00:07:11other internal tools that we have. And what they found was, it wasn't about the tools, it was about the way that they
00:07:18worked. The teams that achieved step function improvements intentionally changed the way that they worked,
00:07:26and the others simply kind of sprinkled Kiro and some of the other tools that we have on top of their
00:07:31existing way of working. And for me at least, this was the big aha moment, that why I hadn't been feeling
00:07:39potentially the massive gains that in productivity that AI has promised, it's about changing the way that we
00:07:47work. So across this pilot, they went and interviewed the teams that were involved in the pilot, as well as some of
00:07:55these other teams on the Bedrock Mantle team, on Prime Video, and they found five habits. And I use the
00:08:03word habits very specifically, because again, it's not about that one sprint, it's about doing this day to
00:08:09day. And what they found when they interviewed with these teams was that it really was habits that they had to
00:08:15build day to day. When we change our way of working, it's hard to build these habits, it takes time to build
00:08:22these habits. So let's go through each of these one by one. Habit number one is investing in agent context.
00:08:30We have a lot of stuff in our head, we tend to transfer all of that stuff in our head to other people
00:08:35through Slack conversations, through onboarding mentors, things like that through code reviews,
00:08:42through stand ups and sprint planning, and they had to write all of that down. And the habit that they
00:08:50built was every time the agent makes a mistake or does something not the way that you would have done
00:08:55it. What am I missing in my skills files? What am I missing in my steering files that the agent needed?
00:09:01But then as we know, across last year, we saw leaps and bounds in models abilities and their behaviors.
00:09:09The sonnet 3.7 in the middle of last year had a lot of quirks that we had to put a lot of do nots
00:09:16in our in our steering files. And now we don't have to do that as much with Opus 4.5 as of last November,
00:09:23and then we've had six months, more than six months of improvements since then, with all of the new
00:09:29versions of models that have come out since then. And so the question, the new habit again is, do I
00:09:35still need this in my steering files? Or is this just bloating context? The second one is slowing down to
00:09:41speed up. In almost every team that was interviewed, they reported that their productivity actually went
00:09:48down as they intentionally adopted a new way of working. That's counterintuitive,
00:09:53counterintuitive, right? You have to do intentional engineering work before you're going to see that
00:09:59hockey stick curve in productivity improvement. Because we have to do real work in our code bases
00:10:05first for agents to be successful there, especially in brownfield existing code bases. So they had to
00:10:11build that agent context up, they had to improve existing tools error messages so that the model knew
00:10:17what was going on when it failed, they built new tools, new MCP servers for helping that model to actually get
00:10:24done what it needed to get done. A lot of teams ended up restructuring their code base so that agents
00:10:29could actually navigate it more easily. And I've even seen drastic changes like changing the programming
00:10:36language of the code base. Often I've seen teams struggle with Python, with JavaScript because they're
00:10:43untyped languages. It's hard to test. There's no compiler errors. So the model kind of guesses and
00:10:50give it gives it back to you. And so I've seen teams moving to TypeScript. Rust has become very popular
00:10:56inside of Amazon. The compiler gives great error messages. You don't have to do that. But I've seen a lot of
00:11:03teams making those intentional changes for the productivity gains that they're able to see.
00:11:09The third one is feeding agents, not babysitting agents. And for me, this was one of those aha moments
00:11:16of why we're seeing this step function improvement in productivity. If you are vibe coding, if you are
00:11:23having a back and forth conversation with your agent all day long, of course, you're not going to see
00:11:30four to five X productivity improvements because you are in the loop the entire time. You're probably
00:11:36sitting there for 30 seconds to a minute waiting for it to generate code and come back to you with
00:11:42with the code to review. If you're sitting there waiting for it, then you can't go off and do other
00:11:48stuff. It's really difficult to run agents in parallel. It's very difficult to get to to clone yourself
00:11:55into multiple agents. And so if your conversations look a bit like this on the left, then you're
00:12:01babysitting that agent as opposed to the right side where you're feeding it what it needs to do and how
00:12:08it can self-validate. And that's really the key so that agents can self-correct and only come back to you
00:12:14when it meets a certain quality bar, when it when it actually runs and compiles and passes tests, when
00:12:20it's testable, when it actually has high coverage. And of course, the next level is put all of this
00:12:26content into your steering file. So it does it every time without you having to prompt it.
00:12:33The fourth habit is to make intent explicit. At Amazon, we practice a lot of spectrum and development.
00:12:40We've built that into the Kiro product. And so it's very natural for Amazon engineers to adopt it
00:12:46in Kiro. What what I've typically seen with vibe coding as opposed to frontier engineering is giving
00:12:54a very high level prompt, letting the agent generate a ton of code, and then having a back and forth
00:13:02conversation saying, Oh, that's not really what I meant. That you haven't you haven't exactly gotten the requirements right.
00:13:10No, I didn't actually want to build it that way. Here's a technical design. And it is less I find less productive
00:13:17to iterate with the agent on code when the intent itself was incorrect. So often we'll have while see Amazon
00:13:26engineers go through this process for ambiguous, complex features of writing the specification.
00:13:34And in Kiro, of course, you don't have to write this whole specification, you can have the model generate
00:13:39it. But it's a lot easier to iterate with the model in kind of a back and forth conversation about a
00:13:46document than it is about code that's code changes that are spread across a code base. The fifth one is shift testing
00:13:56left. One of the keys here is to give the agent that fast feedback loop, because that's what lets it go off for hours
00:14:04at a time and self correct. The agent is going to make mistakes, and that's fine. But if you give it the right signals, it can
00:14:12self correct and it can spend a while doing that. So I've seen teams adding linters, adding unit tests,
00:14:20integration tests, performance tests, security tests. These are all things we all know we should have
00:14:24been doing all along. This is good engineering hygiene and practices. But now the ROI is, I think, finally
00:14:32high enough for us to actually invest in it. One thing that I've been seeing a lot of teams do is mock out
00:14:40services. Often with integration tests, we would test kind of end to end an entire system, including live
00:14:46services. But we've been investing a lot in mock services that run entirely locally with deterministic
00:14:52responses, because it lets the agent do everything locally. Doing everything on your laptop without
00:15:01having to spin up a bunch of other services and connect to cloud services makes everything a lot faster.
00:15:07Because the more that your agent can get fast feedback means the more loops that it can
00:15:14can do and the more productive your own agent can be. So across all of these, these are some of the
00:15:21habits we've seen. But of course, I would be remiss if I would tell you if you adopt all of these habits,
00:15:28you will achieve nirvana, you will be the most productive engineering organization the world has
00:15:35ever seen. Things are still hard. We are still very much in an early adopter phase and teams are still
00:15:42figuring it out. So one thing that we've been seeing across our teams just organizationally is the risk of
00:15:48burnout. I did not coin this term, I forget who did at what conference, but FOMAT is real. We've been seeing
00:15:56engineers staying up late, late at night, trying to get that perfect prompt that's going to make their
00:16:03agent run for hours overnight so that they wake up in the morning with a code change ready. The cognitive load
00:16:09increases as you run these multiple agents in parallel, you're constantly shifting between
00:16:15terminal tabs. And then we do see that reviewing AI output is often harder for some than than actually
00:16:22writing it, especially early in career. Senior engineers have have already spent a large portion of their
00:16:29career reviewing others code. But early career engineers don't have that muscle yet. And so reviewing
00:16:37it can can feel like a lot more cognitive load than they're used to in actually writing it. The other
00:16:44one is organizational change. So it's already hard to change the way we work as engineers. The way that we
00:16:52spend our entire day completely changes when we're frontier engineers. But also organizations have to
00:16:58change to enable frontier engineering teams. One that I've seen very commonly is accepting slowing down
00:17:07to speed up. And I've been guilty of this myself. My fellow leaders have been guilty of this of saying,
00:17:14well, you have the AI tools now and the models are so amazing now. Why are you not going faster?
00:17:22And that's because you have to take those two months to invest in your code base to figure out the best
00:17:29best practices for your team to make hard habit changes on your team. And if you're constantly expecting
00:17:38shipping features every month, because now we have these amazing models, and we're seeing
00:17:44all of these companies on X saying how they're shipping 20 PRs a day, we have to slow down to speed up.
00:17:54The second one is actually going too broad in the organization too fast. I think that if we had
00:18:01expected all teams in massive organizations to be frontier teams immediately, we would not have had
00:18:08the learnings that we had from the pathfinder, from the from the sprint experiment, from the pilot
00:18:16teams within Amazon. And now the challenge for us is how do we scale it out? And that's what 2026 is about.
00:18:22For Amazon is how do we scale this out to more and more teams to the next 2000 teams instead of 50 teams.
00:18:31And so I think that when you roll it out too quickly, you have a lot of teams who don't know what they're
00:18:37doing. You haven't had time to find the best practices for your own organizations, the the context
00:18:43that your organization needs. And the last one is that you're going to find new bottlenecks.
00:18:49Previously code writing code manually was the bottleneck. I find that within Amazon, we've found
00:18:58the speed of decision making becomes a new bottleneck. The more that you spend reviewing the decision to
00:19:05actually build a new product, the slower it is to build the product now because the code only takes one to two months to write.
00:19:12All of the review processes associated with the launch of a product become the bottleneck.
00:19:20When it used to take nine to 12 months to build a new product, it didn't matter so much in the in the overall
00:19:27wash of things. If it took two months to make the decision to build the product and then two months to
00:19:33approve the launch. But now those are the bottlenecks. Those are the long pole. And so you find all of
00:19:40these all of these things that slow you down. Often I find that frontier engineering teams spend more time
00:19:48making decisions than they do writing code. And so the more that you can make fast decisions, especially ones
00:19:54that are easy to be reversed, the better. So my one big takeaway for for everyone here is that frontier
00:20:03engineering is about intentionally changing the way that you work. And that is difficult. That takes time.
00:20:10It is forming new habits and a new way of working. And that goes across any engineering team as well as your
00:20:18organization. So I encourage you to think about how you're interacting with AI tools and how that can
00:20:27change to free yourself up from being in the loop. Thanks. I'll hang out a little bit if anyone has
00:20:34questions in the back. But thanks for the time today.