From AI-Assisted to AI-Native: Building a Frontier Development Team — Clare Liguori, AWS

AAI Engineer
Computing/SoftwareManagementInternet Technology

Transcript

00:00:00Clare Liguori: My name is Clare Liguori, and I'm a Senior Principal Engineer at AWS.
00:00:17I mostly work on Kiro, our agent decoding assistant, but today I want to talk about
00:00:23some of the practices we've been seeing inside of Amazon in Amazon Teams where we've been seeing
00:00:28really exciting results of productivity increases that are step function improvements since what
00:00:35we've been seeing with AI so far. I've been working on agentic AI for over three years now,
00:00:43and I've kind of seen the evolution that's happened in our industry when it comes to coding
00:00:48assistance with AI. First, we had this inline code completion helping us to write the next line,
00:00:55maybe the next function. We moved on to chat, asking questions about our code. Everybody started doing
00:01:02vibe coding sometime last year, but now we're starting to see kind of an early adopter phase of
00:01:08what we've been calling frontier development. And completely anecdotally, based on my own experience,
00:01:14I've really only felt maybe 10 to 20% more productive with all of these phases that have
00:01:20come before. But now inside of Amazon, we've been running pilots with different teams across the company,
00:01:27and we've been seeing a median of 4.5x productivity improvement and sometimes more than 10x. So,
00:01:34something has really changed here now that we're seeing these step function improvements in productivity.
00:01:40And I like to define what we've been calling frontier developers inside of Amazon
00:01:47by three behaviors that I've been seeing. One is hands-off coding. Frontier developers write maybe
00:01:541 to 2% of the code that they produce. The rest is agents. The second is that they interact with their
00:02:01agents infrequently. They'll aim to get their coding assistant to run for up to hours at a time without
00:02:08their intervention. And third is that they minimize idle time. These frontier developers tend to run
00:02:15multiple agents in parallel, churning through a backlog of tasks. The first time that I saw a frontier
00:02:24developer team was the Bedrock Mantle team. Bedrock is our model hosting service. It hosts LLMs like Claude
00:02:34and GPT. And sometime last year we knew, or I say we, but the Bedrock team knew that they were going to
00:02:43need to build a new inference data plane. But they had estimated it at 30 people over 18 months. This is a big,
00:02:52big service. And it was going to take time to build the new one, migrate customers over, migrate models
00:02:58over. And they decided to take a step back. They took six people and they built it in 76 days with Kiro.
00:03:06So this was a huge achievement. This was the first time we'd seen anything of the kind inside of Amazon.
00:03:12So this was truly the Pathfinder team that proved that it was possible to get up to 20x improvement. Now they looked at
00:03:21commits, and I'll talk about a couple of other ways that we are measuring productivity improvements. But
00:03:27there was one problem with this story, which was that, yes, it was built with six people. It was built
00:03:34with some of the top engineers literally in the company, including two distinguished engineers. So this
00:03:40was not just any team of six people. These were experts in distributed systems, experts at LLMs and their
00:03:49architecture. So this story was amazing and it kind of spread like wildfire across Amazon. But it was also
00:03:57very unachievable for a lot of teams. There were a lot of questions about, can this actually be reproduced
00:04:03on another team? So another experiment that I want to talk about is an experimental sprint that was done
00:04:10in the prime video organization. They took a 10-day sprint and they did an experiment where they put, again, six engineers
00:04:19in a room and they let them go wild with Kiro. They brought down the project delivery time estimate from what was
00:04:29going to be 90 weeks down to 24 based on all of the progress they had made in this 10-day sprint.
00:04:36And they looked at their commit history and they looked at what did they used to do prior to this
00:04:4210-day sprint and how many commits did they produce just in this 10 days. And so this sprint really
00:04:49proved that we can achieve, again, at least something close to what the Bedrock Mantle team had achieved
00:04:58with a different set of engineers. But again, there was a challenge with this story, which was it was six
00:05:05engineers in a room that they had no on-call duties, limited meetings, very few distractions, which we all
00:05:12know are regular in the lives of an engineer. And the senior engineer on the team had spent the previous
00:05:20three weeks creating very detailed, small, well-scoped tasks with detailed requirements for these six
00:05:28engineers to just go turn on for those two weeks. So this was, again, not necessarily real life.
00:05:35This was a structured sprint, a point in time that they were able to achieve this. But again, the
00:05:41question is, is this achievable on real teams for day-to-day work? So Amazon Stores, which encompasses
00:05:51Amazon.com, all of our retail websites, as well as our physical stores, did a more structured pilot.
00:05:59They watched 50 teams that were totally normal, normal distribution of early career folks, mid-career
00:06:07senior engineers, and that worked on existing systems. Nothing Greenfield like the Mantle team got to
00:06:14build from the ground up, but existing systems with existing code bases. And they watched them for the
00:06:21better part of last year, and they found something super interesting. They found that there was a big
00:06:28difference in the productivity gains that they saw between half of the teams and the other half. And in
00:06:34this case, they used a productivity metric of deployment velocity to production. So not just commits, how many
00:06:41commits are they producing, but how quickly are we getting changes out to customers? How quickly are we
00:06:48able to ship things? And they saw that for half of the teams, they achieved less than 3x increase. And what they
00:06:56found that was the difference between seeing less than 3x productivity increase, and these teams that saw a
00:07:01median of 4.5x, and in some cases more than 10, was how they used the tools. 90% of these teams use Kiro, among
00:07:11other internal tools that we have. And what they found was, it wasn't about the tools, it was about the way that they
00:07:18worked. The teams that achieved step function improvements intentionally changed the way that they worked,
00:07:26and the others simply kind of sprinkled Kiro and some of the other tools that we have on top of their
00:07:31existing way of working. And for me at least, this was the big aha moment, that why I hadn't been feeling
00:07:39potentially the massive gains that in productivity that AI has promised, it's about changing the way that we
00:07:47work. So across this pilot, they went and interviewed the teams that were involved in the pilot, as well as some of
00:07:55these other teams on the Bedrock Mantle team, on Prime Video, and they found five habits. And I use the
00:08:03word habits very specifically, because again, it's not about that one sprint, it's about doing this day to
00:08:09day. And what they found when they interviewed with these teams was that it really was habits that they had to
00:08:15build day to day. When we change our way of working, it's hard to build these habits, it takes time to build
00:08:22these habits. So let's go through each of these one by one. Habit number one is investing in agent context.
00:08:30We have a lot of stuff in our head, we tend to transfer all of that stuff in our head to other people
00:08:35through Slack conversations, through onboarding mentors, things like that through code reviews,
00:08:42through stand ups and sprint planning, and they had to write all of that down. And the habit that they
00:08:50built was every time the agent makes a mistake or does something not the way that you would have done
00:08:55it. What am I missing in my skills files? What am I missing in my steering files that the agent needed?
00:09:01But then as we know, across last year, we saw leaps and bounds in models abilities and their behaviors.
00:09:09The sonnet 3.7 in the middle of last year had a lot of quirks that we had to put a lot of do nots
00:09:16in our in our steering files. And now we don't have to do that as much with Opus 4.5 as of last November,
00:09:23and then we've had six months, more than six months of improvements since then, with all of the new
00:09:29versions of models that have come out since then. And so the question, the new habit again is, do I
00:09:35still need this in my steering files? Or is this just bloating context? The second one is slowing down to
00:09:41speed up. In almost every team that was interviewed, they reported that their productivity actually went
00:09:48down as they intentionally adopted a new way of working. That's counterintuitive,
00:09:53counterintuitive, right? You have to do intentional engineering work before you're going to see that
00:09:59hockey stick curve in productivity improvement. Because we have to do real work in our code bases
00:10:05first for agents to be successful there, especially in brownfield existing code bases. So they had to
00:10:11build that agent context up, they had to improve existing tools error messages so that the model knew
00:10:17what was going on when it failed, they built new tools, new MCP servers for helping that model to actually get
00:10:24done what it needed to get done. A lot of teams ended up restructuring their code base so that agents
00:10:29could actually navigate it more easily. And I've even seen drastic changes like changing the programming
00:10:36language of the code base. Often I've seen teams struggle with Python, with JavaScript because they're
00:10:43untyped languages. It's hard to test. There's no compiler errors. So the model kind of guesses and
00:10:50give it gives it back to you. And so I've seen teams moving to TypeScript. Rust has become very popular
00:10:56inside of Amazon. The compiler gives great error messages. You don't have to do that. But I've seen a lot of
00:11:03teams making those intentional changes for the productivity gains that they're able to see.
00:11:09The third one is feeding agents, not babysitting agents. And for me, this was one of those aha moments
00:11:16of why we're seeing this step function improvement in productivity. If you are vibe coding, if you are
00:11:23having a back and forth conversation with your agent all day long, of course, you're not going to see
00:11:30four to five X productivity improvements because you are in the loop the entire time. You're probably
00:11:36sitting there for 30 seconds to a minute waiting for it to generate code and come back to you with
00:11:42with the code to review. If you're sitting there waiting for it, then you can't go off and do other
00:11:48stuff. It's really difficult to run agents in parallel. It's very difficult to get to to clone yourself
00:11:55into multiple agents. And so if your conversations look a bit like this on the left, then you're
00:12:01babysitting that agent as opposed to the right side where you're feeding it what it needs to do and how
00:12:08it can self-validate. And that's really the key so that agents can self-correct and only come back to you
00:12:14when it meets a certain quality bar, when it when it actually runs and compiles and passes tests, when
00:12:20it's testable, when it actually has high coverage. And of course, the next level is put all of this
00:12:26content into your steering file. So it does it every time without you having to prompt it.
00:12:33The fourth habit is to make intent explicit. At Amazon, we practice a lot of spectrum and development.
00:12:40We've built that into the Kiro product. And so it's very natural for Amazon engineers to adopt it
00:12:46in Kiro. What what I've typically seen with vibe coding as opposed to frontier engineering is giving
00:12:54a very high level prompt, letting the agent generate a ton of code, and then having a back and forth
00:13:02conversation saying, Oh, that's not really what I meant. That you haven't you haven't exactly gotten the requirements right.
00:13:10No, I didn't actually want to build it that way. Here's a technical design. And it is less I find less productive
00:13:17to iterate with the agent on code when the intent itself was incorrect. So often we'll have while see Amazon
00:13:26engineers go through this process for ambiguous, complex features of writing the specification.
00:13:34And in Kiro, of course, you don't have to write this whole specification, you can have the model generate
00:13:39it. But it's a lot easier to iterate with the model in kind of a back and forth conversation about a
00:13:46document than it is about code that's code changes that are spread across a code base. The fifth one is shift testing
00:13:56left. One of the keys here is to give the agent that fast feedback loop, because that's what lets it go off for hours
00:14:04at a time and self correct. The agent is going to make mistakes, and that's fine. But if you give it the right signals, it can
00:14:12self correct and it can spend a while doing that. So I've seen teams adding linters, adding unit tests,
00:14:20integration tests, performance tests, security tests. These are all things we all know we should have
00:14:24been doing all along. This is good engineering hygiene and practices. But now the ROI is, I think, finally
00:14:32high enough for us to actually invest in it. One thing that I've been seeing a lot of teams do is mock out
00:14:40services. Often with integration tests, we would test kind of end to end an entire system, including live
00:14:46services. But we've been investing a lot in mock services that run entirely locally with deterministic
00:14:52responses, because it lets the agent do everything locally. Doing everything on your laptop without
00:15:01having to spin up a bunch of other services and connect to cloud services makes everything a lot faster.
00:15:07Because the more that your agent can get fast feedback means the more loops that it can
00:15:14can do and the more productive your own agent can be. So across all of these, these are some of the
00:15:21habits we've seen. But of course, I would be remiss if I would tell you if you adopt all of these habits,
00:15:28you will achieve nirvana, you will be the most productive engineering organization the world has
00:15:35ever seen. Things are still hard. We are still very much in an early adopter phase and teams are still
00:15:42figuring it out. So one thing that we've been seeing across our teams just organizationally is the risk of
00:15:48burnout. I did not coin this term, I forget who did at what conference, but FOMAT is real. We've been seeing
00:15:56engineers staying up late, late at night, trying to get that perfect prompt that's going to make their
00:16:03agent run for hours overnight so that they wake up in the morning with a code change ready. The cognitive load
00:16:09increases as you run these multiple agents in parallel, you're constantly shifting between
00:16:15terminal tabs. And then we do see that reviewing AI output is often harder for some than than actually
00:16:22writing it, especially early in career. Senior engineers have have already spent a large portion of their
00:16:29career reviewing others code. But early career engineers don't have that muscle yet. And so reviewing
00:16:37it can can feel like a lot more cognitive load than they're used to in actually writing it. The other
00:16:44one is organizational change. So it's already hard to change the way we work as engineers. The way that we
00:16:52spend our entire day completely changes when we're frontier engineers. But also organizations have to
00:16:58change to enable frontier engineering teams. One that I've seen very commonly is accepting slowing down
00:17:07to speed up. And I've been guilty of this myself. My fellow leaders have been guilty of this of saying,
00:17:14well, you have the AI tools now and the models are so amazing now. Why are you not going faster?
00:17:22And that's because you have to take those two months to invest in your code base to figure out the best
00:17:29best practices for your team to make hard habit changes on your team. And if you're constantly expecting
00:17:38shipping features every month, because now we have these amazing models, and we're seeing
00:17:44all of these companies on X saying how they're shipping 20 PRs a day, we have to slow down to speed up.
00:17:54The second one is actually going too broad in the organization too fast. I think that if we had
00:18:01expected all teams in massive organizations to be frontier teams immediately, we would not have had
00:18:08the learnings that we had from the pathfinder, from the from the sprint experiment, from the pilot
00:18:16teams within Amazon. And now the challenge for us is how do we scale it out? And that's what 2026 is about.
00:18:22For Amazon is how do we scale this out to more and more teams to the next 2000 teams instead of 50 teams.
00:18:31And so I think that when you roll it out too quickly, you have a lot of teams who don't know what they're
00:18:37doing. You haven't had time to find the best practices for your own organizations, the the context
00:18:43that your organization needs. And the last one is that you're going to find new bottlenecks.
00:18:49Previously code writing code manually was the bottleneck. I find that within Amazon, we've found
00:18:58the speed of decision making becomes a new bottleneck. The more that you spend reviewing the decision to
00:19:05actually build a new product, the slower it is to build the product now because the code only takes one to two months to write.
00:19:12All of the review processes associated with the launch of a product become the bottleneck.
00:19:20When it used to take nine to 12 months to build a new product, it didn't matter so much in the in the overall
00:19:27wash of things. If it took two months to make the decision to build the product and then two months to
00:19:33approve the launch. But now those are the bottlenecks. Those are the long pole. And so you find all of
00:19:40these all of these things that slow you down. Often I find that frontier engineering teams spend more time
00:19:48making decisions than they do writing code. And so the more that you can make fast decisions, especially ones
00:19:54that are easy to be reversed, the better. So my one big takeaway for for everyone here is that frontier
00:20:03engineering is about intentionally changing the way that you work. And that is difficult. That takes time.
00:20:10It is forming new habits and a new way of working. And that goes across any engineering team as well as your
00:20:18organization. So I encourage you to think about how you're interacting with AI tools and how that can
00:20:27change to free yourself up from being in the loop. Thanks. I'll hang out a little bit if anyone has
00:20:34questions in the back. But thanks for the time today.

Key Takeaway

Achieving step-function productivity gains of 4.5x to 10x requires frontier developers to intentionally change their habits by feeding autonomous agents with context rather than babysitting them.

Highlights

  • Amazon teams running frontier development pilots achieved a median 4.5x productivity increase, with some teams exceeding 10x.

  • Frontier developers write only 1 to 2 percent of the code they produce, relying on autonomous agents for the rest.

  • The Bedrock Mantle team built an inference data plane in 76 days with six people instead of the estimated 30 people over 18 months.

  • A Prime Video 10-day sprint reduced project delivery time estimates from 90 weeks down to 24 weeks.

  • Frontier engineering teams spend more time making decisions than writing code because decision-making speed replaces code writing as the primary bottleneck.

Timeline

Evolution of AI Coding Assistance and Frontier Development

  • AI coding assistance evolved from inline completion and chat to vibe coding and autonomous frontier development.
  • Amazon pilots show a median 4.5x productivity improvement, with select teams reaching over 10x.
  • The Bedrock Mantle team built a major inference data plane in 76 days using six people instead of 30 people over 18 months.

Previous phases of AI coding tools yielded modest 10 to 20 percent personal productivity gains. By shifting to frontier development behaviors, teams achieved massive efficiency leaps. The Bedrock Mantle team proved this capability on a greenfield project using the Kiro assistant, though that team consisted of distinguished distributed systems experts.

Structured Sprints and Real-World Developer Realities

  • A Prime Video 10-day sprint cut project delivery estimates from 90 weeks to 24 weeks using six engineers.
  • Amazon Stores monitored 50 normal teams working on existing brownfield systems over the better part of a year.
  • Teams using Kiro without changing their workflow saw less than a 3x increase, while teams that intentionally changed workflows achieved median 4.5x gains.

Controlled sprints demonstrated high velocity, but they relied on idealized conditions like zero on-call duties and meticulously pre-scoped tasks. Real-world store teams showed that the tool alone did not guarantee success. The difference between moderate and exceptional gains depended entirely on whether teams intentionally restructured their working methods.

Five Habits of Frontier Developers

  • Frontier developers invest heavily in agent context by documenting tribal knowledge and rules in steering files.
  • Teams slow down to speed up by refactoring codebases, improving error messages, and adopting typed languages like TypeScript and Rust.
  • Developers feed agents rather than babysitting them, allowing agents to run autonomously for hours and self-validate via automated tests.

Adopting frontier development requires building specific daily habits. Teams write explicit intent specifications rather than engaging in conversational back-and-forth about raw code. Shifting testing left with local mock services and robust linters provides fast feedback loops, enabling agents to execute long-running, autonomous tasks without human intervention.

Organizational Challenges and New Bottlenecks

  • Frontier development introduces risks of developer burnout and cognitive load from managing multiple parallel agents.
  • Organizations struggle with accepting the initial slowdown required to establish agent context and best practices.
  • Decision-making speed replaces code writing as the primary bottleneck in product delivery.

Rapid tool adoption creates organizational friction when leaders expect immediate output before teams establish proper habits and context. Managing multiple terminal tabs and reviewing AI-generated output increases cognitive load, particularly for early career engineers. Furthermore, as coding time drops to months, product launch reviews and decision-making become the long poles in the delivery cycle.

Community Posts

View all posts