How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage) — Eyal Blum, Figma

AAI Engineer
Computing/SoftwareManagement

Transcript

00:00:00Eyal Blam: Good afternoon, my name is Eyal Blam, I am a software engineer at Figma, and
00:00:17in my talk today, we're going to talk about how we've adapted or are adapting agent into
00:00:25our workflow at Figma while maintaining a high quality for our code base.
00:00:32So as you may know, Figma is the browser-based editor where design and engineering and now
00:00:40AI agent collaborate together to ship code.
00:00:44This Figma has pivoted very strongly from being a traditional tool to an AI-first tool,
00:00:52but in this talk I'm not going to talk about our product, I'm going to talk more about our
00:00:55internal organization and how our engineering org has been adapting AI agents.
00:01:04And what we found internally is both organizations, companies, and individuals, there's kind of
00:01:11a three-act process of AI adoption.
00:01:16You start with picking up something, whether it was a lot of the people in this room who
00:01:21are very AI-pilled and have been using AI for a while and they picked up something and got
00:01:27some simple things to work very well, 10x faster.
00:01:30Then you start applying those same practices to bigger problems and AI fails pretty badly at that, gives you bad stuff,
00:01:40lots of bugs.
00:01:41The trust that you build breaks down and then from that point you start building the real skill, which is
00:01:50learning how to use AI correctly and put the right guardrails and the right prompting and the right
00:01:55contact and all the stuff that we've been talking all day about here and all the talks in order to actually build a real skill.
00:02:02And one thing that is happening internally is we adopt whether teams or individuals, the adoption is uneven.
00:02:12We have teams that are very AI forward and have already transformed their entire workflows and then we have teams that are still experimenting in the earlier act and/or have lost confidence and they all need to work together in order to ship our product.
00:02:28So, they need to co-exist in the organization and we need to find a way to support them while bringing on everybody along for the journey and getting everybody to the third act of the story.
00:02:44Aside from that main friction point, we have also noticed other friction points that happen as we adopt AI.
00:02:54One thing that we've heard a lot from developers and managers have been noticing is that reduced developer agency causes engineers to lose some of their job satisfaction.
00:03:04So, a lot of people used to take a lot of pride and enjoyment in writing code and getting into the flow and a lot of people feel like that's been lost or they're losing a lot of that element and getting into more of a prompt cycle where they just wait and output from AI and then speak to the AI that's not as much fun as they used to have.
00:03:22We've noticed another interesting thing is actually our best engineer, they want to hold all their contacts in their brain and they end up getting out of the burden and what ends up happening is they know where all the pitfalls are, they are holding together with their mental duct tape all the places that agents are not working well and they're preventing all the really bad stuff from coming in or they have all the institutional contact that haven't been
00:03:52never written down in their head and never written down in their head and they get so much burden and become bottlenecks and get really frustrated so they actually end up being slower to adopt because they see all the problem firsthand.
00:04:04That's another big issue that we've seen and this one I'm sure everybody here will resonate and all of a sudden all the design docs and all the slack messages and all the emails have gotten three or four times as long and we've gotten two or three times as many emails and they say basically as much as they did before.
00:04:28So the communication has gotten quite inefficient and some of the markers of like what is high quality and important things versus not so much high quality and has become challenging to navigate.
00:04:40So I'm going to spend the next few minutes talking about some of the lessons that we've learned and how we've been trying to apply this.
00:04:50This is a journey we have not come out to the other end but we've seen some really interesting progress along a lot of these lines.
00:04:57I think a lot of the speakers here have touched upon this but investing in verification is probably the highest value thing we can do in our code base.
00:05:11Anytime we can left shift anything in our workflow from a human needing to do it to an agent being able to verify it.
00:05:19So for example when playwright MCP came out instead of having humans navigate the code now the agent can explore the code that was a big win unlock for productivity in a lot of our team.
00:05:33That's really that's always a big win for us.
00:05:37The other thing is even better if when you find something that the agent has found to be useful take the time to take it and encode it into a deterministic flow.
00:05:51And deterministic flow that can be easily repeated it's saved on token and saved on time for the and then it also you also know that you're using the LLM when it needs to reason.
00:06:01But when you have something that is already known and basically can be encoded into a test spending that time always always pays dividends.
00:06:11And another tip if you tell your skills or your agent to write the code that you're writing and like at the red to green at the TDD style it almost always gives you better results.
00:06:26Because you set a goal then you tell the agent to strive toward that goal.
00:06:30It will almost always give you better results in writing the code and then writing the test afterwards because then it will fit the test to the code rather than fit the code to pass the verification criteria.
00:06:42And this is the testing pyramid the classic testing pyramid from the previous and just when you think about the testing themselves which you had the end to end test when the integration test and the unit test.
00:06:54This is very similar move as much as you can down to the deterministic analysis where that's linting the compiler and the unit test themselves whatever that can be covered easily you can have engine agent to reviews on it based on criteria.
00:07:12And architectural standards that have been.
00:07:15And architectural standards that have been easily encoded into the code base you can move into the agent.
00:07:19And then only at the very top you need to have some sort of human review which is usually around the functionality and this is the right thing to build that like only leave the human to do what the humans need to actually be involved in.
00:07:31And another really important thing is the planning versus prompting this is really tied into the giving agency back to developers and finding a replacement to the craft of writing code.
00:07:50So spending a lot of time writing the plan and then sending it off to the agent basically as an implementation that can be done automatically is something that we find to really kind of reintroduce the joy of building back into the process.
00:08:07And so it's not uncommon to spend a week writing a very detailed plans making all the decisions flushing it out iterating sending it out to teammates to review.
00:08:17And then only when it's ready and you've flushed out all the decision can send it to the agent the agent will send it back to you when it's implemented.
00:08:26And that that has been really successful also in accelerating and also really restoring some of the joy into the development process.
00:08:36And so what makes a good plan and really important to start with the why at the top.
00:08:43It really helps preventing agent drift if you have like a bold big section of kind of like when you write a design doc.
00:08:49You want to have the executive summary put that in there for the agent otherwise they'll start drifting over time and make sure that the agent don't go back and change that because they feel like it.
00:08:58And so we start with the why make sure that the plan can be broken down into small parts that can each be verified independently.
00:09:08And my personal way of knowing what is a good size would I want to review that the PR that will correspond to that part if it's going to be too big for me to want to review in one sitting.
00:09:18It's kind of like the test is I'm going to need to get a cup of coffee before I read this.
00:09:22That means it's too big and I'm going to want to have it broken down into pieces.
00:09:26And then I make sure that each part can be validated independently because what I don't want to have is have five stages and then the first one is written but not validated and then everything else is built on top of all the assumption.
00:09:41So having kind of a validation gate or an exception criteria for each phase really helps make the plan resilient to drift.
00:09:51And there's all kind of technical how to manage the context and doing a software factory on top of doubt.
00:09:58But once you have the plan you can use whatever loop you want or whatever workflow you want in order to implement it.
00:10:05And this is a screenshot that I randomly picked up a plan but this is what I usually look for.
00:10:12The executive summary at the top.
00:10:14The phases break it down and then each one of them I will go into lots of details so that I can just fit it into a sub-agent.
00:10:20And the sub-agent can independently work on that and not have to worry about it.
00:10:24And there are other workflows that would work or other structures to the plan.
00:10:29I find that part of the things that great about AI workflows is that everybody can set up the thing that works best for them.
00:10:41No thank you.
00:10:43Everybody can very easily set up the workflow that works exactly for them.
00:10:47So the diminution return is trying to centralize everybody on one thing.
00:10:51But as long as it works for their flow and other people can iterate with them I find that it generally works very well.
00:10:58And this is just an example kind of a brag of like this could be a result from a plan.
00:11:04And there are probably 20 PRs here.
00:11:07Some of them would be maybe 10 lines and some of them would be 100 lines but probably nothing bigger than that.
00:11:12And that allows us to -- this is in the pre AI wall this plan probably worked on it for a week.
00:11:19I aligned with three other teams for another week on that and then I just send it to an agent to implement overnight.
00:11:26And it came back -- this is probably from two plans, not one -- but it's basically six weeks of coding work just -- it only took one week.
00:11:36So that's where I get the five X speed up.
00:11:40If I include the review cycle at the end that you always have to remember.
00:11:45And moving on from planning and back to the issue that we had with the skeptics and the people who are burdened with the most work.
00:11:53Make sure that you bring them in and take their feedback really seriously.
00:11:59They are spec haptic because they are seeing where you are lacking validation, where your tools fail.
00:12:05So their feedback is basically the roadmap of how to improve your agent and interacting with the code base.
00:12:12Just make sure to bring them in rather than trying to figure out how to make them use AI.
00:12:19Just have them be in charge of the roadmap to make AI safe in your organization.
00:12:25And they will come along once they see that the improvement that they're making are actually making their life better.
00:12:33And as you can see, they'll not be shy about telling you what you need to fix.
00:12:37This is less than an hour sitting with a bunch of people.
00:12:40And this is the result of brainstorms.
00:12:45Another thing that's been really helpful with my team specifically, and we're working to adopt it in the broader organization as well,
00:12:53is to make sure that you have an attention-aware communication.
00:12:58In the age of AI, human attention is a scarce resource.
00:13:01I think I've heard it for multiple talks, and a lot of people have come to the same conclusion.
00:13:06You can get more human attention.
00:13:08So where you spend your time and what you're reading becomes really important.
00:13:13And so since it's such a scarce resource, marking what was generated by AI versus what was written by a human is really helpful to know how much time you need to spend reading this,
00:13:24and how much slope can you expect in this part of the communication.
00:13:31And that's kind of building a new culture around that style of communication.
00:13:35It really helps.
00:13:36And so for example, the team that I work with, we've decided we always -- every PR description will start with something like that.
00:13:45Something that I wrote by hand could be very short to describe what this is doing, and then the AI description is going to come after that.
00:13:54It's just -- I will probably read it.
00:13:55I will probably edit it to remove some wrong things, but they didn't write every line here.
00:14:00So they should be more suspicious, and they should pay more attention to what I wrote in the top, and they should override it.
00:14:05Things like that in Slack, in email, just like leaning into the fact that everybody knows that you're using AI to craft your communication,
00:14:15but just don't be shy about telling them what you should read and what they should pay less attention to.
00:14:21And I remember early on, maybe like earlier in this year, I tried to -- I had some senior engineers in our org that had kind of were very much AI skeptics,
00:14:35and I tried to reach out to them to see what was the problem, what was going on, and say I tried to run an analysis on some of the PR comments that you've run,
00:14:42and obviously I used AI to do that.
00:14:45And then I didn't distinguish very clearly what I wrote versus what they -- what AI generated, and they got very upset.
00:14:54They're like, why are you sending -- I did not expect somebody that I respect this much to send me something that's clearly this sloppy.
00:15:02And then, like, I -- like, I apologize.
00:15:05I realized I should have marked it clearly and marked my intention.
00:15:08Like, this is what I wrote.
00:15:10This is what the AI wrote, and I need your feedback on it because I don't have the context to know if it's sloppy or not.
00:15:15And that's what I'm asking you for.
00:15:17So, lessons like that and changing the culture is just as important as some of the engineering challenges that we've been facing.
00:15:28Another thing that's really helpful around the adoption is, as you progress to adoption, there's a lot of very fancy tools and a lot of very fancy workflow that we've been implementing.
00:15:41But one of the really effective things is just letting people use AI where they're at.
00:15:47So, it really helps normalize the use of AI for everyday tasks, and it helps reduce the friction.
00:15:54And really, one of the most powerful things is being able to tag an agent in a Slack message with somebody and say, can you just do this for me?
00:16:02And have the agents close the loop in the thread?
00:16:07And that kind of thing is really powerful.
00:16:09And then you can go on top of that and have all this being automated and be all kind of fancy things.
00:16:14But if you're having a conversation with somebody who's not fully bought in, and then you can tag it in a non-passive-aggressive way, you can tag it and say, let's try it and see if the agent can get it this time.
00:16:26And they close the loop, and if it's a good experience, that really helps people try it out on their own in other cases.
00:16:33And our journey continues.
00:16:37We're still learning, even though we're shipping AI externally, our AI adoption, and we're experimenting with so many things all the time, where our automation story is not fully there yet.
00:16:50We're still trying to figure out when we should use, how we can use Cloud Agent effectively, given all the dependencies we have for some of our build system.
00:16:58And so we are continuing to learn.
00:17:01It's a culture shift.
00:17:02It's an engineering shift.
00:17:03And I don't know about you, but I've been working in the Valley for the last 15 years, and this is the biggest change by orders of magnitude of everything that I've seen in terms of culture and technology.
00:17:16So we're all here together, and we're all figuring it out.
00:17:19And that's what I wanted to talk to you today.
00:17:22Thank you.
00:17:28Thank you.

Key Takeaway

Implementing a structured three-act AI adoption process with detailed pre-planning, shift-left verification, and attention-aware communication enables organizations to scale coding agents without degrading code quality or developer satisfaction.

Highlights

  • Figma's internal AI adoption follows a three-act process involving initial 10x gains, subsequent failure on complex tasks, and the development of proper guardrails and contextual prompting.

  • Comprehensive planning before coding involves spending up to a week writing detailed implementation plans and breaking them into small, independently validated parts.

  • Shift-left verification utilizes the Playwright Model Context Protocol (MCP) to allow agents to independently explore code instead of relying on human navigation.

  • Attention-aware communication practices require explicitly marking AI-generated content in PR descriptions, emails, and Slack messages to manage scarce human attention.

  • Engaging AI-skeptical senior engineers by placing them in charge of the roadmap for safe agent usage leverages their firsthand identification of tool failures to improve the system.

Timeline

Phases of Internal AI Adoption and Emerging Friction Points

  • Internal adoption follows a three-act curve moving from initial high productivity to failure on complex tasks and finally the acquisition of proper prompting and guardrail skills.
  • Reduced developer agency leads to lower job satisfaction as engineers shift from creative coding into a repetitive prompt-and-wait cycle.
  • Best engineers act as bottlenecks because their unwritten institutional knowledge forces them to carry the mental burden of patching agent failures.

Figma transitioned into an AI-first tool, prompting an examination of how internal engineering teams adopt AI agents. Adoption is uneven across teams, creating organizational friction between AI-forward groups and experimenting teams. Additionally, design docs, Slack messages, and emails have expanded in volume while maintaining the same information density, which complicates quality navigation.

Investing in Verification and Shift-Left Testing

  • Investing in verification shifts tasks from human execution to agent-driven validation, such as using Playwright MCP for code exploration.
  • Encoding useful agent discoveries into deterministic flows saves tokens and execution time while reserving LLMs strictly for reasoning tasks.
  • Instructing agents to write code using Test-Driven Development (TDD) red-to-green styles yields superior alignment with verification criteria.

Verification represents the highest-value investment in a codebase. The testing pyramid is shifted downward toward deterministic analysis, including linting, compilers, and unit tests, leaving human review strictly focused on high-level functionality and architectural standards.

Pre-Planning and Context Management for AI Coding

  • Writing detailed implementation plans over the course of a week restores the joy of building and prevents agents from drifting off task.
  • Plans must start with the 'why' in an executive summary and break down into segments small enough to review before needing a cup of coffee.
  • Validating each phase independently prevents errors from compounding across stacked sub-agent tasks.

Planning replaces the traditional craft of writing code with architectural design. A well-structured plan allows a developer to coordinate with other teams, write specifications over a week, and then dispatch implementation to an agent overnight, achieving up to a 5x speedup including the review cycle.

Cultural Shifts, Skeptic Management, and Communication

  • Involving AI-skeptical engineers in tool improvement harnesses their critical feedback to build effective safety guardrails.
  • Attention-aware communication requires marking AI-generated text in PRs and messages to signal appropriate reading scrutiny.
  • Allowing teams to use AI organically within existing workflows like Slack integrations reduces friction and normalizes everyday usage.

Managing the cultural shift requires treating human attention as a scarce resource. Explicitly delineating human-written text from AI-generated text prevents friction during peer reviews. Allowing developers to tag agents directly in Slack threads helps bridge the gap for team members who are not yet fully bought into the workflow.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video