Measuring Developer Experience Impact with Cost to Serve | Amazon Web Services
Key Takeaways
Amazon Web Services measures developer experience impact using the cost to serve metric, which captures systemic improvements in software development and delivery, and is calculated as the cost across all development efforts and tooling versus the amount of units of software delivered for value to customers. The metric is used to drive innovation, reduce technical debt, and improve developer productivity.
Full Transcript
[music] Hello and welcome to the Executive Insights podcast powered by AWS. My name is Mark Schwarz. I'm an enterprise strategist with AWS. And with me is Jim Howitt, the VP of Amazon software builder experience. Jim, could you start by telling us a little bit about your background and your role at Amazon? Hello Mark and thank you. Um I've been working in the software industry for about 30 years. I've been leading uh platform and software enablement teams for about 15 uh some of the leading companies and largest customers of AWS in the world. My job at Amazon essentially is to make it easier for all the software engineers that build the software of Amazon retail AWS Prime Video others. make it easier for them to build, deliver, and operate software for customers. >> This sounds like a pretty big mission. >> It is a very big mission. Uh we have tens of thousands of engineers. We uh have hundreds of millions of customers around the world. Uh and but it it the ability to have all of those teams be able to innovate faster. And if you think about the impact of having thousands of teams being able to innovate faster is tremendously rewarding for us. and for what we can do for Amazon and its customers. >> Wow, that's amazing scale. How do you justify the investments that you make in improving the builder experience? Is everybody on the senior leadership team aligned behind investing in these things? >> When we created the organization, uh our leadership team was on board and and very interested in improving the experience based on the anecdotes that they had heard from their teams. Um they wanted us to be able to go faster, to go better. Um and they challenged us to go off and and commence in that mission. We were reporting doubledigit improvements in many numerous aspects from how quickly people can get started on work to how often they were bringing software changes to customers uh to reductions um in manual work uh to improvements in safety. And about a year and a half ago after we had gone through continuously improving uh we got an an interesting question that was like these are all great we're very happy to see that but what do we get from that from Amazon as a business so that it can help us in the investment level of decision and really understand how do these improvements and experience translate to value for our customers. >> That sounds like the CXOs I know. Yes, >> it is a it's a question that uh I have heard um prior to work at Amazon. It's one when I speak of other customers, they have that question and it's people believe it's it's good to improve the experience of engineers and let them do more. But then how do you compare that side by side of other investments you made? >> Sure. >> So we we we took that challenge. We went back and we wanted to look at how do we answer that? And one of the lucky points is working at Amazon is we have lots of inspiration to draw on. And we found that years ago in our retail business, the challenge of how do we bring goods to customers was one. How do we improve that? uh and our retail business at Amazon developed a metric called cost to serve which captured the entire systemic improvements to every aspect of taking a unit that somebody orders and packaging it up and bringing it to their house and we realized we could translate that to software. >> So can you tell me a little bit more about what this metric is exactly? How do you calculate it? So at the highest level, cost to serve is the cost across all of your development efforts, your development tooling that it took to build and deliver software for use by customers uh over the the amount of units of software that you delivered for value to your customers. It is a system level metric that looks at the totality of value that a team has delivered and how are you implying improvements to make it better. The goodness of the cost to serve framing is that it meas it it measures the the quality of the system. So anytime you take friction, sources of delay, sources of loss, sources of defect out of the system, you essentially reduce the cost to serve. So if you bring products to people sooner, you reduce it. If you eliminate defects and returns or reduce them, you reduce cost to serve. We realized software which is complicated and has everything from planning to design to development to operations had some very similar fundamental concepts. At the end of the day an engineer wants to build something that creates customer value. So your units are the code that you build. You package that into applications or services and the value is realized when customers are actually using it. And we looked at the overall system and said anywhere we can take defects, delays, friction, waste out of the system, we give essentially and teams back more time to innovate for customers. >> That's really interesting. So cost to serve encompasses all of these aspects of delivery essentially. So it's it's like the ultimate metric in a way. >> Uh are there things it doesn't measure? >> There are lower level things that it doesn't elevate to the top level surface. So like all complex systems we we measure um multiple we have metrics dozens of metrics across different components. But when you bring it the top level cost to serve will have both output and quality. If you have a reduction in quality your cost to serve will essentially go up because you have to do rework. If you have a team that gets into an unhealthy state and you have people leave, you got to bring more people on board. So, it actually helps you correct for that because it'll show your cost to serve will go up where if you keep the team happy and healthy, your cost to serve goes down. If you let your technical debt get out of uh control, you're going to slow down as well. So, it's it's got good control structures. But that being said, since those are only recognized as second order effects, what we do is we bring tension metrics up alongside cost to serve. So, we can take primary concerns and elevate them for attention. One of the ones that we look at is the sentiment of a team. So you want to make sure as the team is enabled to do more, they're also feeling happier. Um, we find in our surveys that teams that feel like they can get work done for customers regular and often actually are happier, which doesn't surprise me cuz people come to Amazon because they want to create impact. >> Uh, another that we look at is making sure safety. So we look both at the security coverage of our software and risks and we look at the availability impacts of releases. And we want to show that as you release more software, your availability uh and and events that are happening are getting better and not worse. And as you're releasing more software or you're doing more work, your your security is getting better. And we're happy that we've been able to both increase what teams are able to do in tandem of improving safety. So I I guess this helps with the classic problem where if you if you define a metric and use a metric, people find ways to game it. Is this uh tension metrics essentially work against that, right? The power and cost to serve and we discovered this as we started to dig in is the value comes in year-over-year improvement or quarter over quarter improvement. So if you're a team that is releasing software on a weekly basis and you go well if we do it daily um we release in smaller pieces and our cost to serve is lower that's good but the value is as you go quarter by quarter year by year if your team is getting better if they're reducing defects if they're reducing outages essentially they're just going to reward off that baseline and as they improve as a percentage they're you they're generating value. So there's a built-in safeguard against that. Similarly, if you have a manager pushes the team too hard, says don't worry about uh technical debt or uh the team just gets burned out, the systems will slow down, people will leave, they'll have to bring them back on board and it becomes it's a correcting factor. So it captures that. So one of the wonderful discoveries we saw there's built-in resistance and then when you add the tension metrics and you expose additional data to people you you can help them focus on improving the goodness for kids. >> I see. U so this sounds kind of magic. Uh do you actually use this? We started using it last year to prioritize our investments and then um re in the last year we've started to also use it because we can translate the value you get not just into dollars but into effective capacity and as as leaders say came and say how can I help my team do more we can use this to show if you adopt these techniques these technologies you give your team effective capacity to do more of what they want to do which is their road map and the features and capabilities they want to build for customers and we're now seeing teams bring that into their planning one of the as joys and challenges of Amazon is how do you bring data to teams where they work uh so they can make decisions uh and we believe our in our teams being owners so if you give them data in an easy fashion they can make decisions and if we give we also give data sets and guidance to our leaders in the company at the director and uh and VP levels that these are the things you can prioritize for your team so you can get more capacity and your teams can do more um and it allows them to create space or help prioritize uh some of those investments as well. So this really empowers the teams and that's pretty amazing with so many teams. So, uh, a lot of our listeners to this podcast, viewers, um, they're leaders of large organizations. Do you think this is something that can work for other organizations besides Amazon? >> We we do think it can work in in part. Amazon is has teams that work in retail. We have teams that work in hardware. We have teams who launch uh, satellites uh, into space. We have teams that do cloud services uh, and foundational models. So we we cover a broad spectrum. One of the things that we found in our research is the the most important influencer on an individual developer's experience and productivity is their team and how their team builds and ships. And we found in different business divisions, different technology types, people build and ship differently. And when we allow in cost of service, you can plug in the model that your team uses. Each of those teams can say, "This is how we work. we plug this in and we want to see how we improve. Uh and that allows you to not just take an an odd new metric but to take the way you're working and plug it in and show value. Uh that's based on how you deliver value for customers. It's interesting you framed a lot of things in your answer there in terms of teams. Uh you know I think we we hear a lot of questions about developer productivity these days. Uh but you seem to be looking at it from the perspective of teams rather than individual developers. Am I am I getting that right? >> One of the things I'm most proud of that we do is that we not only did we foundationally believe that teams were the unit but we did the science research uh over 5 years of data and we found that how the team ceremonies how they ship how they work are the biggest predictor of an individual's productivity. Uh so that uh led us to focus on how do we enable teams to make the right investments to use the right tools uh to make their lives better and then that rising tide lifts the boats of all members on their team. From that perspective uh individual performance is something between an employee and their manager. Uh but developer experience is a team sport and if you can help the team play on a better field uh they and give them better equipment and tools they can do better. What are some examples of changes you've made to the developer experience to improve the cost to serve? >> One is making it easier for teams to bring their software to customers which is the the build and deployment process. Uh so we did a lot of investments in uh deployment safety and deployment automation and that's that's had two effects is one it's enabled teams to just be able to bring software to customers more easily uh by doing less work but having safety improve they actually have more time to go back and build the things they want to build. So from that we've seen both software gets to customers uh faster uh software takes less work to release the fact and we've actually watched as a result teams are able to release more changes to customers as they move along. The second area is we've looked at different types of abstractions and templates that we can give them. So that essentially instead of building everything from scratch, you can grab a template uh golden path off the shelf, get started and start building from that. And if we know that some of these templates actually are easier to maintain or we can centrally uh manage and patch them on behalf, then you can spend more time building on your differentiated work. And in the last uh few quarters with the rise of AI uh topic on things we're now start we're seeing it growing in its impact uh in its effect to cost to serve and what is very interesting about AI is there's been a lot of discussion of lines of code and how many lines of code did the AI do. Um, lines of writing code is just one of many jobs to be done by engineers because cost to serve is a system metric and it captures the entire experience. If you use AI to help you find information, resolve defects, uh, diagnose problem tickets, go through logs, all of those benefits start to roll up, and as we're bringing AI into more places, uh, we're helping builders and more of their jobs to be done. and it's and it's it's reducing cost to serve. So, we're excited about where it will go. >> Yeah. Before I ask you a little more about AI, uh there's one thing that's bothering me and what you said or worrying me. I once was a software developer and I've led software development organizations and I I don't know my perception is that developers are pretty opinionated. You know, that they want their tool chain, they want their deployment process that's their favorite. Do you um do you mandate what process they should use? How can you get these uh cost to serve improvements without having control over what they're doing? Essentially, >> yeah, you are correct in that uh engineers like to pick their tools. >> They do. >> And we drive our decisions by a mixture of anecdote, data, and expert judgment. And our anecdotes come from uh many from surveys and user research. and the ability to pick your tool is one of the key aspects of autonomy. So the way that we combine giving people choice and helping them reduce is we take the most common tasks uh patching software deploying across dozens of regions uh doing security scans maintaining large code repositories that teams could do but it would take a lot of time away. It's the work. It's the less it's what we call low differentiated or undifferiated work. >> The toil >> and the it's toil. Um but it's also just it's it's not the things that's going to differentiate your value to customers, right? >> And we provide those to our teams to use. Um and then they pick additional tooling on top of that that makes them uh effective. Um and that combination we can tackle the large section of undifferiated work. They can tailor uh things as needed. And we have so many teams that some will have specialized needs. And where we bridge the gap is we do provide guidance uh to people and we say if you want to use this template you can go faster but if you have unique needs you know you can choose your own but we share like teams like yours who do things like this get capacity back. >> All right now to the important topic AI. Um what what's happening? I guess it's changing the way software is developed in in uh in a big way. What kinds of changes are you seeing? What do you expect? How do you think it's going to affect the developer experience? We've seen a pretty significant evolution in the last 18 months and the pace of change is increasing. So early on we saw a lot of autocomplete type work. Um uh last year our team did a lot of work of using transformations to enable teams to do large scale upgrades, language upgrades. Um so instead of them having to do that, we used AI and tools to do that for them. A famous example uses our JDK17 upgrade we did last year. Uh where the average team had 58 se separate dependencies and it would have taken them weeks to manage through them. and we use AI to essentially reduce that work to uh just a day uh for folks and realize a lot of benefits. So in the last six to nine months as agents have really taken off. uh we started to see new patterns of work uh which essentially were where people would give a a task they would describe what they want to the agent it comes back they look at it they say I like this I don't like that they give it additional instructions so they're they're still an engineer they're now auditing critiquing and improving instead of learning a new language learning how to do you know doing kind of more of the repetitive configuration changes and the like. Uh so the some of the examples we've seen is people can learn new code bases faster. They can make improvements to code bases that they don't know that particular programming language. Uh we found usability accessibility improvements can be done very quickly. And part of the things that we do as as a company is we've opened up the space to experiment. So we've had some pioneers who wanted to share their ideas and part of ASBX we did is we gave them a large platform live streams uh video sessions you can play back knowledge series to share that knowledge and we've now seen that picked up by uh 20,000 plus people in the company like in in one single channel and we now have tens of thousands of people who are now riffing on those experiments sharing new ideas and then We look at those best ideas and we bring them into our products. >> How have the developers responded to it? Do they enjoy working with AI? >> Every month we we talk to a few thousand of our engineers and we ask them what do they think? Uh and the three things they've told us say is it they save time, they feel more productive and they'd be unhappy if we took the tools away. When you have anecdote and sentiment and data all going in the same direction, that's a good sign. I love that example of the JDK upgrade because there's nothing more boring than doing that kind of work, right? Exactly. I could imagine the developers would love having some relief on that. >> It's upgrades, patches are fun the first time you do it cuz you learn how to do like high-grade professional software. The third time you do it, they become toil, but there's something you always have to do. Your phone is always updating new apps. So, making that go away. Um, and we're looking at ways within our ASPX to make more and more of that go away because then you just spend more time on the innovation you want to build. >> And how do you think Amazon Curo will affect the developer experience? I think you've got some developers who are using it now, right? uh products like curo and the broader agent problem uh products uh based have a very big aspect of developer experience that's more than just lines of code and the ability to help developers and all those jobs to be done will let them basically be able to take more of their ideas to customers faster and what we're talking about today of cost to serve is that encapsulates the whole system and how do you improve and reduce friction and make things faster and speed getting answers and the entire system of the jobs be done for developers. >> So, uh Jim to close things out could you tell our audience maybe what are maybe three important lessons that they can learn from what we've been talking about? The first I would say is that you can improve developer experience in all of its complexity and demonstrate business value that your seuite can see that can be done um and without sacrificing the experience. The second is that value you can use in three different ways. Simple dollar savings, return on investment to plan your investment or my favorite capacity that you enable teams innovate and build more. And third, you can start small. You can start with a division or a team. You can start with existing frameworks you're using around velocity and plug them in um and start to recognize value and use that as a flywheel to generate value and encouragement in the rest of your company. That's great. That's that's great advice on how to actually transform an organization using a metric like cost to serve software and by improving the developer experience. So, uh, thank you so much, Jim, for being with us and for giving all this helpful advice. [music]
Original Description
Discover how Amazon measures and improves developer experience at enterprise scale in this interview with Jim Haughwout, VP of Software Builder Experience at Amazon. Drawing from Amazon's retail business expertise, Jim reveals how the "cost to serve" metric transforms developer productivity measurement by focusing on system-level efficiency rather than individual performance. Learn how Amazon balances developer autonomy with standardization, implements tension metrics to prevent gaming, and leverages AI to re-imagine software development workflows. This essential discussion provides practical frameworks for quantifying developer experience improvements and connecting them directly to business value, from deployment automation to AI-assisted development at scale.
Learn more about AWS Executive Insights: http://go.aws/4nNKXY9
Subscribe to AWS: https://go.aws/subscribe
Create a free AWS account: https://go.aws/signup
Try AWS for free: https://go.aws/free
Connect with an expert: https://go.aws/contact
Explore more: https://go.aws/more
Next steps:
Explore on AWS in Analyst Research: https://go.aws/reports
Discover, deploy, and manage software that runs on AWS: https://go.aws/marketplace
Join the AWS Partner Network: https://go.aws/partners
Learn more on how Amazon builds and operates software: https://go.aws/library
Do you have technical AWS questions?
Ask the community of experts on AWS re:Post: https://go.aws/3lPaoPb
Why AWS?
Amazon Web Services is the world’s most comprehensive and broadly adopted cloud, enabling customers to build anything they can imagine. We offer the greatest choice of innovative cloud capabilities and expertise, on the most extensive global infrastructure with industry-leading security, reliability, and performance.
#AWS #AmazonWebServices #CloudComputing #DeveloperExperience
#CostToServe #AIInSoftwareDevelopment #AmazonEngineeringCulture #AWSExecutiveInsights
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Amazon Web Services · Amazon Web Services · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Agentic AI Design Patterns Introduction and walkthrough | Amazon Web Services
Amazon Web Services
Galileo on modernizing on banking infrastructure | Amazon Web Services
Amazon Web Services
Alliander Speeds Innovation and Energy Transition Using AWS | Amazon Web Services
Amazon Web Services
AWS and Scuderia Ferrari HP streamline F1 power unit assembly | Amazon Web Services
Amazon Web Services
How AWS machine learning supports Scuderia Ferrari HP pit stops | Amazon Web Services
Amazon Web Services
Nasdaq Builds Market Infrastructure of the Future with AWS | Amazon Web Services
Amazon Web Services
AWS Security Hub Exposure Findings | Amazon Web Services
Amazon Web Services
How do I use Session Manager port forwarding to connect to my EC2 instance through RDP?
Amazon Web Services
How do I extend an EBS volume with LVM partitions?
Amazon Web Services
AWS Graviton makes it easy to optimize performance, cost, and sustainability | Amazon Web Services
Amazon Web Services
Run Cloud Adoption Framework workshops with Miro | Amazon Web Services
Amazon Web Services
Getting Started with AWS Cost Optimization Hub | Amazon Web Services
Amazon Web Services
Why did my Amazon SQS messages get sent to a dead-letter queue?
Amazon Web Services
Declarative Policies for EC2 | Amazon Web Services
Amazon Web Services
How do I troubleshoot IAM permission issues for the Billing and Cost Management console?
Amazon Web Services
Integrity at Scale: Inside the Flo Health Mission | Amazon Web Services
Amazon Web Services
Fueling Success: Small shifts, powerful performance | Amazon Web Services
Amazon Web Services
WEX enhances customer experience with AI-powered chatbot | Amazon Web Services
Amazon Web Services
Accelerate troubleshooting with Amazon CloudWatch investigations | Amazon Web Services
Amazon Web Services
Why is my Windows WorkSpace stuck in the starting, rebooting, or stopping status?
Amazon Web Services
Telemetry Pipelines for AI | Amazon Web Services
Amazon Web Services
Getting Control over Security and Observability Data | Amazon Web Services
Amazon Web Services
The Problem with Telemetry Data Volume | Amazon Web Services
Amazon Web Services
Telemetry Pipelines on AWS | Amazon Web Services
Amazon Web Services
What are Telemetry Pipelines? | Amazon Web Services
Amazon Web Services
Using AI for RegEx on Telemetry Pipelines | Amazon Web Services
Amazon Web Services
Multi-Session Support in the AWS Console | Amazon Web Services
Amazon Web Services
How CloudHedge delivers assessment with AWS ISV Tooling Program at no cost?
Amazon Web Services
How customers speed up migration and modernization to AWS with CloudHedge | Amazon Web Services
Amazon Web Services
Chaos Experiment with Amazon ElastiCache | Amazon Web Services
Amazon Web Services
Amazon S3 Access Points: Easily manage access for shared datasets on S3 | Amazon Web Services
Amazon Web Services
ElastiCache Valkey 8.0 - Savings and Efficiency | Amazon Web Services
Amazon Web Services
Pennymac scales document processing with AWS | Amazon Web Services
Amazon Web Services
AWS | Next Level Innovation | Amazon Web Services
Amazon Web Services
Driving Cloud Innovation: Mindtickle's Partnership with AWS Enterprise Support | Amazon Web Services
Amazon Web Services
A Leader's Edge from Executive Insights | Amazon Web Services
Amazon Web Services
How do I create a custom Amazon WorkSpaces image?
Amazon Web Services
Charles Leclerc tests his AI-generated race track | Amazon Web Services
Amazon Web Services
Redington Scales India’s Cloud Access with AWS Partnership | Amazon Web Services
Amazon Web Services
How do I prevent the resources in my CloudFormation stack from getting deleted or updated?
Amazon Web Services
How do I troubleshoot authentication errors when I use RDP to connect to an EC2 Windows instance?
Amazon Web Services
Exploring the Possibilities of Digital Twin & AI at the Edge | Amazon Web Services
Amazon Web Services
Exploring the Possibilities of Digital Twin & AI at the Edge | Amazon Web Services
Amazon Web Services
AWS at the FORMULA 1 AWS GRAN PREMIO DELL'EMILIA-ROMAGNA 2025 | Amazon Web Services
Amazon Web Services
What's new in RCPs | Amazon Web Services
Amazon Web Services
API Caching using Amazon ElastiCache | Amazon Web Services
Amazon Web Services
Pendula: Amazon Nova Customer Testimonial | Amazon Web Services
Amazon Web Services
InDebted : Amazon Nova Customer Testimonial | Amazon Web Services
Amazon Web Services
Amazon DynamoDB global tables with multi-Region strong consistency | Amazon Web Services
Amazon Web Services
Siemens Mobility uses AWS to operate securely, efficiently on a global scale | Amazon Web Services
Amazon Web Services
How do I reuse a knowledge base session in Amazon Bedrock?
Amazon Web Services
EP5: MBZUAI, CMU : Causal AI, Answering The “Why“ and “What if“ Questions | AWS for AI Podcast
Amazon Web Services
Hema scales time to market developing a data mesh on AWS (Technical) - Cloud Adventures
Amazon Web Services
Hema scales time to market developing a data mesh on AWS (Business) - Cloud Adventures
Amazon Web Services
How Langfuse Scaled Their AI Platform with AWS: From Open-Source to Enterprise | Amazon Web Services
Amazon Web Services
SLMs and LLMs: What’s the Difference? | Amazon Web Services
Amazon Web Services
SLMs and LLMs: When to use them? | Amazon Web Services
Amazon Web Services
SLMs on CPU | Amazon Web Services
Amazon Web Services
Intelligent Model Routing | Amazon Web Services
Amazon Web Services
SLMs, LLMs, and Model Routing in Agents | Amazon Web Services
Amazon Web Services
More on: AI Systems Design
View skill →Related Reads
📰
📰
📰
📰
Your HIPAA Posture, in Version Control
Medium · DevOps
hermes-memory-installer: Avoiding Stale Commit Hashes in Consistency Notes
Dev.to AI
Every AWS project starts with copy-pasting last repo's Terraform. I built a generator instead.
Dev.to · Framz
Kubernetes Health Probes: Liveness, Readiness, and Startup Explained
Dev.to · toothbrush
🎓
Tutor Explanation
DeepCamp AI