{"id":21193,"date":"2025-11-20T09:05:00","date_gmt":"2025-11-20T17:05:00","guid":{"rendered":"https:\/\/www.microsoft.com\/insidetrack\/blog\/?p=21193"},"modified":"2026-08-14T15:41:30","modified_gmt":"2026-08-14T22:41:30","slug":"moving-from-a-scream-test-to-holistic-lifecycle-management-how-we-manage-our-azure-services-at-microsoft","status":"publish","type":"post","link":"https:\/\/www.microsoft.com\/insidetrack\/blog\/moving-from-a-scream-test-to-holistic-lifecycle-management-how-we-manage-our-azure-services-at-microsoft\/","title":{"rendered":"Moving from a \u2018Scream Test\u2019 to holistic lifecycle management: How we manage our Azure services at Microsoft"},"content":{"rendered":"\n

Nearly a decade ago, as we began our journey from relying on on-premises physical computing infrastructure to being a cloud-first organization, our engineers came up with a simple but effective technique to see if a relatively inactive server was really needed.<\/p>\n\n\n\n

They dubbed it the \u201cScream Test.\u201d<\/p>\n\n\n\n

\u201cWe didn\u2019t have a great server inventory and tracking system, and we didn\u2019t always know who owned a server,\u201d says Brent Burtness, a principal software engineer in Commerce Financial Platforms, who was one of the leaders for the effort in his group. \u201cSo, we essentially just turned them off. If someone screamed\u2014\u2018Hey, why\u2019d you turn off my server?\u2019\u2014then we\u2019d know it was still being used.\u201d<\/p>\n\n\n\n

Today, the basic idea behind the Scream Test is being used across the company, but in a more holistic way. Importantly, it\u2019s been incorporated into the overall lifecycle management of our computing infrastructure. And, through the automation tools provided by Microsoft Azure, we have a much more efficient process for making sure that we\u2019re saving time and money by reducing the number of underused machines we operate, monitor, and maintain.<\/p>\n\n\n\n

\"A<\/figure>\n\n\n\n
\n

\u201cWe thought we were going to get rid of a small number of machines that weren\u2019t being used. But we found the actual share was about 15% of all machines, which saved us a lot of effort of moving those unused machines to the cloud. In other words, we downsized on the way to the cloud, rather than after the fact.\u201d<\/p>\nPete Apple, cloud network engineering architect, Microsoft Digital<\/strong><\/cite><\/blockquote>\n\n\n\n

Uncovering more than expected<\/h2>\n\n\n\n

The Scream Test was part of the huge effort to evaluate our on-premises compute resources before we began moving to the Azure cloud. After all, why spend resources moving something that isn\u2019t needed?<\/p>\n\n\n\n

Pete Apple, who helped develop the concept of the Scream Test, is a cloud network engineering architect in Microsoft Digital, the company\u2019s IT organization. Looking back, he remembers the surprising results that emerged when they began shutting down specific servers to see who noticed.<\/p>\n\n\n\n

\u201cWe thought we were going to get rid of a small number of machines that weren\u2019t being used,\u201d Apple says. \u201cBut we found the actual share was about 15% of all machines, which saved us a lot of effort of moving those unused machines to the cloud. In other words, we downsized on the way to the cloud, rather than after the fact.\u201d<\/p>\n\n\n\n

As part of this process, Apple explains, our engineers looked at two related factors to reduce inefficiencies in our usage of computing resources.<\/p>\n\n\n\n

The first was to identify systems that were used infrequently, at a very low level of CPU (sometimes called \u201ccold\u201d servers). From that, we could determine which systems in our on-premises environments were oversized\u2014meaning someone had purchased physical machines according to what they thought the load would be, but either that estimate was incorrect or the load diminished over time. We took this data and created a set of recommended Microsoft Azure Virtual Machine (VM) sizes for every on-premises system to be migrated. <\/p>\n\n\n\n

\u201cWe learned that there’s a lot of orphaned, or underutilized, resources out there,\u201d Burtness says. \u201cThese were cases where the workload was so small on a server\u2014like under 5% CPU\u2014that it didn’t make sense to host it on its own machine. We could then move the task or application and get it down to just one or two CPUs on a virtual machine.”<\/p>\n\n\n\n

At the time, we did much of this work manually, because we were early adopters. The company now has a number of products available to assist with this review of your on-premises environment, led by Azure Migrate<\/a>.<\/strong><\/p>\n\n\n\n

Another part of the process was determining which systems were being used for only a few days a month or at certain busy times of the year. These development machines, test\/QA machines, and user acceptance testing machines (reserved for final verification before moving code to production) were running continuously in the datacenter but were really only needed during limited windows. For these situations, we applied the tools available in Azure Resource Manager Templates<\/a> and Azure Automation<\/a> to ensure the machines would only run when needed.<\/p>\n\n\n\n

Automating with Azure<\/h2>\n\n\n\n

Today, we don\u2019t have to rely on anything as crude as the Scream Test to find unused and underused computing resources. With 98% of our IT resources operating in the Azure cloud, we have much greater insight into how efficient our network is, so much of the process can be automated.<\/p>\n\n\n\n

\u201cWe\u2019ve found this effort much easier to manage in the cloud, because all our computing resources are integrated with the Azure portal<\/a>,\u201d Apple says. \u201cThey have an API system and offer various tools within Azure Update Manager<\/a> and Azure Advisor<\/a> to help with cost efficiency. It’s kind of like a modern version of Clippy\u2014\u2019Hey, it looks like your VM isn’t being used much. Do you want to downsize that or turn it off?'”<\/p>\n\n\n\n

(For the uninitiated, Clippy was the Microsoft Office animated paperclip assistant introduced in the late 1990s. It offered tips and help with tasks, like writing and formatting documents. Clippy became iconic for its quirky suggestions, including recommending that you remove things from your desktop that you weren\u2019t using.)<\/p>\n\n\n\n

\"Burtness<\/figure>\n\n\n\n
\n

“With everything being in the Azure portal or in Azure Resource Graph, it’s much more streamlined, and makes it easier to get that data out to the teams. They can then go into the portal and clean up the resource.”<\/p>\nBrent Burtness, principal software engineer, Commerce Financial Platforms<\/strong><\/cite><\/blockquote>\n\n\n\n

And simply taking the step of turning off stuff that we weren\u2019t using turned out to be very effective. Thanks, Clippy! <\/p>\n\n\n\n

Today, we approach this challenge in a more efficient and sophisticated way, taking advantage of Azure tools like Update Manager and Advisor.<\/p>\n\n\n\n

“With everything being in the Azure portal or in Azure Resource Graph<\/a>, it’s much more streamlined, and makes it easier to get that data out to the teams,” Burtness says. “We can run automated queries with Azure Resource Graph. Then we bring that information into our internal Service 360 tool, which we use to give action items to our developers. Each item gives them a link to Azure portal, and they can then go into the portal and clean up the resource.”<\/p>\n\n\n\n

Managing for the lifecycle<\/h2>\n\n\n\n

One of the most important things we learned by using the Scream Test to identify inefficiencies and moving our systems from on-premises servers to the cloud was that it\u2019s an ongoing process, not a fixed-end project. <\/p>\n\n\n\n

\u201cWe had this idea that it was going to be a one-time event, that we’ll move to the cloud and then we’ll be done,” Apple says. “A better understanding is that it’s a lifecycle. We have integrated this concept of continual evaluation into our processes around everything that’s still on-premises, because we still have labs, we still have physical infrastructure.\u201d<\/p>\n\n\n\n

We continue to do this evaluation on a regular basis with both physical and virtual computing resources, because needs and usage are constantly changing.<\/p>\n\n\n\n

Cutting our cloud costs<\/h2>\n\n\n\n
\"A
In a pilot set of Azure subscriptions, the Commerce Financial Platforms team reduced usage by 233 resources across 36 subscriptions and 17 services in 6 team groups, saving more than $15,000 in monthly operating costs.<\/figcaption><\/figure>\n\n\n\n

“Now we have a basic process around a six-month cycle,\u201d Apple says. \u201cSo, every six months we ask, does this still need to be on-premises or should we start moving it to the cloud? And we do the same thing with our cloud resources. Who’s still using these VMs? And we still go through the same review process to see if it\u2019s needed, or if we can shut it down or move it.”<\/p>\n\n\n\n

This has resulted in significant cost savings for the company. \u201cWe\u2019re up to about 15% to 20% less compute cost, depending on the organization, because of this much better understanding of our business needs,\u201d Apple says.<\/p>\n\n\n\n

Better governance, increased security<\/h2>\n\n\n\n

Another major benefit of this process was establishing much stronger governance of compute resources across the entire organization. <\/p>\n\n\n\n

\u201cWhen we first did the Scream Test, we weren’t always really sure who owned what, in some cases,\u201d Apple says. \u201cWe\u2019ve fixed that as part of this process. This governance aspect is a key part of being more efficient with our resources.\u201d<\/p>\n\n\n\n

Burtness explains why this is so important.<\/p>\n\n\n\n

\u201cIt\u2019s critical to know exactly who to contact when there’s something wrong with the server,\u201d Burtness says. \u201cNow, with clearer ownership, clearer accountability, and better inventory, it\u2019s a much better experience.\u201d<\/p>\n\n\n\n

Better governance also means tighter security, according to both Apple and Burtness. <\/p>\n\n\n\n

\u201cThis is really important when it comes to threat-actor response,\u201d Apple says. \u201cUnused servers can often be an entry point for hackers. Or, say we discover that a machine or server is getting hacked; you need to talk to who owns it. If you don\u2019t know, it takes you longer to track them down and combat the hack. That’s not great. Improving our governance has definitely made securing our environment easier.\u201d<\/p>\n\n\n\n

\n
\n
\"\"<\/figure>\n\n\n\n

Key takeaways<\/p>\n<\/div>\n\n\n\n

Here are some things to keep in mind when managing your own enterprise compute resources for greater efficiency:<\/p>\n\n\n\n