Somewhere between five and twenty percent of a typical estate consists of virtual machines that nobody can account for. They have been running for years. They consume licences, capacity, backup windows, and patching effort, and during a platform migration they consume migration effort too.
Nobody switches them off, because the person who would be blamed if it mattered is the person considering switching it off, and the incentive is entirely one-directional.
Why the usual approach fails
Send a list to application owners asking them to confirm what is still needed. The response rate is poor, and the responses that come back are mostly “keep it, just in case” because that answer carries no risk for the responder.
You end up having done work and eliminated nothing.
The process that works
Gather evidence first. Before asking anyone, find out what the machine actually does. Network connections in and out over thirty days. Logon history. CPU and disk activity. Whether anything monitors it. Whether it appears in any backup restore request, ever.
A machine with no inbound connections, no interactive logons in ninety days, and flat resource usage is almost certainly doing nothing. That evidence changes the conversation from “do you need this” to “here is what we observe, tell us if we are wrong”.
Then a notice period with a default. Announce that machines meeting the criteria will be powered off on a specific date unless somebody claims them. Publish the list. The default is action, not inaction, which flips the incentive.
Power off, do not delete. Leave it powered off for a defined period — sixty or ninety days is typical. This is the entire safety mechanism. If something breaks, power it back on and the incident lasts minutes.
Almost nothing gets reclaimed in practice, and the ones that do are memorable and cheap.
Snapshot or back up before deletion, and keep that copy for longer than the notice period. Then delete.
Document each decision. Machine name, evidence, notice date, power-off date, deletion date, who approved. When somebody asks in eighteen months, you have an answer. This record is what makes the process repeatable, because the next round faces less resistance.
The categories you will find
Genuinely dead. Projects that ended, migrations that completed, test environments from a product evaluation in 2019. The majority.
Dormant but real. Runs once a quarter or once a year. Payroll, regulatory reporting, an annual process. These are dangerous to switch off on a ninety-day activity window, which is why the observation period should span a quarter and why the notice must be published widely.
Alive but silent. Something that listens on a port nobody uses until it is needed. Evidence gathering catches these because the service is running even if traffic is rare.
Someone’s undocumented production. Rare, and the reason the power-off stage exists.
Doing this before a migration
This is the highest-return work available before a platform migration, and it is almost always skipped because it is unglamorous and produces no visible progress.
Eliminating fifteen percent of the estate reduces the licence count, the migration effort, the target platform size, and the duration of the parallel-running period. It is cheaper than every other optimisation available and it requires no technology at all.
Do it first. The migration is a rare moment when the organisation will actually agree to it, because there is a deadline and a cost attached to carrying dead weight across.