Platform

Cutting an MSP's cloud bill without cutting corners

Cloud waste is the quietest margin leak in the business, on your own estate and on every client's. Here is where the money actually goes, and five checks that claw it back without touching reliability.

By Nathan Carroll21 September 20268 min read

The first time I properly audited our own cloud spend, I found that nearly a third of it was doing nothing. Not underused. Nothing. Machines nobody had turned off, disks attached to servers that no longer existed, snapshots from a migration two years earlier, all of it billing quietly, every hour, while we worried about winning the next client.

That is the thing about cloud cost. It does not announce itself. There is no line on the invoice that says "waste." It hides inside a bill that feels roughly right because it grew slowly, a little each month, until "roughly right" was thousands of pounds a year of margin walking out of the door.

For an MSP this cuts two ways. It is your own margin on your own platform, and it is a service your clients need badly and rarely get. Most providers who resell cloud mark up the whole bill, waste included, and call it a day. That works right up until the client's finance director looks closely, and then it stops being a revenue line and becomes a trust problem.

Here is where the money actually leaks, and the five checks that recover it without touching reliability.

A word before you start cutting

Cost optimisation gets a bad name because it is confused with cutting corners: turn things off, downgrade everything, hope nothing breaks. That is not this. Every check below either removes something that is genuinely doing no work, or right-sizes something that was over-provisioned in the first place. Done properly, the estate ends up cheaper and exactly as reliable, often more so, because you finally understand what is actually running.

1.Idle and orphaned resources

This is the biggest and most embarrassing one, because it is pure waste. Virtual machines left running after a project ended. Disks no longer attached to anything but still provisioned and still billed. Old snapshots and images kept "just in case" for years. Public IP addresses reserved and forgotten. Development and test environments running twenty-four hours a day to be used for six.

None of this touches reliability, because none of it is doing anything. Find it with the provider's own cost and advisor tooling, confirm it is genuinely orphaned, and remove it. Put dev and test environments on a schedule so they sleep outside working hours. This one check often recovers ten to twenty per cent of a neglected estate in an afternoon.

2.Over-provisioned by default

Cloud makes it trivial to ask for more than you need, and almost nobody goes back to check. Virtual machines two sizes larger than their workload ever uses. Premium solid-state disk under something that would run happily on standard. Databases provisioned for a peak that happens twice a year and sit idle the rest of the time. Over-provisioning feels safe, and it is quietly expensive.

Right-sizing is not guesswork. The utilisation data is already there, sitting in the console, waiting for someone to read it.

Match the resource to the actual demand across a few weeks of CPU, memory and disk figures. Keep genuine headroom where load is spiky, and use autoscaling for the workloads that need it rather than paying for peak capacity around the clock. You are not making anything slower, you are stopping yourself paying for a ceiling nothing ever reaches.

3.Egress and data transfer you never see

This is the leak that catches people out, because it stays close to invisible until you go looking. Data leaving a region, moving between availability zones, flowing through a NAT gateway, replicating for backup, chattering between services that were placed without much thought. Every gigabyte has a price, and almost none of it appears as a line you would recognise. It just inflates the total.

You cannot eliminate egress, but you can stop paying for the accidental kind. Look at where data actually moves, keep services that talk to each other in the same place, be deliberate about cross-region replication, and check what your backup and monitoring tools are quietly shipping out. The savings are less dramatic than a room full of idle machines, but they are permanent.

4.Licensing left on the meter

This one is close to pure margin, because you are often paying full rate for something you could have at a discount, or paying twice for something you already own. Steady, predictable workloads left on pay-as-you-go pricing when a reservation or savings plan would cut them by a third or more. Windows and SQL Server licences you already hold, never applied through Hybrid Benefit. Deprovisioned users still holding paid seats. Trial add-ons that quietly became paid ones.

Commit reservations or savings plans to anything with a stable baseline, apply the licensing benefits you are entitled to, and reconcile seats against real users on a schedule. This is not a technical change at all, it is a billing one, which means it recovers margin with zero risk to anything running.

5.Storage and backup bloat

Storage is cheap per gigabyte, which is exactly why it is allowed to sprawl. Hot, expensive tiers holding data nobody has opened in a year. No lifecycle rules to move cold data down on its own. Backups and snapshots kept far longer than any policy requires, multiplying every month. Log data retained forever because nobody ever set a limit.

Set lifecycle policies so data ages into cheaper tiers automatically. Match backup and snapshot retention to what recovery and compliance genuinely require, not an unexamined "keep everything." Cold data on cold storage costs a fraction of the hot tier, and it is still there the moment you need it.

What this is actually worth

Put a number on it, because "cloud waste" is easy to nod at and just as easy to wave away.

The maths

Take a cloud estate billing £10k a month, your own or a client's. Across the industry, roughly a quarter to a third of cloud spend is waste, so call it 30%: £3k a month, £36k a year, buying nothing.

You will not reclaim every penny, but half of that is realistic without touching a single production workload, most of it from idle resources, right-sizing and reservations alone.

That is £1,500 a month recovered, on one estate, with zero impact on reliability.

On your own platform that is margin straight to the bottom line. On a client's, it is the most credible trust-builder you have, and the natural opening for a recurring optimisation service.

The advisor's advantage

Here is the part most MSPs get backwards. If you resell cloud, cutting a client's bill feels like cutting your own revenue, so the temptation is to leave the waste where it is and keep the markup. Resist it. The provider who proactively saves a client thousands becomes the trusted advisor they never leave, and the one they call first for the next project. The provider marking up waste becomes a line the finance team eventually questions. One of those relationships compounds. The other has an expiry date.

Package the work: an initial optimisation review, the remediation, then ongoing monitoring so the waste does not creep back. It is recurring revenue built on saving your client money, which is about the easiest thing there is to sell.

Where to start

Start with your own estate, this week, before you offer any of it to a client. Run the provider's cost tools, find the idle and orphaned resources, and turn off what is doing nothing. It is the fastest, safest money you will find all quarter, and it turns the whole exercise into something you have done rather than something you have read about.

Where is your cloud margin leaking?

  • Do you know how much of your cloud bill is idle or orphaned resources, right now?
  • When did you last right-size something, rather than only ever scale it up?
  • Can you see egress and data transfer as its own line, or is it buried in the total?
  • Are your steady workloads on reservations, and are you applying the licences you already own?
  • Does cold data age into cheaper storage automatically, or is everything still in the hot tier?

If those were hard to answer, the waste is there. It always is. The only question is whether you find it before the client's finance director does.

Nathan Carroll

Platform is one of the three levers I score

Want to know where your platform is leaking?

Book a free Scale Audit and I'll give you an operator-to-operator read on your Platform, Protection and Profit, with a one-page action plan you keep either way. Thirty minutes, no pitch.