Stop Showing Me Uptime Percentages—Tell Me How Many Hours You Were Down

Can we stop with the uptime percentages?

Stop Showing Me Uptime Percentages—Tell Me How Many Hours You Were Down

Jim Nielsen argues that uptime percentages are a terrible public interface for status pages. While infrastructure engineers intuitively grasp that 99.99% is ten times better than 99.9%, most users see both as an A grade. He proposes replacing 'GitHub Actions: 98.31% uptime' with '12 hours affected in the last 30 days (98.31% uptime)'—a metric anyone can understand.

The difference between a 6.2 and a 7.8 might not seem that big, but it represents a massive difference in magnitude. Uptime percentages have a similar problem: 99.9% and 99.99% look pretty much the same, but the latter is 10× less!
  1. proxysna

    Wider audience started to use status pages because the service unreliability became so much more noticeable than before and not the other way around. I never had to use a status page for bear blog or protonmail because i never had and issue with it or just i never noticed.

    I am now _required_ to consult status page of github, circleci or MS services etc because i need to know why a build is not passing, why i cannot open a repo, why is my work stalling.

    Percentages matter, it is just so much more obvious why they matter when it comes down to important pieces of the internet like github. And i highly doubt the number of 12 hours in the last month. MS has been downplaying the issues they have with GH performance for a while now and i don't think it is time to start to believe them yet. Maintaining these pieces of infrastructure is responsibility and a burden.

    Overall i would be careful with "nonlinear significance of numbers near 100%" we are talking gh being well into the 90's this year and one number that infra people are also often being reminded about is that "1% is 3.5 days".

    Things are tough for gh people and i feel for them but they are not a startup or a underdog of some sort to receive sympathy in that case.

  2. hdgvhicv

    I’ve worked in systems which at worse have had three nines for years, but I’ve also worked in systems where five nines is a failure.

    This attitude of modern tech claiming 98% is good just doesn’t work in the old tech acceptance. We had individual components fail all the time. We’re still looking at a 230ms outage to a branch office last week caused by a power failure combined with a badly plumbed power distribution.

    Modern software people don’t consider 230ms to be an outage. Glad they don’t work in electricity.

    (The failure we had was only on the services we guarentee at 99.1%, our lowest sla. After that there’s 99.95 and 99.999.

    (In reality we reach five nines year after year on even the lowest levels, but there are major concerns like “large bomb in data centre” which could cause some of our less critical units to drop way more than 5 minutes a year.

  3. teraflop

    Separately from how you present the number, the very concept of "uptime" as a single number is a bit muddy in the context of a distributed system, where different components can be differently available for different users.

    Also, 0.1% downtime in the form of a 45-minute outage per month is very different from 0.1% of requests failing in brief bursts. You often see downtime reported as "increased error rates" which is so vague as to be meaningless.

    Google's "windowed user-uptime" attempts to deal with this a bit better, by exposing different views of the data instead of trying to condense uptime into a single number: https://www.usenix.org/system/files/nsdi20-paper-hauer.pdf

  4. hx8

    > We say something like:

    > GitHub Actions: 12 hours affected in the last 30 days (98.31% uptime).

    This is trying to shine the most favorable possible light onto a deteriorating situation. It doesn't take away from the fact that most businesses have measurable missed revenue in downtime. Customers that shop somewhere else, ads that were never severed, leads that grew a little colder. 12 hours of downed GitHub results in millions of dollars of lost developer productivity that was externalized by Microsoft to other companies.

    We shouldn't be trying to spin downtime as "just a few hours a month." Those hours cost real dollars.

  5. runjake

    I don't really care about percentages, either. But for some industries, the difference between "two 9s" and "five 9s" can be millions of dollars, so that's why they're published that way to the customer.

  6. PaulKeeble

    One thing that is I feel missed about uptime percentage when compared to on premise uptime is when the downtime occurs. Its far more impactful if its in the middle of the working day or during the busy period of shopping. A store that goes offline in the middle of black friday or in the run up to Christmas is harmed a lot more than some down time on a Sunday night/Monday morning at 3am.

    One thing I have noted over time is a lot of these AWS, Azure et el downtimes is they occur in the middle of everyones day, millions of people are impacted by them. Same with github its getting in the way of work. Whereas when we hosted services on our own equipment the downtime was usually out of main hours. The percentages are in many ways the wrong measure of downtime because hours aren't equal in impact to businesses.

  7. procflora

    These service status pages have so many other problems, I can't really get excited about this article's point. As a user of a service, I don't really care that an additional 0.09% of uptime is any more or less difficult to achieve than an additional 0.9%, even if you describe it in terms of fractions of time. I only care about two things: what the service status is right now and your service reliability's impact to me over the long term (get out of here with your 30-day crap).

    In the electric utility world we have a few IEEE standardized metrics (with appropriately IEEE'd acronyms) for tracking service reliability that I like much better and always wish for when I'm looking at a status page. Pie in the sky stuff for sure, nobody wants to do this analysis and publish the results without a regulator telling them have to, but c'est la vie.

    SAIDI - System Average Interruption Duration Index. How many minutes an average customer experienced service interruption in a year. This is the big one I'd want to see on your service status page IMHO. For the power grid, we consider any outage longer than five minutes to be an interruption ("non-momentary outage").

    SAIFI - System Average Interruption Frequency Index. How many total periods of interruption occurred for the average customer in a year.

    CAIDI - Customer Average Interruption Duration Index. How long it takes service to be restored for the average customer when there is an interruption.

    For the US, here is what these numbers look […]

  8. jakevoytko

    These numbers are useful proxies for how likely you are to have your work disrupted outside of your own control.

    If you do something 100 times a day against a four-nines service, you can reasonably expect that everything will succeed.

    If you do something 10,000 times a day against a two-nines service, you can expect to hit a substantial number of errors during that day, or even have long periods where your work cannot happen at all.

    People aren't frustrated with Github because Github has 98% uptime or whatever the specific number is. They're frustrated because it regularly interferes with their ability to work. The 98% number is just a concise way to say it.

More from this day

2026-09-16