A Server Lost Power at 00:32. We Found Out at 08:18

A storage server at Danube Data lost power at 00:32 UTC on 17 August 2026, silently taking down object storage, the container registry, and serverless deployments for over eight hours. The status page detected the incident within three minutes, but no one saw it until an engineer happened to check at 08:18. The postmortem reveals why one machine caused a total outage: erasure-coded chunks were spread across disks without ensuring they were on different machines, and a few kilobytes of bookkeeping records were stored with less redundancy than the hundreds of terabytes of customer data they described. The team has since added a phone-call pager and a watchdog that can power-cycle a dead server, but the storage layout and capacity headroom remain unfixed.

Monitoring that nobody is woken by is a record, not an alarm.
  1. danieka

    The incident is somewhat embarrassing for what appears to be a hosting provider. And the article is clearly written with a LLM. From a PR perspective that’s a peculiar combination. A bit more empathy and a human touch would probably have been a good move. Accidents and outages do happen, but reading the LLM article makes me feel like no one cares about what happened.

  2. walrus01

    Very obviously LLM written.

    As to the extremely lengthy content of the page itself, the absence of the server should have been noticed by even the most rudimentary librenms or opennms type monitoring and alerting system. Alert went to one person by email? Uhhhh....

  3. jefftk

    Clearly LLM written, and describing a design intended for a minimum of 6 servers that was running on 4. They don't say either way, but it seems possible to me that they might happen to write 3 chunks to a single machine and lose data on a one-machine failure

More from this day

2026-08-24