Hyperscale RTFM bug: A Google Cloud engineer dropped part of the cloud and immediately disconnected all fiber optic cables

Hyperscale RTFM bug: A Google Cloud engineer dropped part of the cloud and immediately disconnected all fiber optic cables

Google Cloud revealed a rather strange reason for a partial outage in the us-central1-b region, which occurred between 07:41 and 11:52 Pacific Time (PT) on September 1. Engineers repairing the equipment physically disconnected part of the cloud from the network. Report Register. This resulted in “severe network degradation and resource isolation.”








According to Google, at the height of the incident, traffic through the affected areas of the cloud was “reduced by 100%,” which is why many instances were not available to users at all and packet loss was expected to increase. The company explained that it built in router redundancy when designing its cloud data centers, but that there’s nothing against human “intelligence.” The system is designed for double or even triple failures, using separate equipment power supplies and multiple fiber optic lines. However, it turns out that all of these protections can be bypassed simply by unplugging the cable from the outlet.

It is reported that the direct technical cause of this failure was that the optical cable was disconnected during routine equipment maintenance work. The company reported that engineers consistently shut down 100% of fiber connections in just 13 minutes. The nature of the error and the high working speed of the “active” employee did not allow us to warn him of the erroneous operation in time before complete shutdown.

    Photo credit: Jimmy Nilsson Masth/unsplash.com

Photo credit: Jimmy Nilsson Masth/unsplash.com

After the router is shut down, the virtual machines in an availability zone in the us-central1-b region physically lose contact with the outside world. Although their ability to share data with each other is preserved, users cannot access it or even understand what is going on. After Google discovered the problem, traffic was redirected from the router, which engineers shut down, to “working capabilities elsewhere in the region.” Engineers then put the cables back and restored the network to its original configuration. The engineer who caused the fault was found to have breached the basic requirement to read instructions before starting work.

Human factors have repeatedly caused failures in the United States and abroad. So, in the summer of 2025, the Australian military mistakenly “pass” Wi-Fi and radio on the New Zealand coast. In December 2025, it was reported that the laziness and carelessness of Australian Nokia and Optus engineers resulted in to death People who can’t catch an ambulance. As the Uptime Institute points out, data center failures have become Less frequent, but more significant.

If you find an error, select it with your mouse and press CTRL+ENTER. |Can you write better? We always welcome new authors.

source:

Exit mobile version