GitHub restores services after nearly 8-hour outage disrupts Actions, APIs, PRs and Copilot
GitHub has restored services after a nearly eight-hour outage disrupted several of its core developer tools, including Actions, pull requests, APIs, Git operations, Webhooks, and Copilot, impacting software development workflows across its platform.
“This incident has been resolved. Thank you for your patience and understanding as we addressed this issue,” the company wrote on its status page.
The disruption was first reported at 1:40 PM UTC on August 17, with GitHub initially flagging degraded performance across parts of its platform. Within minutes, the disruption had spread to API Requests, Actions, Webhooks, Issues, and pull requests (PRs).
At the height of the incident, GitHub reported an error rate of about 20% across its web experience and API traffic. Archive downloads and raw repository content downloads were seeing an error rate of approximately 50%, while SAML and OIDC authentication, SCIM, and Team Sync were also affected.
GitHub’s AI coding assistant, Copilot, also experienced degraded availability beginning at 2:31 PM UTC. The company later said that some Copilot authentication problems persisted even after other parts of the platform had recovered.
Recovery was not linear
Finally, at 4:36 PM UTC, the company said that it had identified the problematic component and had taken corrective actions.
However, the recovery of services was not linear.
Git Operations experienced another period of degraded performance, while API Requests briefly returned to a degraded state. GitHub subsequently said it had mitigated the Git Operations issue and restored normal API operation by 7:01 PM UTC.
Still, other services remained affected with problems centered largely on authentication. GitHub said it had partially disabled authentication-token retries after observing sporadic authentication failures and later reported that Copilot authentication issues were still affecting some applications.
“We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation,” the company wrote.
It finally fixed these issues at 8:45 PM UTC and marked the incident resolved at 9:15 PM UTC, nearly 8 hours after it was first reported.
Downtime significant for enterprise teams
While the incident didn’t amount to a complete shutdown of GitHub, the number and nature of services impacted are significant for enterprise development teams as the company increasingly serves as more than a repository for source code.
Development teams use its APIs, Actions workflows, PRs, and integrations as connected parts of their software delivery processes, and a disruption to several of those components simultaneously can therefore affect workflows even when basic repository access remains available.
GitHub has yet to identify the “problematic component” that caused the failure or why its failure took out so many services together.
Until the company’s post-incident analysis is available, it is unclear whether the outage was caused by an infrastructure failure, a software change, an authentication problem, or another issue like an attack.
For now, GitHub’s status page says that a “detailed root cause analysis will be shared as soon as it is available.”
The article originally appeared on InfoWorld.ComputerworldRead More