'We let you down': GitHub pledges to scale up before developers give up
CTO promises architectural overhaul following second outage of the month
DEVOPS
'We let you down': GitHub pledges to scale up before developers give up
CTO promises architectural overhaul following second outage of the month
GitHub's handwringing continued this week as CTO Vladimir Fedorov offered more detail about the August 17 outage – while carefully avoiding the word "sorry."
The outage lasted 7 hours and 47 minutes and disrupted developers worldwide. Actions, pull requests, issues, Copilot, and APIs were among the services affected as the platform failed to scale with demand. "If you were trying to ship software that day, we let you down," wrote Fedorov.
The incident followed another outage involving Actions on August 6. The platform has been wobbly for some time, something it acknowledged in April, but work to address the underlying issues has not kept pace with the relentless rise in traffic.
In April, monthly commits were at 1.4 billion. GitHub says it now handles 2.9 billion commits, 24 million new repositories, and 130 million merged pull requests each month.
Microsoft Azure currently handles approximately 58 percent of GitHub's platform load and half of all Git operations. According to Fedorov, GitHub has accelerated the migration of more workloads to its parent company's cloud.
"Our next milestone is an architecture that scales read capacity linearly with the number of readers, enabling unlimited read operations," he stated. "We will roll it out gradually, beginning with the largest monorepos."
Before that architecture arrives, GitHub must address the scaling weaknesses exposed by retry storms and misconfigured limits. Fedorov stressed that "neither outage was caused by a code or configuration change" – in other words, the failure modes were already lurking in the platform rather than introduced by a fresh deployment.
The company is also working to isolate critical systems to reduce the blast radius of future failures, tighten retry limits, and add alerts for early signs of traffic spikes.
GitHub's repeated outages have rattled at least some developers. Responses on social media mixed sympathy for the challenge of operating at such scale with frustration from paying customers who say the service is falling short.
Fedorov concluded: "The developer community depends on GitHub to build, ship, and operate their work. That is only possible if you can rely on us, and on August 17, you couldn't. It is our responsibility to fix that. We'll earn your trust through the scaling and reliability of the platform." ®
Originally published on The Register
