GitHub Actions was down yet again
Another outage puts last week's promises to an early test
devops
GitHub Actions was down yet again
Another outage puts last week's promises to an early test
GitHub Actions stumbled again on Wednesday, days after the code host renewed its promises to improve reliability.
Wednesday’s problem, as has so often been the case, hit Actions, GitHub’s CI/CD platform for automating software builds, tests, and deployments. According to the incident report GitHub put out for the disruption, things started going south at 1511 UTC.
The source code host identified an issue with a database primary and failed over to a replica, but said the move "did not fully mitigate" the degradation. GitHub then throttled inbound traffic while investigating upstream Vitess issues before gradually restoring traffic. By 1800 UTC, it said Actions was operating as expected and inbound queues had recovered.
The latest disruption isn’t particularly reassuring given that GitHub claimed last week that it’s now serving double the commits it was dealing with in April, which wasn’t exactly a good month for GitHub either. No month this year has been great at the ‘Hub, really.
The history archive on GitHub’s status website indicates there were 26 issues with the platform in April. There were 23 in May and June, 26 in July, and there’ve been 23 so far in August with just under a week left to go. Whether this month can top March, with 32 incidents, or February’s 37, remains to be seen. There were 25 in January, too, meaning GitHub has suffered at least 23 reliability issues every month this year.
GitHub Actions is arguably a central part of the platform for many developers using CI/CD workflows and other forms of automation, and it has been among the services hardest hit by GitHub’s ongoing reliability problems.
As everyone who uses a software-as-a-service product knows, uptime is a key element in measuring reliability, and Actions isn’t exactly at triple nines right now - as of Wednesday, GitHub’s uptime page for Actions shows it at just 98.13 percent for August - nearly in danger of slipping into 97 percent reliability territory. That’s a bad place to be when you’re supposedly dealing with 2.9 billion commits, 24 million new repos, and 130 million merged pull requests a month, as GitHub claims it is.
While GitHub’s issues this year have been many and frequently reported on here at The Register, August has been a particularly bad month for the operation. August 17 saw GitHub suffer from a nearly eight-hour outage that hit multiple services, with Issues, Pull Requests, APIs, Actions, and Copilot all producing elevated errors and hamstringing customers’ ability to do work.
GitHub has pointed the finger at AI for many of its issues, blaming bots and agents for skyrocketing usage it hasn’t been able to cope with. The same went for that August 17 outage, with GitHub CTO Vladimir Fedorov issuing a mea culpa for the incident, saying that his operation had let users down and promising, just like he did back in April when GitHub admitted it was having issues, to scale enough to support its growing user base, human or otherwise.
“We'll earn your trust through the scaling and reliability of the platform,” Fedorov wrote in last week's postmortem of the August 17 outage. Six days after he promised to fix things, here we are with reliability slipping and the issue count growing.
GitHub didn’t respond to questions for this story. ®
Originally published on The Register
