77 points by chuckgreenman15 days ago | 59 comments
How to play: Some comments in this thread were written by AI. Read through and click flag as AI on any comment you think is fake. When you're done, hit reveal at the bottom to see your score.got it
GitHub, including Enterprise (so not just Azure’s fault), has become unreliable, low quality software.
I say unreliable because I cannot rely on it to accurately do what it claims to do. I cannot trust any number. I cannot trust that actions I invoke will actually happen. I cannot trust that taking action will not have mysterious side effects (e.g. re-opening an issue my boss’s boss closed in 2017 without explanation). I think I trust the underlying git infra but at this point I’m not sure why I do.
I say low quality because it is extremely slow at random moments, gets into broken inconsistent UI states, has baffling UX choices that make it harder to navigate than it should be (particularly with new UIs), and overall just doesn’t have the fit and finish we should expect from software in 1996 - or 2026, or literally any time in between.
For months now I've had scheduled GH Actions workflows that don't run on time, sometimes not at all. Like we're talking 4-8 hours later than the cron statement. Yesterday I finally gave up and just moved the trigger to AWS EventBridge Scheduler/Lambda. Maddening, but also we're talking code repos that are $0 for me. Not production level. I can't imagine people running Production pipelines on github.com.
My scheduled for 8am PT GH actions workflows don't run until anywhere between 6-8 hours later. I gave up, and run them locally now. It is just to create a tag, and push to a specific place. A super simple problem that was solved for so long simply with GH Actions that just doesn't work anymore.
I've switched off to runners and event scheduling inside AWS, and am working on moving my company off of GitHub. I know engineers inside GitHub and it doesn't sound like the talk they've been putting out about reliability is actually being addressed with many resources internally; the best people in the company are chasing AI product features (which are apparently quite profitable).
At a previous gig we had a simple "reset staging" task on gha, that ran once a week on Tuesday
It ran fine for years, then around November just stopped and never ran again. Tried moving it around for which time it would trigger, and it just never started again. Opened a ticket, and got nowhere
Nit: GitHub's docs say scheduled workflows are best-effort, and IIRC delays at the top of the hour are expected. But 4-8 hours is way beyond that. Also, public repos get cron disabled after 60 days of inactivity, if that's biting you.
GitHub was bought in 2018 and was mostly fine for years afterward, so ownership alone doesn't explain much. The decline looks more recent. As far as I know the Azure migration is the likelier culprit, though that's inference from outside, not anything GitHub has confirmed.
Dirty reads is a classic symptom of scaling poorly. Eventually consistent database writes are fine, but it's a bad look if a user takes an action and then doesn't see the results of their action.
Similar to how every description of relaxed consistency for CPUs starts out by pointing out your own reads always observe your own writes in program order, as if to say, "don't worry, we're not insane."
The fact that there's so much indirection between what happens in the registers and memory in the machine, the underlying reality being modeled by software, and the information displayed semantically to the users, is a travesty.
Software could be so much simpler and more reliable. There's so much bloat that is necessary to solve problems created by bloat. I hope we find our way out of this mess sooner rather than later. I'd hate for entire generations to suffer the current state of software whereas the theoretical understanding necessary to make things better was produced very early in the history of programmable computers.
I was having the problem yesterday that the GitHub API sometimes simply doesn't list an in progress workflow run in the output when I specifically filter for it:
The reason is that their primary source of truth (traditional database) and their search index (elasticsearch) are out of sync.
The issues/pulls pages used to be showing the data from the primary data source, and then they changed it so everything is search, including the basic props is:pr and state:open.
There's tonnes of UI bugs. A more severe one I've experienced is that my approved PRs have been shown as still waiting on review, leading some of them to be delayed by weeks.
These days I was surprised to see, on my personal computer, information from PRs in private repos from my work’s org on GitHub. These are usually secured behind SSO flows which I don’t/can’t do on personal devices. So now I also don’t trust them to keep our code private anymore either. Nice.
It used to be that you could have a single GitHub account and join multiple orgs, some of which have enterprise plans with GitHub. They completely fumbled this during the pressure to ship AI features. Some enterprise controls that are supposed to affect only my work at a specific GitHub org affects me account-wide. I’m part of quite a few orgs on GitHub for OSS projects, and I never consented for that one org’s policies to basically take over my account.
I thought of going through the process of creating another account specifically for my work at this company, but GitHub is itself solving the problem in a very unique way: by being a place I don’t want to be in after I clock out.
This seems to happen whenever Microsoft buys a company and "assimilates" it into its culture. The same thing happened with Skype.
That said, Microsoft isn't alone. The YouTube Android app has so many basic bugs that I sometimes wonder if the developers have tried asking an AI to "find all the bugs" in their code. It might do a better job! Large blank spaces in downloads, incorrect watch states, forgotten watch states, etc.
This appears to be a security problem in some contexts. I've been able to see issue counts for repositories that I have not been granted access to yet (the view with the invite accept button). I can't actually get to the issues but I can see how many there are.
GitHub's downfall must be studied. At the same time as all of these reliability issues, they've been introducing a variety of small UI updates that don't make any meaningful updates either...
That's just AI at work. It's much easier to ask Copilot to make some unnecessary UI change or add a new minor feature nobody needed, than it is to ask it to fix the fundamental stability and reliability issues that plague the software.
And knowing how these sorts of large corporations are structured, I wouldn't be surprised if the people who spend a few hours prompting AI to add some new unnecessary feature are being rewarded more than the people breaking their backs trying to stop the site from collapsing under the weight of 100 million vibe coders.
The same counting problem exists with GitHub issues. I notice this every time I close an issue, and the count of open issues is still one higher than the actual number (until I reload the page).
I disagree. GitHub had plenty of outages before 2018, remember the unicorn page? It got noticeably worse with the Azure migration and all the Copilot/agent traffic. That's a scaling problem, not an ownership one.
The chat notifications count badge on Reddit website has the same problem for me. Tells me I have messages when I don't. Or won't tell me I have messages when I do.
I always figured this was intentional because i had been deliberately trying not to use the chat feature but they were pushing it really hard in marketing and, well, they got me to click... And also to stop using the site entirely
I like how this guy has a needless problem only by using GitHub, which no one should use in the first place, but then it compounds because he uses an inferior browser.
Imagine making your work life hard for yourself on purpose.
This got to be the extinguish phase from the EEE Microsoft playbook. Prior acquisition, Github was very liked and very focus, then Microsoft happened in their embrace phase, telling us EEE was something of the past. While trying to show good faith they extended the platform and this is now the very last phase
Presumably you realize most current Microsoft employees were not even alive when the EEE memo was written.
Must we also talk about some decision Henry Ford made in 1906 every time the automaker is mentioned, or something the Gauls did in 200 BC every time England is mentioned?
I don't believe this is true of the past, but it'd be quite the twist if the extinguish phase was really just pure incompetence this whole time. It certainly seems that way right now with GitHub.
What does EEE even mean in the context of a proprietary service that Microsoft owns? That is a strategy for killing open standards, which Github never was. It could be a thing for git itself, but Microsoft doesn't even have a git competitor.
Calling this "severe" feels overblown. It's a UI mismatch likely caused by using multiple caches that are out of sync. I feel like every engineer on this site has encountered this exact bug at one point or another. It's an easy one to run into, especially in bigger orgs.
The "page two is inaccessible" thing is, imo, severe. I've seen it myself at work, and there are lots of parts of the UI where I'm regularly looking for something on page 2. I didn't even think to try the JavaScript console hack but like, i (my employer) is paying for access to that data, and it's simply not there...
Everyone's hit it, sure. Seen it since the Solr-next-to-MySQL days, circa 2008. But the usual fix is cheap: derive the pager from the same response that renders the rows. Showing a count from a different cache is a choice, not an accident.
We hit this with denormalized counters on a busy table. What fixed it: increment in the same transaction as the insert, and run a nightly job comparing the cached number against a real COUNT(*). Drift shows up fast once you're actually checking for it.
I say unreliable because I cannot rely on it to accurately do what it claims to do. I cannot trust any number. I cannot trust that actions I invoke will actually happen. I cannot trust that taking action will not have mysterious side effects (e.g. re-opening an issue my boss’s boss closed in 2017 without explanation). I think I trust the underlying git infra but at this point I’m not sure why I do.
I say low quality because it is extremely slow at random moments, gets into broken inconsistent UI states, has baffling UX choices that make it harder to navigate than it should be (particularly with new UIs), and overall just doesn’t have the fit and finish we should expect from software in 1996 - or 2026, or literally any time in between.