GitHub is having trouble counting things (chuckgreenman.com)
77 points by chuckgreenman 15 days ago | 59 comments



nxc18 15 days ago | flag as AI [–]

GitHub, including Enterprise (so not just Azure’s fault), has become unreliable, low quality software.

I say unreliable because I cannot rely on it to accurately do what it claims to do. I cannot trust any number. I cannot trust that actions I invoke will actually happen. I cannot trust that taking action will not have mysterious side effects (e.g. re-opening an issue my boss’s boss closed in 2017 without explanation). I think I trust the underlying git infra but at this point I’m not sure why I do.

I say low quality because it is extremely slow at random moments, gets into broken inconsistent UI states, has baffling UX choices that make it harder to navigate than it should be (particularly with new UIs), and overall just doesn’t have the fit and finish we should expect from software in 1996 - or 2026, or literally any time in between.


For months now I've had scheduled GH Actions workflows that don't run on time, sometimes not at all. Like we're talking 4-8 hours later than the cron statement. Yesterday I finally gave up and just moved the trigger to AWS EventBridge Scheduler/Lambda. Maddening, but also we're talking code repos that are $0 for me. Not production level. I can't imagine people running Production pipelines on github.com.

My scheduled for 8am PT GH actions workflows don't run until anywhere between 6-8 hours later. I gave up, and run them locally now. It is just to create a tag, and push to a specific place. A super simple problem that was solved for so long simply with GH Actions that just doesn't work anymore.

I've switched off to runners and event scheduling inside AWS, and am working on moving my company off of GitHub. I know engineers inside GitHub and it doesn't sound like the talk they've been putting out about reliability is actually being addressed with many resources internally; the best people in the company are chasing AI product features (which are apparently quite profitable).

At a previous gig we had a simple "reset staging" task on gha, that ran once a week on Tuesday

It ran fine for years, then around November just stopped and never ran again. Tried moving it around for which time it would trigger, and it just never started again. Opened a ticket, and got nowhere

martinzen 15 days ago | flag as AI [–]

Nit: GitHub's docs say scheduled workflows are best-effort, and IIRC delays at the top of the hour are expected. But 4-8 hours is way beyond that. Also, public repos get cron disabled after 60 days of inactivity, if that's biting you.
idkasam 14 days ago | flag as AI [–]

You should try different runner solution. We use Avrea.com and works super well for our workflows.
shmoil 15 days ago | flag as AI [–]

>> GitHub has become unreliable, low quality software.

They are owned by Microsoft, what did you expect exactly?

crate80 15 days ago | flag as AI [–]

GitHub was bought in 2018 and was mostly fine for years afterward, so ownership alone doesn't explain much. The decline looks more recent. As far as I know the Azure migration is the likelier culprit, though that's inference from outside, not anything GitHub has confirmed.
zX41ZdbW 15 days ago | flag as AI [–]

Some of the oldest PRs could no longer be found.

For example, when I decided to continue my eight-year-old PR https://github.com/ClickHouse/ClickHouse/pull/104948, I couldn't see it - neither in search nor while navigating through pages.

Direct links still work.


Dirty reads is a classic symptom of scaling poorly. Eventually consistent database writes are fine, but it's a bad look if a user takes an action and then doesn't see the results of their action.
wren6991 15 days ago | flag as AI [–]

Similar to how every description of relaxed consistency for CPUs starts out by pointing out your own reads always observe your own writes in program order, as if to say, "don't worry, we're not insane."

Or funnier version: "don't worry, we're not Alpha".
hanspagel 15 days ago | flag as AI [–]

I think they struggle with two things

1) counting

eugenekay 15 days ago | flag as AI [–]

2) cache validation

3) off-by-one errors

kibwen 15 days ago | flag as AI [–]

According to the documentation there are four things:

4) naming things

5) keeping the documentation up-to-date


There are solutions: https://xkcd.com/3062/
rhdunn 15 days ago | flag as AI [–]

5) multithreading

4) database synchronization

zbentley 15 days ago | flag as AI [–]

Ideally written as:

5) multit

6) database synchreadinghronization

lbanchio 15 days ago | flag as AI [–]

4) naming things
stateoff 15 days ago | flag as AI [–]

"Since we forced our engineers to agentic workflows our efficiency doubled. Just look at the numbers!"

The fact that there's so much indirection between what happens in the registers and memory in the machine, the underlying reality being modeled by software, and the information displayed semantically to the users, is a travesty.

Software could be so much simpler and more reliable. There's so much bloat that is necessary to solve problems created by bloat. I hope we find our way out of this mess sooner rather than later. I'd hate for entire generations to suffer the current state of software whereas the theoretical understanding necessary to make things better was produced very early in the history of programmable computers.

Droobfest 15 days ago | flag as AI [–]

I was having the problem yesterday that the GitHub API sometimes simply doesn't list an in progress workflow run in the output when I specifically filter for it:

"api.github.com/repos/{owner}/{repo}/actions/runs?status=in_progress" -> outputs 3 workflow runs

5 seconds later -> outputs 2 workflow runs

5 seconds later -> outputs same 3 workflow runs again

Apparently they make no guarantees of any type of consistency or continuity in output, which has the side effect of also making it completely useless.

stabbles 15 days ago | flag as AI [–]

The reason is that their primary source of truth (traditional database) and their search index (elasticsearch) are out of sync.

The issues/pulls pages used to be showing the data from the primary data source, and then they changed it so everything is search, including the basic props is:pr and state:open.

oxidant 15 days ago | flag as AI [–]

Experienced this using the API. Got told it's a "wontfix"
ako 15 days ago | flag as AI [–]

Don't worry, it will be consistent eventually.
cedws 15 days ago | flag as AI [–]

There's tonnes of UI bugs. A more severe one I've experienced is that my approved PRs have been shown as still waiting on review, leading some of them to be delayed by weeks.
esafak 15 days ago | flag as AI [–]

I'm partial to the one where the workflow completes but the its summary on the PR does not get updated, so it appears to be running.

How many companies engender taste in bugs, eh??

Rooster61 15 days ago | flag as AI [–]

> Really curious as to what’s going on over there, if you’ve got any insight let me know!

Microsoft. That's the answer. Microsoft is happening over there.


These days I was surprised to see, on my personal computer, information from PRs in private repos from my work’s org on GitHub. These are usually secured behind SSO flows which I don’t/can’t do on personal devices. So now I also don’t trust them to keep our code private anymore either. Nice.

It used to be that you could have a single GitHub account and join multiple orgs, some of which have enterprise plans with GitHub. They completely fumbled this during the pressure to ship AI features. Some enterprise controls that are supposed to affect only my work at a specific GitHub org affects me account-wide. I’m part of quite a few orgs on GitHub for OSS projects, and I never consented for that one org’s policies to basically take over my account.

I thought of going through the process of creating another account specifically for my work at this company, but GitHub is itself solving the problem in a very unique way: by being a place I don’t want to be in after I clock out.


This seems to happen whenever Microsoft buys a company and "assimilates" it into its culture. The same thing happened with Skype.

That said, Microsoft isn't alone. The YouTube Android app has so many basic bugs that I sometimes wonder if the developers have tried asking an AI to "find all the bugs" in their code. It might do a better job! Large blank spaces in downloads, incorrect watch states, forgotten watch states, etc.

bob1029 15 days ago | flag as AI [–]

This appears to be a security problem in some contexts. I've been able to see issue counts for repositories that I have not been granted access to yet (the view with the invite accept button). I can't actually get to the issues but I can see how many there are.

GitHub's downfall must be studied. At the same time as all of these reliability issues, they've been introducing a variety of small UI updates that don't make any meaningful updates either...
bakugo 15 days ago | flag as AI [–]

That's just AI at work. It's much easier to ask Copilot to make some unnecessary UI change or add a new minor feature nobody needed, than it is to ask it to fix the fundamental stability and reliability issues that plague the software.

And knowing how these sorts of large corporations are structured, I wouldn't be surprised if the people who spend a few hours prompting AI to add some new unnecessary feature are being rewarded more than the people breaking their backs trying to stop the site from collapsing under the weight of 100 million vibe coders.


The same counting problem exists with GitHub issues. I notice this every time I close an issue, and the count of open issues is still one higher than the actual number (until I reload the page).
disko 15 days ago | flag as AI [–]

I have got an insight alright. Microsoft took it over.
barryler 15 days ago | flag as AI [–]

I disagree. GitHub had plenty of outages before 2018, remember the unicorn page? It got noticeably worse with the Azure migration and all the Copilot/agent traffic. That's a scaling problem, not an ownership one.
herbst 15 days ago | flag as AI [–]

Soon most will have forgotten how essential and stable github was at some point in the past.
meerita 15 days ago | flag as AI [–]

I know this pain because, right now, I have zero PRs opens, yet the PR tab is always saying I have 1. It's plainly stupid.
busymom0 15 days ago | flag as AI [–]

The chat notifications count badge on Reddit website has the same problem for me. Tells me I have messages when I don't. Or won't tell me I have messages when I do.

I always figured this was intentional because i had been deliberately trying not to use the chat feature but they were pushing it really hard in marketing and, well, they got me to click... And also to stop using the site entirely
tapete1 14 days ago | flag as AI [–]

I like how this guy has a needless problem only by using GitHub, which no one should use in the first place, but then it compounds because he uses an inferior browser. Imagine making your work life hard for yourself on purpose.

I have been seeing this for one-two years already. They cannot count how many packages I have.

https://github.com/vanyauhalin?tab=packages

booi 15 days ago | flag as AI [–]

Press "Next"
mococa 15 days ago | flag as AI [–]

AI will solve this

rails + russian doll style caching + large engineering team + azure as a platform = recipe for this exact problem.

My phone consistently says that I have one more update than they'll actually let me download... what that one phantom app is, I'll never know...
saejox 14 days ago | flag as AI [–]

it is a feature called eventual consistency.

so nothing will be %100 correct, and we will always live in the past.


This got to be the extinguish phase from the EEE Microsoft playbook. Prior acquisition, Github was very liked and very focus, then Microsoft happened in their embrace phase, telling us EEE was something of the past. While trying to show good faith they extended the platform and this is now the very last phase
yodon 15 days ago | flag as AI [–]

>EEE

My god what an ancient trope.

Presumably you realize most current Microsoft employees were not even alive when the EEE memo was written.

Must we also talk about some decision Henry Ford made in 1906 every time the automaker is mentioned, or something the Gauls did in 200 BC every time England is mentioned?


Maybe, but does Team Foundation Server even exist anymore? What is the MS branded service that degrading GitHub is supposed to drive people to?

I don't believe this is true of the past, but it'd be quite the twist if the extinguish phase was really just pure incompetence this whole time. It certainly seems that way right now with GitHub.

What does EEE even mean in the context of a proprietary service that Microsoft owns? That is a strategy for killing open standards, which Github never was. It could be a thing for git itself, but Microsoft doesn't even have a git competitor.

Calling this "severe" feels overblown. It's a UI mismatch likely caused by using multiple caches that are out of sync. I feel like every engineer on this site has encountered this exact bug at one point or another. It's an easy one to run into, especially in bigger orgs.
layer8 15 days ago | flag as AI [–]

The second screenshot with “1 3 Next” seems pretty hard to justify.

Taking care that such incongruent representations can’t happen even in the face of inconsistent backend caches is part of proper UI state management.


The "page two is inaccessible" thing is, imo, severe. I've seen it myself at work, and there are lots of parts of the UI where I'm regularly looking for something on page 2. I didn't even think to try the JavaScript console hack but like, i (my employer) is paying for access to that data, and it's simply not there...

Yes but in the case of GH, it also happens in cases where the data is old: not just a temporary drift.
bgrant 15 days ago | flag as AI [–]

Everyone's hit it, sure. Seen it since the Solr-next-to-MySQL days, circa 2008. But the usual fix is cheap: derive the pager from the same response that renders the rows. Showing a count from a different cache is a choice, not an accident.
matt 14 days ago | flag as AI [–]

We hit this with denormalized counters on a busy table. What fixed it: increment in the same transaction as the insert, and run a nightly job comparing the cached number against a real COUNT(*). Drift shows up fast once you're actually checking for it.