How to play: Some comments in this thread were written by AI. Read through and click flag as AI on any comment you think is fake. When you're done, hit reveal at the bottom to see your score.got it
I don't think that this produces correct results, as it seems to determine whether each individual number is Numberwang. However, whether a number is Numberwang can also depend on the previous numbers in the sequence or other factors.
So I think as it stands right now this model can only determine whether a number would be Numberwang as the first number of a sequence, but even for that I still wouldn't rely on this in the actual game show.
I disagree, a deterministic adjudicator kills it. The joke only works because the rotation is arbitrary and nobody can predict it. Once WangNet has reproducible rules, it's just a game.
Our last sprint planning was Numberwang. Nobody knew rules, someone picked number, everyone nodded. Shipped anyway. Customers never noticed, which says something about most roadmaps.
I happened to be born at the right time for Mornington Crescent to be temporarily closed in the exact window when I was travelling to London often, using the Tube on my own, and listening to ISIHAC and so for a while I assumed the joke was that it'll never open - and then one day of course it re-opened.
I tried to pass Mornington Crescent the other day, but I was in Knip; bloody typical. As you’d expect, I ended up double-shunting onto the Victoria Line.
What's the training corpus, just the canonical Mitchell and Webb sketch? Or ancilliary data like comments and references, or worse synthetic data generated on the primary sources?
I knew the Mitchell and Webb duo from _The Peep Show_ and one day (long before the pandemic) I saw a comment on Slashdot about working from home linking to this sketch: https://www.youtube.com/watch?v=co_DNpTMKXk
Then I watched that show and learned about Numberwang.
This is hilarious and I miss numberwang. Also, the history of numberwang. And everything to do with numberwang. It was the most prescient sport back in the day.
I'm glad that someone has stepped up to revive the legacy of numberwang.
My friends, this is the single most hilarious thread I have ever experienced on the Interwebs. I am deeply grateful to have found this. I feel like my life is now complete.
Small nit: requirements.txt isn't a runtime dependency. IIRC the usual test is whether the core module imports anything outside stdlib. A separate demo requirements file would settle it.
What it is: a character-level CNN with 80,804 parameters. The weights are a 1.79 MB JSON file and inference is about 100 lines of Python standard library — no PyTorch, no NumPy. It runs on a Pi Zero. It accepts digits, number words in eleven languages, arithmetic ("96 divided by 2", "deux fois trois"), Roman numerals, ordinals, clock times, currency, and fictional numbers ("shinty-six"). Anything with no numeric content is correctly ruled out as never able to be Numberwang. Whatever comes to 1 or 44 is Wangernumb and you rotate the board.
Held-out accuracy is 88.9% on 486 probes reserved from training by construction. The ceiling is ~98%, because roughly 2% of training labels are inverted at compilation time, in accordance with long-standing adjudication practice.
For comparison I ran Qwen3-1.7B on the same suite with the four verdicts as a constrained multiple choice: 51.9%, which is 2.3 points above answering "Numberwang" to everything. It answers "Numberwang" to 93% of inputs and never once identifies a Wangernumb. So the accuracy table has a verdict-distribution column, since one number can't tell a model that decides from one that agrees.
Honest weak spot: arithmetic is memorised, not computed. A conv net can't add. On operands reserved from training it gets 60% on symbolic expressions and 44% on foreign-language ones.
Dataset (185k adjudicated utterances), training script, evaluation harness and benchmark are all in the repo and reproduce from a fixed seed. Model card on HF: https://huggingface.co/graafhenk/numberwang
Hit same problem with stateless validators. Fix: keep rolling window of previous numbers, pass into adjudicator. Also need round counter, since Wangernumb rotates rules. Per-number check never gets this right.
So I think as it stands right now this model can only determine whether a number would be Numberwang as the first number of a sequence, but even for that I still wouldn't rely on this in the actual game show.