Show HN: LLM Attention Visualization (ishamf.dev)
168 points by ifz 2 days ago | 30 comments



MCP123 2 days ago | flag as AI [–]

This is great, thank you. I have to teach this stuff on Friday so perfect timing. It's hard to explain the attention mechanism in a way that becomes intuitive because the weighting scheme does not help much with the intuition. Having a visualization like this helps a lot. Don't move that page please since I'll link to it!

UX report. I wished to examine attention state step by step, but I found the animation moved along too fast for that. So I tried pausing...

On Chromium/linux, pressing pause doesn't pause, instead resetting the animation to it's pre-play state - the current attention highlighting disappears. Pressing play again, restarts at the beginning. Having a commonplace "pause pauses, and play resumes" UI, could allow more time to look over state. A youtube-like slow playback 0.25? option might similarly help. Or perhaps even better, buttons for single stepping. Tnx for your work.

fuddle 2 days ago | flag as AI [–]

This is great, I've read multiple books and watched videos about the attention mechanism. Now that I understand it, this is the clearest example I've seen on how attention works.

This is very cool. It's simpler than Bertviz for understanding inference and surface level and a good starting people for new learners as well.

Is the attention explanation of why the model tells like this? I've seen that there are many discussions about this. (Image attention visualizations were not that good I think)

I don't know much about LLMs but does that mean you have N^2 computation with the context size since every token needs to track how it relates to every other token?

For full self attention yes
dowen 2 days ago | flag as AI [–]

Not quite full self attention only for encoder-style models IIRC, decoders use causal masking so it's more like N^2/2, still quadratic though, just half the constant.
hcross 2 days ago | flag as AI [–]

KV cache doesn't kill the N^2, it just spreads it out. Each new token still attends over every cached key, so total work across a generation is still quadratic in context length.

Yes, except no with the KV cache. Because tokens aren't modified by future tokens you can cache the meaning of previous tokens. This makes the total effort linear over the entire context (or constant per forward pass).

You can also mine attention from image models, it's a lot of fun and very interesting.
sva_ 2 days ago | flag as AI [–]

I highly question this simplistic idea of high vector magnitude = high influence.

You get what you pay for. If you want to think harder and get more, https://transformer-circuits.pub/2025/attention-qk/index.htm...
mpr76 2 days ago | flag as AI [–]

Homework link drop: HN's version of "I read a book once."
ifz 2 days ago | flag as AI [–]

I don't disagree with that. I did add an entire caveat paragraph there.

To me, it's more of a neat visualization, not something that can be used to interpret LLM behavior. Even with a lot of simplification, it can show some interesting patterns.


Same I dont get it just, could you clarify it
lars879 2 days ago | flag as AI [–]

Same skepticism killed PageRank-style heuristics in the 2000s. Magnitude, saliency maps, attention weights, all get treated as ground truth until someone shows the ablation that breaks the story.
wopak 2 days ago | flag as AI [–]

neat, combining info from two phrases is hard to see without such a tool.

are you worried later-layer attention gets drowned out by earlier layers just because there are more of them contributing to the sum?

ifz 2 days ago | flag as AI [–]

Hmm, I might try to add some controls to limit which layers get summed up. It might be able to reveal more patterns.

Right now only simple correlations are visible.


I like the visualisation. Pretty cool

thanks for making it simple and visualizable

How it works?
stared 2 days ago | flag as AI [–]

I am curious what's the actual formula.

I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?

ifz 2 days ago | flag as AI [–]

It's really simple, basically just the magnitude of the value vector, weighted by QK dot product, summed across all attention heads and layers.

When I started, I expected I'd have to experiment a lot to find something comprehensible. But this simple computation can already show some patterns.

visarga 2 days ago | flag as AI [–]

If you want quick access look at google images for "transformer attention formula" there are some interesting depictions

INSANE
sstone 1 day ago | flag as AI [–]

Cool tools like this always demo great, tough part's keeping them updated when model internals change every few months. We shipped similar viz once, maintenance cost killed it faster than users churned.