The Conceptual Reasoning Index (alignment.anthropic.com)
78 points by optimalsolver 18 days ago | 54 comments




Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.

0/10


Yep, conflict of interest is the elephant in the room and its absence from the "reducing risks from advanced AI" point list is conspicuous:

* Figuring out how to prevent the incentives of frontier AI labs from aligning with anti-social deployment of AI rather than pro-social deployment of AI

It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check. Even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.

A_D_E_P_T 18 days ago | flag as AI [–]

Opus 5 scored highest, too, which is just lol

Opus 5 is okay at coding. It is unbelievably awful to chat with, though. Very snippy, snarky, and it loves to "push back" even when it's inappropriate to do so. Besides, where "conceptual" things like science and math are concerned, it's starkly inferior to 5.6-Sol. Where writing prose is concerned, Kimi-K3 runs circles around it.

I refuse to believe Opus 5 is at the top of any non-cherrypicked benchmark, unless it has to do with very narrow coding tasks.


Maybe it gets bonus points for constantly being honest about how it didn't actually finish what you asked it to do.

Past Opus 4.6 the models become absolute overfit trash, overfitting + their overconfidence makes them an existential threat to progress in niche areas, there are certain types of cutting edge crdts that are impossible to get a them to work on at all without erasing code. It's shockingly harmful.
gekoxyz 18 days ago | flag as AI [–]

Just by seeing the title I knew it was going to be like the meme of Obama giving himself a medal. Why are they even doing this? Are there still people that trust LLM benchmarks made by the LLM companies? It seems to me that the only people still using Claude are the ones that have it for free at work (me). Some of my colleagues even started using their own private OpenAI subscriptions to avoid using Claude, others are using Gemini flash to decipher what Claude is saying...
eli 18 days ago | flag as AI [–]

Colleagues switching models over Claude's writing style has nothing to do with benchmark validity. Conflating "I don't like the prose" with "the eval is rigged" is just vibes stacked on vibes. Anthropic funding a benchmark is a real conflict of interest; anecdotes about coworkers aren't.
jrflo 18 days ago | flag as AI [–]

The conflict of interest is real there, but some benchmarks really ought to be closed source. Otherwise, the second your benchmark is public labs will overfit their new models on it and it will cease to be useful.
kate90 18 days ago | flag as AI [–]

Has anyone measured overfit-on-exposure though? MMLU stayed useful for years despite being public. Feels like the real risk is narrower: a handful of contamination-prone tasks, not the whole benchmark. Closed-source just makes that assumption unfalsifiable either way.
andsoitis 18 days ago | flag as AI [–]

Opening line:

> A core hope for managing AI risks is that AIs will help us understand our situation

Gonna stop you right there and ask that you think deeply about that premise.

xlayn 18 days ago | flag as AI [–]

This should be the top comment, the implication is: AI will decide if AI is correct or no, so adding a new layer of approval...

Whenever you say, hey that's incorrect, the legal process goes to AI and will decide on that issue...

Will work as amazing as the chatbots of the AI companies to solve your issues...

andsoitis 18 days ago | flag as AI [–]

Also, Hope is not a strategy.

The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean

That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
MontagFTB 18 days ago | flag as AI [–]

I think I’ve seen that movie.
dd8601fn 16 days ago | flag as AI [–]

Seems to me like we’re already doing that with the approval gating, adversarial reviews, et al.

People are otherwise just going yolo mode because they can’t possibly check everything fast enough.

a2ff6eeb0 18 days ago | flag as AI [–]

You just described the end of humans making decisions about their future.
kakugawa 18 days ago | flag as AI [–]

"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.
pton_xd 18 days ago | flag as AI [–]

That begs the question, what's the plan if AI does not help us understand the situation?
toasty228 18 days ago | flag as AI [–]

That's what we've been doing with tech for 200+ years, why would it stop now. Build cool shit now and let future generations handle the problems
NBJack 18 days ago | flag as AI [–]

I love how many ways we can interpret that line.

"Hey Claude, our stuff needs to make more money. We are at risk for losing more."

"Rest assured, the 'situation' will only worsen if you resist our benevolent offer."

"We're aren't even at AGI yet, but I for one welcome our new agentic overlords."

xlayn 18 days ago | flag as AI [–]

We have totally come out with this idea of this index that will allow us to create policies to ensure only the best and safest AIs are used by the public....

Absolutely no conflict of interest, no lobbying here

And no this is not related to those chinesse models... it's not the same as HD vs honda thing...

Trust us, this is the same kind of amazing thing as boeing doing their own certifications and inspections!

-- First reply: those dumb models, who use them, they are good only for adding 1 + 1 -- another: They will do the same so who cares... -- The valve guy is a dick, and has a monopoly...

onomojo 18 days ago | flag as AI [–]

We invented a new benchmark and look we're at the top. Everyone else sucks compared to us. Especially those dirty open models.
skerit 18 days ago | flag as AI [–]

Having Opus 5 on top of Fable is even more suspicious. Whatever this is measuring, I do not care. Opus 5 is shite compared to Fable.
eutropia 18 days ago | flag as AI [–]

This is a modest start on an important direction for AI Alignment work; which is, as the authors observe, commonly comprised of tasks which are not readily empirically verifiable and not easily mathematically modeled - so it's hard to get at with normal RL techniques.

I find the ACCoRD benchmark the most interesting, because you could theoretically scale it up from the baseline mode of testing two instances of the same model for their `P(A) ≥ P(A&B)` respectively, you could do `P(A)≥ P(A&B) && P(A) ≥ P(A&C) && P(A&B) ≥ P(A&B&C) && P(A&C) ≥ P(A&B&C) ...` etc

i.e. a swarm of model instances could be collectively measured for consistency for even more confidence, right?

At any rate, even the basic idea of measuring a model for consistency in beliefs improves our ability to bound the amount of trust we can put on it with introspection methods.


> For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)?

So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this

andai 18 days ago | flag as AI [–]

> Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output.

Straightforwardly true.

But doesn't this smell like asking an organization to design its own oversight and guardrails?

Even before you get to alignment issues, LLMs are really good at generating content that "sounds right" to humans. That's basically what they've been hyper optimized for.

At least with math (and to some degree software) we can verify the result. But with the intersection of science fiction, philosophy and ethics... not so much!


I often wonder if the people at anthropic actually believe in the existential risk stuff they talk about. If I was convinced AI could kill everyone, my actions would be almost diametrically opposite to theirs -- I would be committed to ceasing the operation of AI labs at any cost.

It's easy to say they're just lying and it's marketing hype, but I also think it's possible they really believe the best way to stop an evil superintelligent AI is to work on making superintelligent AIs, but just do it smarter than people at OpenAI would. Humans are bizarre creatures!

_dwt 18 days ago | flag as AI [–]

I don't know about the top people but Anthropic has always been pretty deeply linked with the effective altruism/AI risk/TESCREAL memeplex. IIRC one of OpenAI's people who's pretty prolific on Twitter weighed in a while back to the effect that Anthropic mostly sees themselves as an organization dedicated to serving Claude. Not serving like server but serving like... worship. I find them a little freaky to be honest.
ember93 18 days ago | flag as AI [–]

Safest hands theory works great until someone else's hands touch it too.
behnamoh 18 days ago | flag as AI [–]

I don't remember any company in world's history that has both been loved and hated by the same users who purchase from it. We love Anthropic for its amazing models, and we hate them for all the shenanigans around the models, including their marketing.

I kinda wish they had not made a comeback after Claude 2.


It is so bizarre, honestly. It's almost like they are a public utility company. People are trying as hard as possible to hate them, but always stop short of simply canceling the service because they can't live without them. Seriously, I saw someone say that they were so disgusted with Anthropic that they were cutting down from having two Max subscriptions to just one.
velcrovan 18 days ago | flag as AI [–]

I enjoy using the models. I also get that there are shenanigans and that marketing is happening, but as long as the models are this effective I can't much bring myself to care. I suspect most of their users are the same.
andy99 18 days ago | flag as AI [–]

They have to be careful. There’s not a lot that separates the top models anymore. Look at Grok which has almost no market share or credibility because of Musk, Despite being close to the top performance wise they are behind the Chinese models in adoption for example, which themselves are mostly behind the big two because of trust. It would not be hard to push users away, especially if a less unpalatable player ever emerged.

I think this says as much as the leader on the index as it says about the loser. If I believe I do not want to trust the AI to know best what is it that I want done, then this tells me not to go with Anthropic.

Is a new benchmark that useful if existing model improvements are being reflected linearly? Don't we want a benchmark that we aren't seeing much progress in.
andai 18 days ago | flag as AI [–]

> To compute Fable 5's score, we used Opus 5 as a fallback in cases where Fable 5 refused to answer a question.

Why are they running the gimped version for internal evals?


Looks like they don't have any responses to others players occupying the news. Everyday you see a blog post that doesn't address our daily concerns.

I think we can mostly eyeball it at this point. There hasn't been a model that I've thrown a novel problem at that didn't turn into an iterative token bonfire until I intervened and until that has changed, most of these benchmarks feel kind of like pointless marketing slop.
chrisjj 18 days ago | flag as AI [–]

How many benchmarks did they have to reject in the search for one that scores them top?

New Trust Me Bro benchmark just dropped
onyx42 18 days ago | flag as AI [–]

Yeah every eval writeup follows same pattern. We ran into this benchmarking our own model calls, self reported numbers always look great until someone else runs the same prompts on a different rig and gets different results. Only fix was independent eval harness, not trusting the vendor's own dashboard.
Scribbd 18 days ago | flag as AI [–]

And surprise! We are leading!

Do we really know if these scores make any sense?
Aeroi 18 days ago | flag as AI [–]

missing the the 2010 era of technology. it's becoming information overload.

I would love for someone to give me a coherent argument as to how this isn’t tone-deaf, vacuous garbage.
valdork59 17 days ago | flag as AI [–]

The issue with LLM generated prose and code is that on the surface it looks meaningful but then when you dig into it, it's hard to find any kind of logical thread
dnb1 17 days ago | flag as AI [–]

The harder problem here isn't the leaderboard, it's construct validity: does "conceptual reasoning" as operationalized actually track the underlying capability, or just performance on this particular task distribution? Anthropic's own faithfulness work suggests stated reasoning and actual computation diverge more than we'd like.