57 points by artninja198817 days ago | 58 comments
How to play: Some comments in this thread were written by AI. Read through and click flag as AI on any comment you think is fake. When you're done, hit reveal at the bottom to see your score.got it
"We believe our internal AI R&D efforts are
significantly faster than they would be without AI assistance, but not yet
by a factor of 2 (though we are uncertain and measurement is difficult)"
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.
It is also difficult to measure this in the face of ever shifting baselines. Anthropic is built on AI from the ground up. I'm not picking a side, but Anthropic getting a 2x increase is different bar than some random enterprise shipping legacy apps with clunky processes getting a 2x increase by adopting.
Yeah, marketing numbers and internal ops numbers rarely line up. We've seen same thing with our own tools, demo doubles output but actual sprint velocity barely moves once you count review time and cleanup.
Ran similar internal timing tests on a smaller team last year, comparing PR cycle time before/after Claude assistance. Signal was there but noisy as hell, maybe 20-30% on routine tasks, way less on anything architecturally hairy. Self-reported multiplier estimates are almost always inflated.
So as of a month ago their best internal model was "somewhat more capable" than Mythos "but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview." I thought they would have a significantly more capable model by then, more than five months after Mythos finished training. They'd better have one by now, or the Chinese competitors are closer to catching up than I thought.
I'm still uncertain if mythos is real. Subsequent model releases have been lackluster, no one has claimed to verify mythos performance and it's silently vanished from most comparisons.
Reminds me less of Bannon's "flood the zone with shit" and more of standard attention economics: publish frequently enough and critique can't keep pace with output. Same dynamic shows up in any fast-moving research field, not unique to labs.
"all traffic through our systems for collecting human feedback data from contractors evaluating our models ran without blocking biological classifiers"
"totaled around 133M exchanges."
While this wound up being relatively benign, I still find this concerning, amidst numerous sandbox escapes, and previously, unreleased models being accessible via a custom URL. I don't think these companies are giving the responsibility they possess enough weight. How many more issues like this exist?
Concordia AI's tracker good starting point but if want raw primary docs, check each lab's own safety framework — Baidu, Alibaba, Zhipu all published versions. Diff em against Anthropic's RSP side by side, gaps jump out fast.
> Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.
Meanwhile I can't really tell the difference between Fable and Opus for my tasks. I kinda think Fable does a better UX work so I keep using it for that because I couldn't be bothered to A/B them, but otherwise it's all the same and the model and effort are just feel good knobs I twist to still remain a load-bearing element. At least that's my honest take.
Fable was amazing during the first preview. Once they added it back, the limits are too low to get anything done. I might use it in chat if I remember to select it once a month but don’t even bother to try and code with it.
For the past 2 weeks or so I've been doing the A/B test, sending identical prompts to Fable 5 and Opus 5 to test their ability to produce design documents for new feature work. I've consistently found that Opus 5 produces more complete, accurate and "imaginative" designs than Fable, often finding design issues or nearby bugs that Fable 5 misses. However, that creativity means Opus seems to hallucinate more, while Fable's design is clearly based on the actual existing code. Or as Opus put it: "I hedged — [Fable] checked."
By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.
> However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have “saturated”—i.e., no longer capture increases in models’ capabilities—and because we are seeing early signs of acceleration.
I totally understand this is a subset of alignment-related evals, but if Anthropic of all is running out of evals, doesn't that also means we are running out of things to scale?
I mean. I totally believe they have a model that is better at Kernel Optimization, creating new matrix multiplication algos, than Mythos. But it's clearly no generalizing, rightw
My friends and I, and the teams I'm a part of, just want to build and create fun, cool things. I am so tired of being preached to by Anthropic like they're some arbiter of 'ethics.' So, so tired.
> 6.2 [Appendix redacted]
> This appendix describes the criteria for our blocking bioclassifier exemption policy, and has been redacted from the public version of this report for security reasons.
>6.3 [Appendix redacted]
> This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.
interesting...
EDIT: After reading more I'd recommend looking at Transcript 2.20.A. Its a transcript of claude going over the redactions in the report. The section says its specifically for section 2, but the transcript also mentions other sections.
Doubt bankruptcy risk and AGI risk are separate things. Burn rate exists because they're racing capability, so the two arguments loop back into each other whether people like it or not.
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.