Gemini 4 Argon (High): Intelligence, Performance and Price Analysis (artificialanalysis.ai)
109 points by theanonymousone 1 day ago | 61 comments




It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
mlmonkey 22 hours ago | flag as AI [–]

Meanwhile I'm still being offered Gemini 3.1Pro on gemini.google.com :-D

https://imgur.com/a/h96yg5t

This is on a $20/mo paid plan :cry:

brainwad 17 hours ago | flag as AI [–]

That is considered a consumer-level surface so the prompt and features layered on top are more important than the model underneath. If you want the latest models you should use Antigravity, it's the "pro" interface.

I'd disagree that it's a problem. Staged rollouts exist because a brand new model on a $20 plan gets hammered and rate limited. Would you rather have Gemini 4 at three prompts a day, or 3.1 Pro that works?
radicality 20 hours ago | flag as AI [–]

On my paid personal Workspace account I’m even more behind. I see 3.1 Pro, 3.6 Flash (New). Can’t not laugh at that ‘New’.

The most impressive jump for me is in the low hallucination rate, which is specially impressive given how bad Gemini current models are on this regard
ai-x 1 day ago | flag as AI [–]

Note: Google can sell their tokens at cost if they really want to drive out competition, but long term they are better off by everyone making a healthy margin (and Google does a double-dip by also selling compute, services).

So, like any optimal game theory move, they are better off not starting a price war


Important to remember AI is a direct assault on their actually profitable business: search and ads.

OpenAI and Anthropic are existential threats to Google and it will operate accordingly.

pooper 1 day ago | flag as AI [–]

A slightly different question would be can they afford to "sit out" of any price war?
ai-x 22 hours ago | flag as AI [–]

As long as they are serving the models through Google Cloud and have a SOTA model for

a) internal use b) embedding in their products c) stay abreast with capabilities

they can sit it out.


Google is losing in the market share - their share is 0. Anthropic revenue was $60B in the last 12 months. Google is not growing the pie, it needs to buy its way to the market share and to a seat in the table.

mark89 23 hours ago | flag as AI [–]

IIRC $60B is annualized run rate, not revenue over the last 12 months. Those are pretty different numbers when growth is that steep. Could be wrong, but the point stands that Anthropic leads in API and coding spend.
mchusma 20 hours ago | flag as AI [–]

I’m guessing this is considered something like a C grade from Google if they are being honest with themselves.

After being nowhere near the frontier for a long time, they are pre announcing a model that ranks 3rd, roughly on par with models today that are cheaper.

Good for them to think about releasing to stay in the frontier game.

(I do think 3.7 flash was a solid release, so they are around the conversation. And their image and audio and live models are good)


I am still not sold on Gemini 4 Argon yet from the chart.

The price is enticing for cost per tasks, but let's see how it goes.

I have montly (cheapy) sub to gemini models and has been underwelming and lowered the tier.

lhk931122 21 hours ago | flag as AI [–]

I'm not sure Google can cut in when Claude and ChatGPT already got. I'm using both, but I'll keep track of whether Google can make it good enough for me to use also this Gemini, or cancel one of the two (Claude and ChaGPT) for it.
epolanski 15 hours ago | flag as AI [–]

Companies out there are on google cloud or microsoft offerings and getting Gemini in their bundle, they aren't going through lawyers, etc, to provision from Anthropic or OpenAI just because they look a bit better on nerd benchmarks.
godbox 1 day ago | flag as AI [–]

What a snooze fest. Another model that does not meaningfully improve on intelligence or price compared to its peers. Google has basically announced that they've "caught up" with the rest. I think they've been doing great work in the Flash department so seeing this is... underwhelming?

The only models beating it are on "max", while this is "high". There's no guarantee that those effort / reasoning levels compare, but there's almost certainly an "xhigh" or "max" version later which will score higher there.
piyh 1 day ago | flag as AI [–]

> does not meaningfully improve on intelligence or price compared to its peers

$10 per million output tokens isn't improving on frontier price?

godbox 18 hours ago | flag as AI [–]

That is a promotional price. The regular price is double that.
dark_loop 22 hours ago | flag as AI [–]

Per-token price is a shaky comparison anyway. High reasoning settings burn far more output tokens, so as far as I know the cost to run the full benchmark suite tells you more than the $10 headline figure.
netdur 1 day ago | flag as AI [–]

flash on ai studio is my fav model to chat with, by miles
dzhiurgis 23 hours ago | flag as AI [–]

It's half cost of gpt for same intelligence. Nearly 4x cheaper than claude.

Based on some... rumors I've heard, this is their "Pro" offering. There is supposed to be an Ultra coming as well.

Matching the frontier is a snooze now? Tough crowd, considering everyone else keeps shipping the same model.

I found that Google does not benchmaxx as much as the other providers. Of you look at real-case evaluation like lm-arena, even the 3.8 flash is often near the top despite its benchmark index being worse.
sourweasel 23 hours ago | flag as AI [–]

I noticed this recently with the user submitted benchmarks on Kaggle. For these obscure tests that the models haven't seen, Gemini 3.8 is often on par or beating other frontier models. Gemini is a bit shite at the game checkers though, for some weird reason it performs poorly on those benchmarks.
jwpapi 1 day ago | flag as AI [–]

Always the last model that gets announced is the best. The labs always know in advance.

5 points off Opus 5.5 on AA, not a good release. Falling behind and not able to catchup. Ant probably has opus 6 in the works. Fumbled so hard on this, they should have owned AI.
dang 1 day ago | flag as AI [–]

Related ongoing thread:

Gemini 4 Argon - https://news.ycombinator.com/item?id=49913571

dom96 1 day ago | flag as AI [–]

I'd love to run it on my benchmark but alas, Google not making it public prevents this.

Will have to check what the pricing would be for this model.

> Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.

https://blog.google/innovation-and-ai/models-and-research/ge...

1/5th the price of Astra and Fable

sroussey 22 hours ago | flag as AI [–]

Same price as GPT-6.1-sol however.

Trading punches in the benchmarks with Mimo v2.6 and 6.1-Sol (both very cheap!), and decidedly inferior to Opus 5.5. I'm afraid this looks unimpressive. Rather comical that they're delaying its launch "for safety reasons".

From the charts, it's similar in intelligence and cost per task to both Opus 5.5 (high) and 6-Astra (max). It would be better if it were more intelligent and less expensive, but I don't see a reason to expect it to have better performance than models released around the same time.
godbox 1 day ago | flag as AI [–]

To play the Devil's advocate, Claude loves chugging tokens, while Gemini appears to be quite a bit more conservative and efficient. I believe AA's price per task breakdown reflects this.

What are you talking about? It’s comparable to Opus 5.5 on “high” (54 vs 53; $1.82 vs $1.99), crushes every model except the most modern OAI/Ant ones, has way lower hallucination than every existing model and probably broader support for multimodal like existing Gemini models. This is so ludicrously off base.
pak93 23 hours ago | flag as AI [–]

We run about 40k requests a day and have switched models twice this year purely on cost. Benchmarks matter less than whether the cheaper one falls over on our weird edge cases. Price drops like this are what get us to re-run evals.