Gemini 4 Argon (High) Tops Artificial Analysis Intelligence Index

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis

Gemini 4 Argon (High) Tops Artificial Analysis Intelligence Index

Artificial Analysis has published its intelligence, performance, and price analysis for Gemini 4 Argon (High). The model is evaluated on the Intelligence Index v4.3.2, which combines ten benchmarks including AA-Briefcase, Terminal-Bench 4.0, and Humanity's Last Exam. The analysis also covers cost per task, token usage, context window, and cache pricing, offering a detailed look at the model's efficiency and capabilities.

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
  1. mchusma

    I’m guessing this is considered something like a C grade from Google if they are being honest with themselves.

    After being nowhere near the frontier for a long time, they are pre announcing a model that ranks 3rd, roughly on par with models today that are cheaper.

    Good for them to think about releasing to stay in the frontier game.

    (I do think 3.7 flash was a solid release, so they are around the conversation. And their image and audio and live models are good)

  2. aliljet

    It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king

    Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...

  3. mlmonkey

    Meanwhile I'm still being offered Gemini 3.1Pro on gemini.google.com :-D

    https://imgur.com/a/h96yg5t

    This is on a $20/mo paid plan :cry:

  4. godbox

    What a snooze fest. Another model that does not meaningfully improve on intelligence or price compared to its peers. Google has basically announced that they've "caught up" with the rest. I think they've been doing great work in the Flash department so seeing this is... underwhelming?

  5. lhk931122

    I'm not sure Google can cut in when Claude and ChatGPT already got. I'm using both, but I'll keep track of whether Google can make it good enough for me to use also this Gemini, or cancel one of the two (Claude and ChaGPT) for it.

More from this day

2026-09-30