LLM AssBench pits Claude, GPT, and Gemini against each other in a cheeky new benchmark

LLM Ass Bench

LLM AssBench pits Claude, GPT, and Gemini against each other in a cheeky new benchmark

AssBench is a playful benchmark that ranks frontier LLMs — Claude Fable 5.1 Max, Claude Opus 5.5 Max, GPT Astra 6 Ultra, Gemini 3.1 Pro Extended, Grok 4.6 X-High, and more — with dated thumbnail entries. One model, Claude Sonnet 5 Max, "initially refused, then agreed to middle ground," hinting the test probes how models handle edgy requests.

Initially refused, then agreed to middle ground.
  1. drywater2

    Finally, I was getting tired of seeing pelicans on a bike.

  2. athrowaway3z

    6 months from now we'll be discussing if model training is assmaxxing.

  3. ravenstine

    I'm impressed that Luna got as far as it did on max reasoning. Don't get me wrong, that's one weird looking ass, but even the top GPT model a year ago probably would have done worse based on my experience (with asking it to create meshes in general, not modeling asses... yeah, that's the ticket). Then again, I'm also not, because I've found it to be a great deal at its relatively low price point when the reasoning is set to high.

  4. variety8675

    The idea that they're going to benchmaxx this in the future amuses me

  5. sethkim

    Folks, we've reached the top.

More from this day

2026-09-22