LLM AssBench pits Claude, GPT, and Gemini against each other in a cheeky new benchmark
LLM Ass Bench

AssBench is a playful benchmark that ranks frontier LLMs — Claude Fable 5.1 Max, Claude Opus 5.5 Max, GPT Astra 6 Ultra, Gemini 3.1 Pro Extended, Grok 4.6 X-High, and more — with dated thumbnail entries. One model, Claude Sonnet 5 Max, "initially refused, then agreed to middle ground," hinting the test probes how models handle edgy requests.
Initially refused, then agreed to middle ground.
- drywater2
Finally, I was getting tired of seeing pelicans on a bike.
- athrowaway3z
6 months from now we'll be discussing if model training is assmaxxing.
- ravenstine
I'm impressed that Luna got as far as it did on max reasoning. Don't get me wrong, that's one weird looking ass, but even the top GPT model a year ago probably would have done worse based on my experience (with asking it to create meshes in general, not modeling asses... yeah, that's the ticket). Then again, I'm also not, because I've found it to be a great deal at its relatively low price point when the reasoning is set to high.
- variety8675
The idea that they're going to benchmaxx this in the future amuses me
- sethkim
Folks, we've reached the top.