Contrastive Language Models: 9× faster than Jev with SOTA agentic coding

CLM-8B is a new System One model trained with a contrastive objective to connect states and actions. It matches Jev on computer-use, gaming, and tool-calling tasks while running up to 9× faster, and with fine-tuning sets new SOTA on DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%). The model uses frozen LLM backbones with trainable projection heads, enabling cached embeddings and efficient scaling.

We find that Jev fails to serve as a verifier for long-horizon tasks, performing below the random-selection (Pass@1) baseline.
  1. khalic

    The tech is very cool but for the love of Gaia stop calling it system one, even Kahneman said this (S1/2) is a framework for understanding the brains inner workings. There is no autonomous system to speak of.

  2. vessenes

    >"CLM-8B also sets a new SOTA on challenging agentic coding benchmarks, including DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%).

    I didn't see any details on this on the announce page. And I don't believe it. Astra x-high pass@1 on DeepSWE is 74% +/- 3%. (https://deepswe.datacurve.ai).

    That said, love seeing some of these new architectures get people exploring. But, surely somebody is incorrect here inre: those numbers.

  3. amluto

    I’m fascinated by this thing and by the way it’s interpreting Jev. It’s very cool, but is it actually a classifier?

    IIUC they took an already-trained “frozen” LLM and trained a little model on top that takes both a question and the hidden states after processing the input data and produces answer “probabilities”. (In contrast, the original LLM would have been run in AR mode to generate multiple output tokens representing its answer.) But then they used it for a purpose that isn’t really classification.

    IMO there is a rather large difference between “is this email spam” and “what character should I type in this agentic workload”. The former is classification: there is hopefully a ground truth (is the email spam?) and the model is trying to classify the email. You would score it with a proper scoring rule. The latter is a strategy: there usually isn’t a correct answer, now or in the future. The model is playing a game consisting of repeated rounds, and the only way to evaluate it is to see how well it plays. You can’t even usefully compare it to the optimal solution because you may not know the optimal solution and you don’t actually need the model to produce an optimal solution.

    I do think this approach is really cool, and it does suggest that one might be able to use a modern LLM to process an input and then extract the model’s next agentic step in a very fast, non-AR manner, with results comparably good to the usual AR decoding. And I think it’s very interesting to de […]

  4. mugul

    Very interesting insight on the training process, it's pretty cool to have some experimental justification for why they took these exact steps, what they tried and did not work, etc. Feels a bit less like dark magic.

    However I agree the latency argument doesn't hold much value with Jev because it runs on a remote server. Seeing how many open Jev-like models came out recently it would be much more interesting to have a comparison with them.

  5. fxwin

    I really hope that "System One" won't stick around as a new buzzword simply meaning "fast".

More from this day

2026-09-24