Unreal Agent's async harness cuts coding agent costs by 40%

Unreal Agent's async harness cuts coding agent costs by 40%

Unreal Labs built Unreal Agent with a fully asynchronous tool-calling harness that removes waits, polls, and heartbeats from the model's workload. The result: up to 40% lower cost than Codex and 20% lower than Pi on real workloads and agentic benchmarks like Terminal-Bench 4.0, SWE-Atlas, and DeepSWE 1.1, with comparable pass rates. The SDK ships as a Go library, a runner executable, and a Harbor-compatible benchmark runner.

The Unreal Agent harness manages tool calls in a completely asynchronous way, relieving the underlying model of the need to manage waits, polls, and heartbeats for tools.
  1. dvt

    I think this space is very untapped. Models are interesting, but I am absolutely obsessed with some things I've been researching/working on for the past few years:

    Fractal tool discovery: tool taxonomy where an agent can "drill deeper" to find what specific tool it's looking for. Helps if/when polluting context with a zillion (mostly unnecessary) tools.

    Leveraging splay trees: this is my favorite data structure and I think relatively unused in the context of agents/harnesses. A lot of times, recently-used workflows/tool-chains will be used again, so having those at the top of the search hierarchy is an awesome optimization.

    Virtual containerized notebooks: models working in sandboxed (WASI) Python notebooks is incredible. Even local models (if given enough time) will usually converge on a good solution. Being able to mount tools/resources/fs is again, imo quite untapped. Some problems here are running native things (thing numpy/pandas) in containers is a nightmare (or impossible).

    Anyway, happy to see other folks seriously doing stuff in this space. If anyone wants to collaborate on anything don't hesitate to reach out :) I'm also actively looking for a job or some contract gigs.

    Fun times ahead.

  2. ricardobeat

    Crush [1] has had async tools for a long time, as has Claude Code.

    They work well, except that the agent will often simply call wait() immediately after, and also do it with a longer timeout, causing crashes/hangs to really slow down the process; whereas immediate tool calls have a built-in 30s timeout in most harnesses.

    I imagine this one will suffer from similar problems.

  3. tekacs

    The headline graph is kind of bizarre.

    For some reason they're comparing their harness running on Astra xhigh to Codex with Astra max?

    ---

    Also worth noting that OpenAI just added support for async tool calling to their harness, which isn't 1:1 with this approach, but is slowly ramping up in being able to provide something similar.

    A big part of why Codex uses so many tokens is that it basically hot loops on polling tasks it starts for... absolutely no good reason: https://www.reddit.com/r/codex/comments/1wdlp7q/weve_discove...

    I fixed it on my fork of Codex too, also back in Jan/Feb – I keep this patch rebased, for anyone who wants it: https://github.com/tekacs/codex/commit/9ffcf8db9078eae43d411...

    It results in token savings similar in scale to those displayed here by Unreal.

    ---

    My harness has used a slightly fancier version of the approach that Unreal is using since ~Feb, and... it definitely works excellently, but it's also assuredly smoother with Astra and other recent models that are more aware of async tool calling.

  4. tapoxi

    Sounds like a trademark issue when Epic ships a wildly popular Unreal Engine

  5. pyrolistical

    Funny, I solved this problem by having Claude write a pi extension

    https://github.com/Pyrolistical/pi-notify

    Now my pi agent setups its own trigger to notify itself when a background process is done

  6. westoque

    not a fan of README's where there's no clear way to run the program, like a getting started or how to build. i think this is the opposite of a docu-monster where AI documents everything, but clearly more documentation is better than lack of.

  7. ozozozd

    I suppose GPT models throw shade when your input is less favorable to Codex:

    “On the surface, Unreal Agent achieves the same outcomes with fewer model turns and fewer input tokens.”

    (Emphasis mine)

  8. tontinton

    Oh very nice, can you also compare it to https://maki.sh?

    Would be interesting to compare to a harness optimizing for cost reduction too.

More from this day

2026-09-22