Codex on AWS Bedrock bug causing 10x charges

Codex on AWS Bedrock bug causing 10x charges

A GitHub issue reports that Codex CLI 0.147.0 on Amazon Bedrock lacks explicit prompt cache controls for GPT-5.6 Sol, causing high cache-write spend. One user saw cache-write tokens account for 85% of estimated model cost, while another observed a 5x daily cost increase after upgrading, with cache write-to-read ratio jumping from 0.08 to 8.84. The issue requests support for prompt_cache_options and prompt_cache_breakpoint fields.

After upgrading from 0.146.0 to 0.147.0 with the Amazon Bedrock provider, the cache write-to-read ratio increased from 0.08 to 8.84, and my daily cost rose to at least 5x its previous level.
  1. amluto

    Wow, that whole thread is borderline incoherent, presumably generated by an AI without adequate oversight.

    Here are the docs:

    https://developers.openai.com/api/docs/guides/prompt-caching...

    The thread has little explanation as to what weird thing they’re doing to Codex that is making the default work poorly, and it kind of seems like it’s getting confused about whether it wants to set the caching mode or the breakpoint or both.

    In any case, I find the behavior change interesting. It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.

  2. ryanjshaw

    The fact that companies with access to SOTA non-public AI keep having these kinds of dumb bugs in something as simple as a chat app helps validate my disbelief of people claiming AI is ready to replace all software development.

  3. TheP1000

    Our codex on AWS Bedrock read / write cache ratio was less than 5%. Cache writes are very expensive and they were never being used. This results in codex on Bedrock causing ~10x what it should due to no caching and massive writes.

    The workaround in issue resolved for me:

    web_search = "disabled"

  4. prtmnth

    Codex usage feels exorbitantly high since today. They [0] are denying it, but the number of anecdotal users who decided to raise this as an issue (as a result it's trending on X) says otherwise.

    [0] https://x.com/thsottiaux/status/2090675027670978569

  5. spacedoutman

    Something is wrong with the codex app too, burning usage like crazy lately.

More from this day

2026-08-21