Meta's Muse Spark 1.3 targets agentic coding with higher first-attempt accuracy

Meta's Muse Spark 1.3 targets agentic coding with higher first-attempt accuracy

Meta has released Muse Spark 1.3, a model trained for long-horizon, agentic workflows and competitive coding. It tracks context and prior results, handles messy inputs, and asks for clarification when needed. With native multimodal perception, it can process video, images, and documents, and its visual reasoning runs through a real execution environment. Benchmarks show competitive performance with frontier models on coding evals. Pricing starts at $0.10 per million input tokens for the contributor version, with a 1M context window.

Feed it a screenshot or a clip and let it build.
  1. simonw

    llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"

    https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

    4.2266 cents, 38 seconds.

    For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

    The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.

    UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

    The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.

    And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

  2. superfrank

    I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.

    I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.

    I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.

    Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused mode […]

  3. bertili

    DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap!

    Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

  4. Lucasoato

    A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).

    Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.

  5. jmward01

    muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?

  6. apodolny

    I like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent.

  7. Gecko4072

    Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.

  8. IIIIIllIIII

    Im a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.

  9. 7734128

    Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.

  10. wxw

    “contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol.

    Definitely shows how important a user data flywheel is for RL and model improvement.

More from this day

2026-09-02