Meta’s Muse Code beats Codex but trails Claude Opus 5 on benchmarks

- Meta shipped Muse Code, a terminal coding agent powered by Muse Spark 1.2.
- Meta’s own benchmarks show it beating OpenAI’s Codex and Google’s Antigravity on most tests but trailing Anthropic’s Claude Opus 5.
- Meta is competing on cost and features like a crash-safe event log.
Meta released Muse Code (beta), a terminal-based coding agent, on Wednesday. The company’s own charts show its model beating OpenAI’s Codex and Google’s Antigravity on most coding tests but losing to Anthropic’s Claude Opus 5 on every benchmark Meta released.
Muse Code trails Opus 5 but adds a crash-safe log
Muse Code is powered by Muse Spark 1.2, an update to the coding model Meta opened up to U.S. developers in July.
Muse Spark 1.2 lags behind Opus 5 on the coding benchmarks Meta presented but outperforms Codex and Antigravity on most of them. Meta said the model was its “next step toward the frontier, with larger and much more capable models on the way.”
The company says that Version 1.2 is better at code generation, debugging, and understanding large codebases after it scaled up compute spent on coding tasks during training.
In one case study, the model rewrote GPU kernels for NVIDIA Hopper chips over more than 1,000 tool calls and for as long as 24 hours, working from a baseline it was told not to copy from existing libraries.
Muse Code maintains a local event log that records all model calls, tool runs, approvals, and edits. “This single source of truth makes the runtime replay-exact and restart-safe,” Meta wrote in its blog post referring to the log.
If a crash occurs, the agent can pick up where it left. The agent also keeps background subagents running for an entire session.
Mark Zuckerberg, Meta CEO, wrote in a post, “When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees.” He continued, “Your working copy is never touched. In testing we had it build six features for a game simultaneously with no collisions.”
Releasing Muse Code in beta today. It’s a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update. pic.twitter.com/xqavk41w6v
— Mark Zuckerberg (@finkd) August 5, 2026
/plan, /grill, and Meta’s bet on cost
Muse Code has several built-in commands. /plan turns a request into an approval-gated plan, /grill stress tests that plan until it sticks, and /goal drives towards the completion of the stated goal.
Developers are able to install the agent on macOS or Linux with a single curl command. Muse Spark 1.2 is also available via the Meta Model API.
The model and the agent work as a pair. The Muse Code toolset keeps them compatible.
Some of the training data was generated by the older Muse Spark 1.1, generating hard coding problems and grading candidate answers. This is a loop Meta says helped the newer model closely follow instructions.
Meta is telling investors it expects to spend between $125 billion and $145 billion this year on chips, data centers, and other infrastructure.
With Opus 5 out of reach on the benchmarks, the company is competing on price instead. A larger Meta AI model, codenamed Watermelon, is still in training.
If you're reading this, you’re already ahead. Stay there with our newsletter.
FAQs
What is Muse Code?
Muse Code is a terminal-based AI coding agent from Meta, released in beta on August 5, 2026, that plans changes, writes code, and validates results across large repositories using the Muse Spark 1.2 model.
How does Muse Code compare to Claude Code and Codex?
On Meta's own benchmark charts, Muse Spark 1.2 beats OpenAI's Codex and Google's Antigravity on most coding tests but trails Anthropic's Claude Opus 5 on every coding benchmark shown.
What makes Muse Code different from other coding agents?
It logs every model call, tool run, and edit, so the agent resumes exactly where it stopped after a crash. It also runs persistent background subagents in parallel.
Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses
Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.
















