Meta’s Muse Spark 1.1 targets agentic coding with 1M context

Meta has just rolled out Muse Spark 1.1, bringing a 1M-token context window and a new Meta Model API for agentic, tool-using workflows. Meta’s benchmarks tout gains in computer use and coding, while pricing signals an aggressive push on cost.

meta cover

TL;DR

  • Meta announced Muse Spark 1.1 via Mark Zuckerberg’s X post; available in Meta Model API and Meta AI
  • Focus on agentic performance, tool use, and computer use; trained for desktop, mobile, and browser interfaces
  • 1M token context window for long-running tasks; can delegate work to parallel sub-agents
  • Benchmarks highlight: MCP Atlas 88.1, OSWorld-Verified 80.8, Terminal-Bench 2.1 80.0, CharXiv Reasoning 88.4
  • Additional scores: SWE-Bench Pro 61.5 and BabyVision 76.3; competitors lead some benchmarks
  • Pricing cited in replies: $1.25 input, $0.15 cached input, $4.25 output

Mark Zuckerberg’s X post introduced what Meta is calling “Muse Spark 1.1,” a model the company describes as “strong” at agentic and coding tasks and available through its new Meta Model API as well as in Meta AI.

In the post, Zuckerberg claims Muse Spark 1.1 is strongest in “agentic performance, tool use, and computer use,” and adds that the model handles long-running tasks with a 1M token context window, can delegate work to sub-agents running in parallel, and is trained to use computer interfaces on desktop, mobile, or browser. Meta also portrays the API as a first step toward letting developers build with Muse Spark, alongside a broader push for “strong agentic and multimodal models at very low cost.”

The benchmark table attached to the announcement compares Muse Spark 1.1 with Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), and GPT 5.5 (xhigh) across agentic, coding, and multimodal tests. On Meta’s chart, Muse Spark 1.1 posts its strongest numbers on MCP Atlas for “scaled tool use” at 88.1, OSWorld-Verified for “agentic computer use” at 80.8, Terminal-Bench 2.1 for “agentic terminal coding” at 80.0, and CharXiv Reasoning for “chart QA” at 88.4. It also shows 61.5 on SWE-Bench Pro and 76.3 on BabyVision, while some of the competing models still appear ahead on selected benchmarks.

A user reply highlighted pricing as $1.25 for input, $0.15 for cached input, and $4.25 for output, suggesting Meta is positioning the model aggressively on cost. The announcement also drew a stream of jokes and jabs on X, including remarks about Zuckerberg posting from his @finkd account and Elon Musk replying, “Jinx.”

Source: X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community