Muse Code and Muse Spark 1.2
Article URL: https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2 Comments URL: https://news.ycombinator.com/item?id=49187575 Points: 266 # Comments: 167
Muse Code, a terminal coding agent, has been released as a beta version, powered by Muse Spark 1.2—the company's newest model. This release represents a significant step forward in the development of larger and more capable models. Muse Code is designed to tackle intricate software engineering tasks across extensive code repositories, including planning modifications, writing code, and validating results.
It can manage multiple persistent subagents for each task, enabling faster, more accurate, and less intervention-heavy problem-solving.
The core of Muse Code's functionality is a simple agent loop enhanced by a series of background agents, which operate asynchronously to augment the main agent's capabilities. These background agents remain active throughout each session, avoiding redundant information gathering and reducing latency by handling subsequent actions and signaling the main agent when necessary. Their persistent nature allows Muse Code to manage complex, multi-step tasks without requiring constant steering.
The runtime of Muse Code is logged in a local event log that records every model call, tool execution, approval, and edit. This comprehensive log serves as a single source of truth, enabling the runtime to replay the exact sequence of events and restart safely after any crashes. This ensures that Muse Code can take on long-running tasks without being interrupted by failures.
Muse Code comes with various default skills, including /plan, which converts a task into an approval-gated plan; /grill, which stress-tests the plan until it meets predefined criteria; and /goal, which focuses on achieving a specified objective. Additionally, Muse Spark 1.2 has been introduced, a coding-focused update to Muse Spark 1.1, featuring enhancements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows.
Muse Spark 1.2 has undergone a substantial increase in training compute for coding tasks and expanded training environment diversity. The model's performance has been bolstered in areas such as general agents while maintaining its strength in coding-related tasks. The training process involved co-training Muse Spark 1.2 with Muse Code, utilizing rejection sampled harness trajectories, recipe optimizations for goals, compaction, and subagents, and integrating the Muse Code toolset to maximize harness compatibility.
The model was extensively tested on long-horizon coding tasks, including whole-repository generation, large end-to-end projects, and auto-research. It employs planning to sequence work, goal conditioning to maintain direction, and context compaction to retain essential knowledge for sustained progress. Muse Spark 1.2 also benefits from training on challenging coding environments generated by Muse Spark 1.1, which then grades candidate solutions based on their adherence to specified requirements, generating a scalable training dataset for Muse Spark 1.2.
To demonstrate the model's capabilities, Muse Spark 1.2 was tested on iteratively optimizing GPU kernels over 1,000+ tool calls, spanning up to 24 hours. By leveraging Muse Code's agentic coding environment, the model was able to write, compile, profile, and progressively improve kernel performance relative to a provided baseline implementation.
Benchmarks were conducted on KDA and MLA kernels for NVIDIA Hopper GPUs, with the model achieving substantial improvements over the baseline implementation, which is the FLA Triton implementation of KDA. The model was prohibited from importing third-party kernel libraries; instead, it had to apply specialized kernel-optimization knowledge to implement the algorithm directly in Triton.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.