Back to AI Research

AI Research

Malena study tests whether stronger coding agents still need elaborate ML harnesses

Key Takeaways

  • A controlled study finds that a minimal coding-agent session can match elaborate autonomous ML engineering harnesses under matched model, hardware and time budgets.
  • Its cost analysis also exposes a trade-off in long-running context.
  • A study of autonomous machine learning engineering asks how much an orchestration layer adds once a model can inspect files, execute code and debug its own work.
  • In [the Malena paper](https://arxiv.org/abs/2609.40303), the authors compare a minimal, single-session coding agent with more elaborate systems while holding the model backbone, hardware and available time fixed.
  • Their finding is specific to these evaluations: extra harness machinery does not deliver a statistically significant advantage over the capable coding-agent baseline.

A study of autonomous machine learning engineering asks how much an orchestration layer adds once a model can inspect files, execute code and debug its own work. In the Malena paper, the authors compare a minimal, single-session coding agent with more elaborate systems while holding the model backbone, hardware and available time fixed. Their finding is specific to these evaluations: extra harness machinery does not deliver a statistically significant advantage over the capable coding-agent baseline.

Compare the environment before the orchestration

The researchers first separate a chat interface from a coding-agent environment. The chat version produces a solution script as text and receives execution errors through an outer program. The coding agent can explore the working directory, write files and run commands within its own session. Both receive a six-hour limit for the initial comparison.
Giving models direct execution access produces the largest measured improvement in the study. The authors then test search policies that either refine the latest solution, favor the best validation score, balance exploration and exploitation, or start independent attempts. At a matched 24-hour budget, their paired analysis finds no statistically significant advantage for one of those search strategies.
Direct repository access is also part of ML.ai's coding-agent workflow, whose documented editor reads files and runs commands. That shared mechanism explains the connection; this paper does not evaluate ML.ai or establish its performance on the research benchmarks.

Let one session manage the search

Malena runs a single session across the available time, with tools for registering submissions, checking resources and managing background jobs. Hidden test scores are not returned to the agent. The researchers compare it with delegation, three parallel agents and a broadcast channel for peer communication. None significantly improves on the base session in the reported ablation.
The broader comparison includes MLEvolve, AiScientist, Arbor and ScienceFlow. The authors report results across 30 MLE-bench tasks and 40 NatureBench tasks with matched resources. Their trace analysis suggests that Malena reuses checkpoints, preprocessing code and validation splits while adjusting its pipeline, rather than restarting every attempt. That explanation comes from the observed trajectories; it is not a general guarantee that agents will organize arbitrary projects well.

Account for the cost of a growing context

A simpler harness can still have a more expensive inference profile. With GLM 5.2, the paper models Malena's per-run inference cost at $12.12, compared with $1.95 for AiScientist. The authors attribute the difference mainly to repeatedly sending a growing session context, including tokens counted as cache reads. Hardware remains the dominant overall cost in their setup.
The paper also reports a different picture with the smaller Gemma 4 31B backbone, where MLEvolve leads by mean estimate while medal-rate confidence intervals overlap. The practical lesson is to evaluate a simple baseline under the same model and budget before adding orchestration. These benchmark results do not settle the value of multi-agent coordination for every model, task or production workflow.

Comments (0)

No comments yet

Be the first to share your thoughts!