Terminal-Universe is a framework designed to solve the scarcity of realistic, executable environments for training terminal-based code agents. While agent trajectories—the recorded history of an agent's actions—are abundant, they are "frozen" demonstrations that cannot be interacted with or verified. Terminal-Universe addresses this by "inverting" the process: it uses the tool-execution history within existing trajectories to reconstruct the original workspaces, turning them into reusable, verifiable environments where agents can be trained to solve new tasks.
Reconstructing Environments from History
The framework reconstructs environments in three distinct stages. First, it performs a deterministic replay of the file operations recorded in a trajectory, restoring files to their state before the agent modified them. Because this initial workspace is often incomplete, a "completion agent" is used to supply missing files and dependencies. Finally, the system filters these workspaces, retaining only those that provide enough context for an agent to successfully perform a task. This approach allows the researchers to generate thousands of task-sufficient environments without needing to build them from scratch or rely on manual curation. The ai agents story also surfaces in Google AI Introduces EnvHarness for Adaptive..., adding another angle.
Scaling Tasks Through Breadth and Depth
Once an environment is recovered, Terminal-Universe scales the utility of that workspace by generating new tasks along two axes. For "breadth," the framework identifies dependency relationships between different environments to create cross-workspace tasks, requiring an agent to navigate and integrate code across multiple projects—a common real-world development scenario. For "depth," it extends single-turn tasks into multi-round sessions. In these sessions, a user agent provides iterative feedback, introduces new requirements, or requests bug fixes based on test failures. This forces the coding agent to handle evolving requirements and learn from its own mistakes. The same ai evaluation question is explored in StarHarness, which adds a research perspective.
Empirical Performance
The researchers applied Terminal-Universe to publicly available terminal agent trajectories, resulting in 37.3k task-sufficient environments. To validate the framework, they used this data to fine-tune the Qwen3.5-27B model. The results showed significant improvements: the model’s performance on the Terminal-Bench 2.1 benchmark increased by 11.9 points for single-round tasks, and its performance on the multi-round EvoCode-Bench v2 improved by 13.8 points. These findings suggest that re-solving tasks in reconstructed environments is a highly effective strategy for training capable code agents compared to simply imitating raw trajectories. The same ai evaluation question is explored in TransMeme, which adds a research perspective. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!