Back to AI Research

AI Research

RunningTab keeps workspace requirements and read evidence outside an agent's context window

Key Takeaways

  • An environment-maintained record reduces omissions in multi-file deliverables, while its completion checks still depend on requirements and cited content.
  • An agent can open the right document and still omit its contents from a finished report.
  • RunningTab addresses that failure by having the execution environment retain requirements, exact read observations and files listed but never opened.
  • The [RunningTab paper](https://arxiv.org/abs/2610.10444) studies direct workspace interaction: an agent uses terminal tools to inspect raw files and create a deliverable without a prebuilt retrieval index.
  • The additional tab persists alongside the task, outside the model's context window and separate from the workspace files.

An agent can open the right document and still omit its contents from a finished report. RunningTab addresses that failure by having the execution environment retain requirements, exact read observations and files listed but never opened.
The RunningTab paper studies direct workspace interaction: an agent uses terminal tools to inspect raw files and create a deliverable without a prebuilt retrieval index. The additional tab persists alongside the task, outside the model's context window and separate from the workspace files.

The environment records what the agent encountered

The agent begins by stating individual requirements. The environment records each file read as an excerpt with its source path and command, and retains listed-but-unopened paths as candidates. It stores observations as shown rather than having the model summarize them into a new note.
Requirements remain open until the agent cites a read-log entry containing their content or marks them unavailable with a reason. Listing requirements also displays matching excerpts and unopened candidates, ranked with BM25. This connects an obligation to evidence already captured and places to search next.
At the first stop, the environment returns a finish check covering unresolved requirements and unopened files. It gives the agent another opportunity to repair the deliverable. It does not continuously prevent every possible premature finish or certify the whole artifact as correct.

The diagnosis separates discovery from loss after reading

The authors analyze plain direct interaction with GPT-5.4 nano on 100 Workspace-Bench tasks over three runs. They report that, among attempts producing a deliverable, 19.8% of expected numeric values that appeared in the agent's observations were missing from its output.
That finding concerns information the agent had already seen. Better search alone would not address all those omissions, because the value was present in a tool observation before the final report lost it. RunningTab gives the agent a persistent path back to the original excerpt.
Failure labeling uses language-model annotators, with reported agreement against human labels of Cohen's kappa 0.83. The resulting categories describe this analysis and should not be generalized into a universal distribution of office-agent failures.

Benchmark gains depend on the task and metric

The study compares three models across Workspace-Bench, TheAgentCompany and OfficeQA Pro. RunningTab improves over plain interaction and model-maintained tracking baselines in the reported comparison. For example, GPT-5.4 nano's Workspace-Bench rubric pass rate rises from 41.63 to 44.73, while Gemini 3.8 Flash rises from 31.42 to 46.91. These are rubric-level scores, not claims that those percentages of complete business projects succeed.
TheAgentCompany evaluation uses a file-server-only subset, so its findings exclude tasks requiring other services. Workspace-Bench also distinguishes rubric satisfaction from task completion at a 70% threshold. These metric definitions matter when estimating what the tab improves.
The environment captures reads, but the agent still supplies the requirements. A missing or poorly formulated requirement can evade that record. Citation validation checks shared terms in the requirement, path and excerpt; lexical overlap alone cannot establish that a report interpreted a fact correctly. RunningTab offers evidence retention and completion support, with semantic checking of the finished deliverable still necessary.

Comments