OpenForgeRL: Train Harness-native Agents in Any Environment
Modern AI agents are increasingly powerful, but they are often wrapped in complex "inference harnesses"—software scaffolds that manage multi-turn reasoning,...
Modern AI agents are increasingly powerful, but they are often wrapped in complex "inference harnesses"—software scaffolds that manage multi-turn reasoning,...
MIRROR: Learning from the Other View for Multi-Modal Reasoning Vision-language models (VLMs) often struggle with reasoning, even when the information provide...
MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning Electroencephalogram (EEG) signals are complex, containin...
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning In long-horizon reinforcement learning for LLM agents, weak policies often struggl...
The Boundaries of Automation: A Theory of Persistent Human Participation The rapid advancement of AI has led to a common assumption: that humans are only inv...
SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning introduces a new methodology to solve the complex, fragmented, and often slow...
Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning This paper explores why large language models (LLMs) often change their opinio...
Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry As organizations increasingly allow non-engineering staff to build AI ag...
AREX: Towards a Recursively Self-Improving Agent for Deep Research introduces a new approach to AI-driven research.
Logical Regression for Planning with Axioms In automated planning, logical regression is a technique used to determine the most general conditions required f...