LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports Large language models are increasingly used to help make decisions about the future, b...
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports Large language models are increasingly used to help make decisions about the future, b...
Deep learning-based semantic hashing is a powerful technique for organizing high-dimensional data into short binary codes, enabling fast and efficient search...
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating Language models with limited memory must consta...
Efficiency Matters in Autonomous Research AI-driven autonomous research (AR) systems are increasingly used to solve complex problems in science and engineeri...
SceneActBench: Can Agents Act on the 3D Scenes They See? Vision-language model (VLM) agents are increasingly being used to interact with 3D environments rath...
A Roadmap to Impactful Pluralistic Alignment Research argues that while the field of pluralistic AI—the effort to make models that reflect diverse human valu...
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation Large Language Models (LLMs) are increasingly used to assist in scientific discovery,...
The Regression Tax: Decomposing Why Skills Help — and Hurt — LLM Agents This paper investigates a hidden cost in AI development: the tendency for "procedural...
This paper investigates whether current AI agent benchmarks actually measure the capabilities they claim to test.
Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal Reasoning introduces TRACTA, a new synthetic benchmark designed to test how AI syste...