Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators Large language models (LLMs) are increasingly used to simulate human behavior in surveys...
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators Large language models (LLMs) are increasingly used to simulate human behavior in surveys...
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports Large language models are increasingly used to help make decisions about the future, b...
Large language models (LLMs) are often used in high-stakes environments where providing an incorrect answer can have serious consequences.
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating Language models with limited memory must consta...
TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs Security Operations Centers (SOCs) frequently use automated tools to map...
SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents addresses the limitations of current weather-forecasting systems,...
A Roadmap to Impactful Pluralistic Alignment Research argues that while the field of pluralistic AI—the effort to make models that reflect diverse human valu...
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation Large Language Models (LLMs) are increasingly used to assist in scientific discovery,...
Agentic Root Cause Analysis through Evidence-Grounded Reasoning Industrial systems are complex, and when they behave abnormally, identifying the specific phy...
TRACE-Router: Task-Consistent and Adaptive Online Routing for Agentic AI Modern enterprise AI often relies on a mix of large, powerful models and smaller, fa...