Back to AI Research

AI Research

SlackDrive: Reclaiming Runtime Slack for Adaptive D... | AI Research

Key Takeaways

  • SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference Modern autonomous driving models, which combine multimodal reasoning with future predicti...
  • Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of vehicle control.
  • We observe that the largest admissible compute budget varies systematically with the residual runtime state, while recent realized latency provides a direct signal of the available compute slack.
  • Motivated by this observation, we propose \textbf{SlackDrive}, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution.
  • SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference
Paper AbstractExpand

Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of vehicle control. Existing acceleration methods reduce tokens, layers, or sampling steps with policies selected prior to deployment, yet leave residual runtime variation largely unexploited after offline profiling and static scheduling on shared onboard compute. We observe that the largest admissible compute budget varies systematically with the residual runtime state, while recent realized latency provides a direct signal of the available compute slack. Motivated by this observation, we propose \textbf{SlackDrive}, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution. SlackDrive profiles the latency and planning utility of a small discrete budget set once, estimates online compute state from completed forwards, and selects the highest-utility budget predicted to remain within the admissible latency envelope, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator. On NAVSIM v2 with DriveDreamer-Policy, SlackDrive improves latency-constrained EPDMS by $21.7\%$ over the strongest baseline under a stringent latency regime, while the full-budget model and preconfigured token-pruning baselines exceed the admissible latency envelope under runtime contention.

SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference
Modern autonomous driving models, which combine multimodal reasoning with future prediction, are becoming increasingly powerful but also computationally expensive. This creates a conflict with the strict real-time latency requirements of vehicle control. While existing methods attempt to speed up these models by pruning tokens or layers, they typically rely on static configurations chosen before the vehicle is deployed. SlackDrive addresses this by recognizing that the available computing power on a vehicle fluctuates during operation. It introduces a system that monitors the actual latency of completed tasks to dynamically adjust the model's compute budget in real-time, ensuring the vehicle stays within its safety-critical timing constraints. The same computer vision question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.

Adapting to Real-Time Fluctuations

The core insight of SlackDrive is that inference latency is not just an outcome to be measured, but a signal that can be reused to manage future performance. Even if a model is optimized offline, the actual time it takes to process information changes depending on the current workload of the vehicle's onboard computer. SlackDrive acts as a "pre-inference" allocator. Before the model begins a new calculation, the system checks the recent history of completed tasks to estimate the current "compute state." If the system is running smoothly, it allows the model to use a higher-quality, more compute-intensive setting. If the system is under heavy load, it proactively scales back the budget to prevent the model from missing its latency deadline.

How the System Works

SlackDrive operates through a three-step process that remains independent of the specific driving model being used: 1. Budget Profiling: The system performs a one-time analysis of the model to map out different "budget" configurations, recording how much latency each one consumes and how much utility it provides for planning. 2. Compute-State Estimation: As the vehicle drives, the system tracks the latency of each completed forward pass. It uses this data to calculate a normalized "runtime load," which helps predict whether the next step will be fast or slow. 3. Dynamic Allocation: Before each new inference, the system selects the highest-utility configuration that is predicted to finish within the required time limit. This allows the model to be as accurate as possible without violating the strict timing constraints necessary for safe vehicle operation. The same reasoning question is explored in Tracing and Coordinating Cross-Layer Influence for..., which adds a research perspective.

Performance and Flexibility

In testing on the NAVSIM v2 platform using the DriveDreamer-Policy, SlackDrive demonstrated significant improvements over existing methods. Under a stringent latency regime, it achieved a 21.7% improvement in performance compared to the strongest baseline. Crucially, while other methods—such as full-budget models or static token-pruning—often exceeded the allowed latency envelope when the system was under contention, SlackDrive successfully maintained its performance within the required limits. Because it functions as a controller that sits outside the main model, it is "actuator-independent," meaning it can be applied to various types of driving models, whether they rely on token pruning, query-based reasoning, or iterative generation, without needing to retrain the underlying model. The same reasoning question is explored in Partner-Specific Affective Precision in Social Active..., which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!