SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference
Modern autonomous driving models, which combine multimodal reasoning with future prediction, are becoming increasingly powerful but also computationally expensive. This creates a conflict with the strict real-time latency requirements of vehicle control. While existing methods attempt to speed up these models by pruning tokens or layers, they typically rely on static configurations chosen before the vehicle is deployed. SlackDrive addresses this by recognizing that the available computing power on a vehicle fluctuates during operation. It introduces a system that monitors the actual latency of completed tasks to dynamically adjust the model's compute budget in real-time, ensuring the vehicle stays within its safety-critical timing constraints. The same computer vision question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.
Adapting to Real-Time Fluctuations
The core insight of SlackDrive is that inference latency is not just an outcome to be measured, but a signal that can be reused to manage future performance. Even if a model is optimized offline, the actual time it takes to process information changes depending on the current workload of the vehicle's onboard computer. SlackDrive acts as a "pre-inference" allocator. Before the model begins a new calculation, the system checks the recent history of completed tasks to estimate the current "compute state." If the system is running smoothly, it allows the model to use a higher-quality, more compute-intensive setting. If the system is under heavy load, it proactively scales back the budget to prevent the model from missing its latency deadline.
How the System Works
SlackDrive operates through a three-step process that remains independent of the specific driving model being used: 1. Budget Profiling: The system performs a one-time analysis of the model to map out different "budget" configurations, recording how much latency each one consumes and how much utility it provides for planning. 2. Compute-State Estimation: As the vehicle drives, the system tracks the latency of each completed forward pass. It uses this data to calculate a normalized "runtime load," which helps predict whether the next step will be fast or slow. 3. Dynamic Allocation: Before each new inference, the system selects the highest-utility configuration that is predicted to finish within the required time limit. This allows the model to be as accurate as possible without violating the strict timing constraints necessary for safe vehicle operation. The same reasoning question is explored in Tracing and Coordinating Cross-Layer Influence for..., which adds a research perspective.
Performance and Flexibility
In testing on the NAVSIM v2 platform using the DriveDreamer-Policy, SlackDrive demonstrated significant improvements over existing methods. Under a stringent latency regime, it achieved a 21.7% improvement in performance compared to the strongest baseline. Crucially, while other methods—such as full-budget models or static token-pruning—often exceeded the allowed latency envelope when the system was under contention, SlackDrive successfully maintained its performance within the required limits. Because it functions as a controller that sits outside the main model, it is "actuator-independent," meaning it can be applied to various types of driving models, whether they rely on token pruning, query-based reasoning, or iterative generation, without needing to retrain the underlying model. The same reasoning question is explored in Partner-Specific Affective Precision in Social Active..., which adds a research perspective. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!