InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Modern intelligent vehicles rely on understanding the driver’s emotional and physical state to improve safety and the overall driving experience. However, existing research tools often focus only on visual data, ignoring the rich information found in human conversation. This paper introduces InCarEmo, a comprehensive, multimodal dataset designed to capture the linguistic, audio, and visual cues that define a driver's state. By integrating RGB and infrared video, audio, and text, the researchers aim to provide a foundation for building more empathetic and responsive in-cabin systems.
A New Multimodal Resource
InCarEmo is built upon scripted, realistic in-cabin scenarios that cover topics like navigation, traffic conditions, and vehicle status. The dataset includes 3,600 Chinese-language emotional video clips and an auxiliary English benchmark to support cross-lingual research. It is specifically designed to support three critical tasks: recognizing the driver’s emotions, detecting fatigue, and monitoring for distractions. By capturing data across diverse lighting and noise conditions, the dataset provides a more realistic environment for training AI models than previous, more limited collections.
The CAMEL Baseline
To demonstrate the utility of the dataset, the authors developed a lightweight baseline system called CAMEL. This system uses a three-stage approach: fine-tuning individual encoders for text, audio, and video; aligning these modalities through contrastive learning to ensure they work together effectively; and using a "multi-classifier voting" scheme to make final predictions. This architecture is designed to be efficient enough for real-time deployment in vehicles while remaining robust enough to handle the noise and low-light challenges typical of an in-cabin environment.
Key Findings and Performance
Experimental results confirm that combining multiple modalities—such as audio and video—consistently leads to better performance than relying on a single source of data. The researchers tested their models under various conditions, including scenarios where certain data streams were missing or obscured by noise. While the models showed strong performance, the study highlights that real-world challenges, such as low-light conditions and environmental noise, remain significant hurdles for current technology.
Practical Considerations
The InCarEmo dataset and the accompanying CAMEL baseline are intended to help developers create safer, more human-centric vehicle interfaces. By providing a unified benchmark, the authors hope to encourage further research into robust, interpretable AI that can accurately monitor a driver's state. The focus on both emotion and safety-critical behaviors like fatigue and distraction makes this a versatile tool for the next generation of intelligent, adaptive automotive systems.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!