Back to AI Research

AI Research

InCarEmo: A Multimodal Dataset for In-Cabin Emotion... | AI Research

Key Takeaways

  • InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring Modern intelligent vehicles rely on understanding the driver’s em...
  • Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction.
  • To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring.
  • The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring.
  • In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation.
Paper AbstractExpand

Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring. InCarEmo integrates RGB and infrared video, in-cabin audio, and dialogue text collected from scripted in-cabin scenarios designed to simulate realistic driver behaviors, covering diverse lighting conditions and driving contexts. The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring. In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation. We provide a unified benchmark with extensive baseline results across unimodal and multimodal methods, including analyses under modality-missing and noise conditions. Experimental results demonstrate the benefits of multimodal fusion and reveal remaining challenges under real-world noise and low-light conditions. By releasing InCarEmo, we aim to establish a comprehensive foundation for robust, interpretable, and human-centric in-cabin affective understanding, promoting safer and more empathetic driver-vehicle interaction.

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Modern intelligent vehicles rely on understanding the driver’s emotional and physical state to improve safety and the overall driving experience. However, existing research tools often focus only on visual data, ignoring the rich information found in human conversation. This paper introduces InCarEmo, a comprehensive, multimodal dataset designed to capture the linguistic, audio, and visual cues that define a driver's state. By integrating RGB and infrared video, audio, and text, the researchers aim to provide a foundation for building more empathetic and responsive in-cabin systems.

A New Multimodal Resource

InCarEmo is built upon scripted, realistic in-cabin scenarios that cover topics like navigation, traffic conditions, and vehicle status. The dataset includes 3,600 Chinese-language emotional video clips and an auxiliary English benchmark to support cross-lingual research. It is specifically designed to support three critical tasks: recognizing the driver’s emotions, detecting fatigue, and monitoring for distractions. By capturing data across diverse lighting and noise conditions, the dataset provides a more realistic environment for training AI models than previous, more limited collections.

The CAMEL Baseline

To demonstrate the utility of the dataset, the authors developed a lightweight baseline system called CAMEL. This system uses a three-stage approach: fine-tuning individual encoders for text, audio, and video; aligning these modalities through contrastive learning to ensure they work together effectively; and using a "multi-classifier voting" scheme to make final predictions. This architecture is designed to be efficient enough for real-time deployment in vehicles while remaining robust enough to handle the noise and low-light challenges typical of an in-cabin environment.

Key Findings and Performance

Experimental results confirm that combining multiple modalities—such as audio and video—consistently leads to better performance than relying on a single source of data. The researchers tested their models under various conditions, including scenarios where certain data streams were missing or obscured by noise. While the models showed strong performance, the study highlights that real-world challenges, such as low-light conditions and environmental noise, remain significant hurdles for current technology.

Practical Considerations

The InCarEmo dataset and the accompanying CAMEL baseline are intended to help developers create safer, more human-centric vehicle interfaces. By providing a unified benchmark, the authors hope to encourage further research into robust, interpretable AI that can accurately monitor a driver's state. The focus on both emotion and safety-critical behaviors like fatigue and distraction makes this a versatile tool for the next generation of intelligent, adaptive automotive systems.

Comments (0)

No comments yet

Be the first to share your thoughts!