Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision explores how adding visual context from first-person video can improve the detection of freezing of gait (FOG) in patients with Parkinson’s disease. While traditional wearable sensors track movement, they often struggle to distinguish between pathological freezing and intentional actions like stopping or interacting with objects. This research evaluates whether egocentric video can provide the necessary environmental context to disambiguate these movements.
Addressing the Limitations of Wearable Sensors
Standard FOG detection relies heavily on inertial measurement units (IMUs) worn on the body. A primary challenge with this approach is that IMU data alone cannot always differentiate between a freezing episode and a voluntary stop. Because FOG is often triggered by specific environmental factors—such as narrow doorways, cluttered pathways, or turning—the researchers investigated whether egocentric vision could capture these contextual triggers. By using smart glasses to record the wearer's perspective, the study aims to determine if visual data can complement kinematic sensors to create a more robust detection system.
Methodology and Data Collection
The study involved 13 participants with Parkinson’s disease who experienced daily FOG. Researchers collected synchronized data using smart glasses for egocentric video and five wearable IMUs on the pelvis and lower limbs. Participants performed three types of activities in their homes: walking through doorways, completing daily tasks like washing dishes, and navigating "hotspots" where they frequently experienced freezing.
To analyze the data, the team evaluated "frozen" representations from various foundation models. For IMU data, they used UniMTS and Chronos-2, while for egocentric video, they tested DINOv3, VideoMAE-v2, V-JEPA 2, and EgoVideo. These models were compared against a Temporal Convolutional Network (TCN) trained from scratch on IMU data. All models were evaluated using a leave-one-subject-out cross-validation protocol.
Performance Results
The IMU-based TCN achieved the highest performance, reaching an F1-score of 42.3 and an AUROC of 83.0. In comparison, the V-JEPA 2 ego-video features achieved an F1-score of 32.6 and an AUROC of 77.2. While the egocentric video models did not outperform the IMU-based sensing, they demonstrated above-chance discrimination. Qualitative analysis indicated that egocentric vision captures FOG-relevant information that is independent of the data provided by IMUs, suggesting that visual context has potential value for clinical motion understanding.
Considerations for Future Research
The researchers identified several factors that influence the effectiveness of these models. The dataset showed significant inter-subject variability in the frequency and duration of FOG episodes, which complicates the training of subject-independent models. Additionally, the study noted that while foundation models offer a way to leverage large-scale pretraining for clinical tasks, the performance of these models is still evolving. The findings support the integration of pretrained ego-video representations as a supplementary source of contextual information for wearable-sensor-based systems, rather than a standalone replacement for kinematic sensing.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!