Time-Series Foundation Models for Cognitive Workload Classification using Eye-Tracking Data
Cognitive workload (CWL) assessment is relevant to a range of applications, such as monitoring driver fatigue and pilot attention, surgeon workload during complex procedures, and astronaut cognitive fatigue during long-duration missions. Eye-tracking datasets are generally small, which hinders the generalizability of the AI models. Time-series foundation models (TSFMs) have shown promise in…
The article explores the use of time-series foundation models (TSFMs) to classify cognitive workload (CWL) using eye-tracking data. CWL assessment is crucial for various applications, including driver fatigue monitoring, pilot attention tracking, surgeon workload measurement during complex procedures, and astronaut cognitive fatigue monitoring during long-duration missions.
However, eye-tracking datasets are typically small, which limits the generalizability of AI models. TSFMs, which are pretrained on extensive data, can be fine-tuned with limited task-specific data, potentially overcoming this limitation.
The study compares two TSFMs, MOMENT and Moirai, against convolutional neural network (CNN), fully connected neural network (FFN), and long short-term memory (LSTM) baselines on two publicly available eye-tracking datasets. Subject-level five-fold cross-validation was employed, where data from each test subject was held out during training. The performance was evaluated using accuracy, area under the curve (AUC), and F1-score, with 95% confidence intervals.
On the class-balanced EM-COGLOAD dataset, pretrained TSFMs demonstrated significant generalization to unseen subjects, with Moirai achieving a remarkable 0.918 AUC and MOMENT reaching 0.882 AUC. In contrast, task-specific baselines remained around 0.70 AUC. In the class-imbalanced COLET dataset, while the performance of all models decreased, MOMENT showed the highest robustness, attaining a 0.670 AUC.
Moreover, both TSFMs exhibited greater stability when evaluated across held-out subjects, suggesting that pretrained representations generalize more consistently across individuals in data-scarce eye-tracking scenarios.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.