Short eye-fixation recordings may contain enough information to distinguish people with Parkinson’s disease from healthy controls using machine learning. The findings, published in the International Journal of Medical Informatics, come from an analysis of eye-tracking data collected during prosaccade and antisaccade tasks. Three time-series classifiers were trained on raw fixation intervals rather than predefined saccade measures. The best-performing approach reached 95% subject-level accuracy on participants who were excluded from training and validation. Model pruning retained only a small fraction of the generated features while slightly improving performance and supporting interpretation of the signal patterns used for classification.
Raw Fixation Signals Replace Predefined Features
The analysis used data from 84 participants, including 54 people with Parkinson’s disease and 30 healthy controls. Patients were non-demented and were at stages 2 or 3 on the Hoehn and Yahr scale. The eye-tracking protocol included horizontal prosaccade tasks, in which participants followed a target with their gaze, and antisaccade tasks, which required them to suppress a reflexive movement and look in the opposite direction.
Rather than analysing the saccades themselves, the models used fixation periods from the preparatory phase before the target moved. Each interval lasted about 1.5 seconds and included horizontal and vertical gaze positions and velocities. This design tested whether brief periods of sustained gaze contained disease-related information without relying on hand-crafted features such as latency, duration, amplitude, reaction time or error rates. After data cleaning and segmentation, the final dataset contained 1,584 fixation intervals.
Must Read: Benchmarking Dataset Tests LLM Symptom Detection
The data were divided into training, validation and test sets that were mutually exclusive at the participant level. This separation was intended to prevent models from recognising individual-specific patterns learned during training. The test set included 24 unseen participants and was structured to age-match the Parkinson’s disease and healthy control groups. Because the dataset contained fewer healthy-control trials, random oversampling was used during training to balance the classes without creating artificial observations. Where possible, patients completed recordings both on medication and after a 12-hour withdrawal.
Three Models Show Different Levels of Accuracy
The models compared were InceptionTime, ROCKET and a pruned ROCKET variant called Detach-ROCKET. InceptionTime applies parallel convolutions at different temporal scales. ROCKET transforms time-series inputs through numerous random convolutional kernels before classification. Detach-ROCKET removes less informative transformed features through sequential feature detachment, producing a smaller model.
Performance was assessed at both trial and participant levels. Trial-level predictions were combined for each person by averaging probability scores and applying a binary threshold. InceptionTime achieved about 78% subject-level accuracy. ROCKET reached about 93%, while Detach-ROCKET performed best at 95%. At trial level, accuracy was lower, at about 56% for InceptionTime and 73% to 74% for the ROCKET approaches. Results were averaged across five runs with different random initialisations.
For Detach-ROCKET, participant-level performance included high precision and recall for identifying Parkinson’s disease. The pruned model retained an average of only around 7% of the original transformed features. Despite this reduction, it performed slightly better than the full ROCKET model at both evaluation levels, although the difference was not statistically significant.
The results also varied modestly by task type. Fixation periods before prosaccades were classified more accurately than those before antisaccades. By contrast, accuracy and confidence did not differ significantly between recordings made while patients were taking dopaminergic medication and those made after withdrawal.
Pruning Supports Interpretation but Validation Remains Limited
The feature-pruning process provided a way to examine which components of the eye-tracking signal contributed to classification. The retained features showed an over-representation of kernels operating at the lowest dilation value. This pattern suggests that high-frequency components of fixation data carried relevant information for distinguishing Parkinson’s disease from healthy controls.
The model’s confidence scores were also compared with patient characteristics. Confidence showed positive correlations with disease duration and motor symptom severity measured by the Unified Parkinson’s Disease Rating Scale Part III in the off-medication condition. No significant relationship was found with age. These associations indicate that the model’s output reflected aspects of disease progression and severity in addition to the binary classification.
Several limitations restrict the current findings. The models were developed from a single dataset with a limited number of participants and relatively uniform recording conditions. Patients were restricted to stages 2 and 3, so the results do not establish performance in earlier or later disease. The small difference between prosaccade and antisaccade preparatory periods was considered indicative rather than conclusive.
Further assessment across other eye-tracking experiments and recording conditions is needed to test generalisability. Longitudinal datasets could also determine whether fixation-based features contain information about subsequent clinical worsening. The present results support continued investigation rather than immediate use as a diagnostic replacement.
Brief fixation recordings from standard saccade experiments enabled machine-learning models to distinguish Parkinson’s disease from healthy controls in unseen participants. The strongest performance came from a pruned ROCKET model that combined high subject-level accuracy with a substantial reduction in transformed features. Its confidence was associated with disease duration and motor symptom severity, but not age or medication condition. The findings support fixation data as a low-burden source for biomarker discovery. Validation in larger and more varied datasets remains necessary before the approach can be assessed across clinical settings.
Source: International Journal of Medical Informatics
Image Credit: iStock
References:
Uribarri G, von Huth SE, Waldthaler J et al. (2026) Deep learning models accurately classify Parkinson’s disease from eye-tracking fixation data. International Journal of Medical Informatics; 220:106626.