Precision medicine is narrowing patient groups by genetic, phenotypic, environmental, treatment-response and outcome-based characteristics, creating smaller datasets for machine learning and deep learning. These rare-disease-sized cohorts, or RDSCs, challenge approaches that depend on large datasets for performance and generalisability. A recent analysis in The Lancet Digital Health considers this tension in digital pathology and other medical domains. As care becomes more tailored, available cohorts can shrink sharply, leaving data infrastructure and algorithmic methods under pressure as medicine moves towards increasingly specific subgroups and, ultimately, cohort sizes approaching a single patient.
Must Read: Data Mesh Reshapes Healthcare Data Strategy
From Broad Cohorts to Narrow Groups
Traditional medical research has relied on large, apparently homogenous cohorts to identify broadly applicable treatments. Precision medicine changes that model by dividing patients into more specific groups according to genetic, phenotypic and environmental characteristics. This stratification supports personalised care but also fragments populations into small RDSCs. Highly restricted pools of eligible participants can emerge in conditions such as paediatric gliomas with particular molecular alterations, metachromatic leukodystrophy associated with ARSA mutations or cancers with NTRK alterations.
Genotype is only one source of narrowing. Clinical course and treatment decisions can further reduce available cohorts. In stage IV melanoma, a broad screened population can become much smaller after applying successive filters for BRAF mutation status, complete response to BRAF and MEK inhibitors and treatment discontinuation. The final group eligible for relapse-risk assessment represents only a very small fraction of the original cohort. As targeted therapies expand and their clinical use evolves, sequential pathways can create very small analytical groups even before combinations of molecular features add further sparsity.
Why Small Datasets Challenge Machine Learning
Machine learning and deep learning models help identify personalised biomarkers and treatment strategies, yet small datasets create major technical limits. With fewer data points, models can overfit by capturing noise or patient-specific features rather than patterns that generalise across a cohort. Performance can then fall when models encounter new, unseen data. Computational pathology illustrates the scale of the problem. Digital pathology images can contain millions of pixels, producing high-dimensional feature spaces in which models can overfit when sample numbers are inadequate. Similar challenges arise in omics domains, particularly where cell-level measurements are used.
Rare disease fields offer a direct warning for precision medicine. Deep learning can detect subtle digital pathology patterns that may escape human observation, but progress in rare diseases has remained limited compared with non-rare diseases because large enough cohorts are difficult to assemble. Diagnostic applications across several modalities have advanced, yet prognosis and therapy-response prediction remain harder. Diagnostic labels are generally clearer, while prognostic and therapeutic outcomes are longitudinal, multifactorial and affected by incomplete or censored information. Early diagnostic success can therefore mask weaker readiness for predictive clinical tasks.
Technology and Data Systems Need Redesign
Small-cohort precision medicine needs approaches that use limited data more efficiently. Zero-shot, one-shot and few-shot learning aim to learn from minimal examples, but their value depends on whether relevant signals are already present and learnable in the training data. Pathology foundation models offer another route by using large pretrained networks to capture morphological patterns in histological images and adapt them to new tasks. Their scale gives them substantial representational power but also creates major compute and data demands, while fine-tuning and transfer learning for RDSCs remain underdeveloped.
Other options include retrieval-augmented generation, federated learning and synthetic data generation. Retrieval-augmented generation can incorporate relevant external data during inference, even after initial training. Federated learning can support cross-site model training without centrally pooling raw data, although privacy, security, administrative approvals and data usage agreements remain barriers. Synthetic data can augment limited datasets, but it cannot fully replace real data and may generate unrealistic outputs. Better algorithms also need stronger infrastructure. Next-generation biobanks require broader biological samples, patient demographics and clinical data, supported by shared standards, interoperable platforms, clinic-facing infrastructure, broad consent, strong privacy measures and policies that enable cross-institutional data sharing.
Precision medicine is moving towards increasingly small datasets as patient-specific and population-specific subgroups become more narrowly defined. This shift creates direct challenges for machine learning and deep learning, especially when models must address prognosis or therapy response rather than diagnosis alone. Foundation models, federated learning, synthetic data generation, transfer learning and retrieval-augmented generation may mitigate some obstacles, but methods designed specifically for small specialised cohorts remain limited. Progress depends on both data-efficient algorithms and stronger data collection, biobanking, standardisation and sharing frameworks. Without these elements, patient-specific models may remain confined to narrow or experimental use.
Source: The Lancet Digital Health
Image Credit: iStock
References:
Janowczyk A, Merkler D, Michielin O & Madabhushi A (2026) Precision medicine’s inevitable trajectory toward raredisease-sized cohorts: implications for machine learning and deep learning. Lancet Digit Health: Online first.