Artificial intelligence models for diabetic peripheral neuropathy show promising results in detection, risk prediction and severity classification, but the evidence is not yet strong enough for routine clinical use. A systematic review in BMC Medical Informatics and Decision Making assessed 26 studies published from 2021 to September 2025. The models used imaging, physiological measurements, gait and plantar pressure data, nerve tests and routine clinical records. Age, diabetes duration and glycated haemoglobin often appeared as important predictors. However, many studies used small or single-centre datasets, applied different definitions of neuropathy and lacked external validation, limiting confidence in their reported performance. 

 

Detection Models Focus on Existing Neuropathy 

Diagnostic models aim to identify patients who already have diabetic peripheral neuropathy or distinguish them from those without it. Their inputs usually reflect current nerve damage or its physical effects. Some systems analyse corneal confocal microscopy images, while others use plantar pressure, gait, balance, movement sensors, foot temperature, microcirculation, nerve conduction studies, electromyography or selected clinical variables. 

 

Several imaging and physiological models reported strong results within their own datasets. Corneal imaging systems performed particularly well in internal testing. Pressure, gait and movement-based approaches captured changes linked to altered weight distribution, reduced sensation and abnormal movement. Electrophysiological models used measures such as nerve conduction velocity, signal amplitude and muscle activity, which are closely related to large-fibre nerve dysfunction. 

 

The wide range of inputs reflects the difficulty of finding early disease. Standard screening based on symptoms and bedside examination can be subjective and may miss subclinical neuropathy. Nerve conduction testing is more reliable for large-fibre damage but is resource-intensive and may not detect earlier small-fibre changes. AI models therefore address different parts of the diagnostic pathway rather than offering one universal screening method. Models based on direct physiological or imaging measurements may help identify current neuropathy, but strong internal results do not confirm that they will work equally well in broader patient groups or routine clinical settings. 

 

Risk Models Use Routine Clinical Information 

Prognostic and risk-stratification models use a different type of information. Instead of measuring existing nerve damage, they mainly analyse electronic medical records and routinely collected clinical data to estimate future neuropathy risk or identify patients who may need further assessment. Common inputs include age, diabetes duration, glycated haemoglobin, fasting blood glucose, body mass index, blood pressure, lipid levels, renal function, albuminuria, uric acid and inflammatory or metabolic markers. 

 

Must Read: LLMs Require Cautious Integration in Healthcare 

 

Age, diabetes duration and glycaemic control appeared repeatedly across these models. Kidney-related measures, including serum creatinine and albuminuria, were also common, together with inflammatory markers such as C-reactive protein and the neutrophil-to-lymphocyte ratio. Some models included physical activity, comorbidities, treatment information and cardiovascular risk factors. However, the importance of individual predictors differed between datasets, and several studies did not clearly explain which variables influenced the model most. 

 

These tools are intended for risk identification rather than diagnosis. Their possible role is to flag higher-risk patients, support triage or guide decisions about confirmatory testing and closer follow-up. They should not be treated as substitutes for models designed to detect current nerve damage. Although many studies reported good internal performance, differences in neuropathy definitions and limited external testing remain major concerns. Models developed from hospital inpatients or specialist clinics may also perform differently in community care or primary care populations. 

 

Severity Models Still Need Stronger Evidence 

Severity classification models are designed for patients with established neuropathy. They aim to assign disease stages or identify different types of nerve damage. Their inputs include structured examination findings, symptom scores, vibration perception, monofilament testing, foot appearance, nerve conduction measurements, electromyography, plantar pressure and gait features. Performance was generally strongest when models used electrophysiological data that directly reflected the degree or pattern of nerve dysfunction. 

 

Models based on clinical examination data could often distinguish absent, moderate or severe neuropathy, but mild disease remained more difficult. One system using a structured screening tool misclassified some mild cases, showing that bedside measures may be less sensitive to early or small-fibre damage. Severity models may therefore support staging, specialist assessment or follow-up, but they are not designed for broad screening of asymptomatic patients. 

 

Important methodological weaknesses were common across all three clinical tasks. Nineteen of the 26 studies had a high risk of bias, mainly because of small samples, poor handling of missing data or weaknesses in model analysis. Many datasets came from single hospitals, inpatient groups or specialist clinics. Reference standards also varied, ranging from nerve conduction testing and consensus criteria to symptoms and physical examination. Performance measures were not reported consistently, prospective studies were uncommon and only a few models were tested in external populations. These limitations make it difficult to compare models or judge how well they would perform in routine practice. 

 

Artificial intelligence may support diabetic peripheral neuropathy assessment, but each model must be evaluated according to its intended task. Detection tools mainly use imaging and physiological markers, risk models rely on routine clinical information, and severity systems focus on established nerve dysfunction. Shared predictors such as age, diabetes duration and glycaemic control do not make these models interchangeable. Larger multicentre datasets, consistent outcome definitions, clearer reporting and prospective external validation are needed before routine use. The current evidence supports continued development and testing rather than immediate clinical implementation. 

 

Source: BMC Medical Informatics Decision Making 

Image Credit: iStock 


References:

Ying TJ & Gary AY (2026) Artificial intelligence for detection, prediction, and severity classification of diabetic peripheral neuropathy: a systematic review. BMC Med Inform Decis Mak. https://doi.org/10.1186/s12911-026-03676-x




Latest Articles

AI diabetic neuropathy, diabetic peripheral neuropathy, artificial intelligence healthcare, neuropathy diagnosis, machine learning diabetes, clinical validation, predictive analytics AI models for diabetic neuropathy show promise, but limited validation, bias and small datasets prevent routine clinical adoption today.