Interpretable clinical prediction models remain difficult to build when key information is buried in unstructured clinical notes. Human+Agent Co-design for Healthcare Instruments, known as HACHI, offers a structured way to combine AI-led concept discovery with repeated clinical expert review. A 2026 analysis in npj Digital Medicine applied the framework to traumatic brain injury prediction in children and acute kidney injury prediction after general surgery. In both tasks, the approach helped refine model development, identify clinically relevant concepts and expose issues such as data leakage, imprecise definitions and differences between clinical sites.

 

A More Inspectable Development Process
Clinical prediction models depend on many design decisions, including which data are used, which concepts are included and how performance is evaluated. These decisions become more complex when models draw on unstructured clinical notes, where relevant information may appear in many different forms. HACHI addresses this challenge by assigning lower-level tasks to an AI agent while keeping clinical experts involved in higher-level review.

 

The AI agent searches clinical notes for potentially useful concepts, turns them into yes-no questions, checks whether those concepts appear in notes and tests different combinations against the prediction target. Clinical AI teams then review the results and adjust the process. They can change prompts, refine concept definitions, remove problematic data, adjust model evaluation or alter how groups of patients are weighted.

 

The framework is built around transparency, steerability and reciprocal learning. Transparency allows experts to inspect the concepts and development steps. Steerability allows experts to guide the AI agent through prompts and other changes. Reciprocal learning means that experts can also learn from what the AI agent identifies, then use that insight to shape the next round. This structure keeps the model-building process auditable while limiting claims to the development and evaluation setting.

 

Lessons From Brain Injury Prediction
The traumatic brain injury task focused on children presenting to paediatric emergency departments after head trauma. The aim was to build a small, interpretable model to support risk prediction using clinical notes. Early results appeared strong, but expert review showed that some performance gains came from problematic signals rather than clinically appropriate information.

 

One issue was that the AI agent learned a concept linked to whether notes mentioned Glasgow Coma Scale. That meant the model was partly using documentation style rather than patient condition. Another issue involved brain bleed, which is usually known only after imaging. Further inspection found that some transferred patients already had a diagnosis or imaging results, while some imaging findings appeared in notes despite filtering based on time.

 

Must read: Hybrid Intelligence for Clinical AI

 

These findings led to changes in the data and prompts. The model was redirected towards patient attributes rather than note-writing patterns, and cases with leakage risks were removed. Later rounds also checked whether model coefficients aligned with clinical expectations and whether performance differed between hospital campuses. Equal weighting across the two campuses improved balance in the reported evaluation. The final model used concepts including loss of consciousness, altered mental status, headache, head trauma and normal gait. Comparisons with PECARN and a non-iterative AI-supported baseline favoured HACHI in the reported dataset, while the findings still remain tied to the evaluation context.

 

Refining Acute Kidney Injury Risk
The acute kidney injury task focused on general surgery patients and used preoperative anaesthesia notes to support prediction. The outcome was acute kidney injury within seven days after surgery, using established kidney disease criteria. The model was designed to include a limited number of clear concepts, in line with the goal of interpretability.

 

The first round identified several patient-related factors, including renal impairment and diabetes mellitus. It also detected other features such as haematocrit levels, leukocytosis and cardiac dysfunction. Expert review found that the prompts had narrowed the AI agent’s attention too much. Surgical risk factors were also important, so the next round expanded the search to include both patient and surgery-related concepts.

 

This change increased measured performance and produced factors that were more closely aligned with the clinical task. Further review then raised a practical concern: some concepts were too vague and could be interpreted differently by different clinicians. The final round therefore required more precise concept questions, with examples included in the wording. The AI agent was also encouraged to consider medication-related concepts. The final version performed better than earlier rounds in the reported internal and later-period evaluations. Baseline approaches, including versions of the Kheterpal model and a one-time AI-generated list of predictors, had lower measured performance in the comparisons described.

 

HACHI illustrates a structured approach for combining AI-guided exploration with clinical oversight during prediction model development. The two clinical tasks show how repeated feedback can refine concepts, identify data problems and improve evaluation results within the analysed settings. The framework should not be read as evidence of deployment readiness. Further validation remains necessary before clinical use, including prospective testing, usability assessment and attention to fairness, bias and equity. Its main contribution is an inspectable process for developing interpretable models from clinical notes.

 

Source: npj Digital Medicine

Image Credit: iStock


References:

Feng J, Kothari A, Vossler P et al. (2026) Human-AI co-design for clinical prediction models. npj Digit Med. https://doi.org/10.1038/s41746-026-02838-5




Latest Articles

HACHI, clinical AI, interpretable models, clinical prediction models, unstructured clinical notes, healthcare AI, human AI co-design Explore HACHI human–AI co-design framework for interpretable clinical prediction models from clinical notes with AI concept discovery and expert review