The emergence of large language models (LLMs) has transformed many aspects of medical decision-making once thought to require human judgment. Proprietary models such as GPT-4 have garnered attention for their high performance in clinical reasoning, including generating differential diagnoses in complex cases. However, the arrival of advanced open-source models, such as Meta's Llama 3.1 with 405 billion parameters, invites scrutiny regarding whether they can compete at the same level. A recent review published in JAMA Health Forum sought to directly compare the performance of these two types of models, focusing on diagnostic accuracy using real-world case records.
Study Design and Evaluation Approach
To investigate the comparative effectiveness of open-source and proprietary LLMs, researchers tested the models on 92 complex diagnostic cases. Seventy of these cases had previously been used to evaluate GPT-4, while the remaining 22 cases were newly selected, having been published after the open-source model's training ended. This design aimed to prevent any undue advantage from prior exposure.
Each model was presented with a summarised clinical case and required to generate a differential diagnosis. Evaluators then assigned a quality score to the models' outputs based on their ability to include the final diagnosis, the precision of their suggestions and their clinical usefulness. Scores ranged from 0, indicating no close suggestions, to 5, indicating that the actual diagnosis was present in the differential. Discrepancies between evaluators were resolved through discussion, and results were analysed using standard statistical methods.
Must Read: Comparing GPT-4’s and Gemini’s Performance in Oncological Reporting
The process adhered to established guidelines for observational studies, and the data analysis did not require oversight from an institutional review board. The models were not permitted internet access during the evaluation, ensuring that responses were generated from internal model knowledge alone. This control enhanced the credibility of the comparison and aligned the assessment with clinical decision-making constraints, where real-time access to external databases may not always be feasible.
Results and Comparative Performance
The open-source model displayed promising performance, rivalling the proprietary GPT-4 in key metrics. On the original 70-case set, the open-source model included the correct final diagnosis in its differential 70% of the time, compared to 64% for GPT-4. Furthermore, its top-ranked diagnosis matched the final diagnosis in 41% of cases, slightly outperforming GPT-4’s 37%. Although these differences were not statistically significant, they underscore the narrowing performance gap between the models.
On the newer 22-case set, published after the open-source model’s training, performance remained strong. The open-source model identified the correct diagnosis in the differential for 73% of cases and named it as the top diagnosis in 45% of them. These results indicate consistent diagnostic reasoning even in unfamiliar cases, suggesting that memorisation did not drive its performance.
The quality of the open-source model's differentials also showed higher inter-reviewer agreement, with consensus reached in 78% of cases, compared to 66% in the prior GPT-4 evaluation. This higher agreement level may reflect the model's ability to generate more coherent or interpretable reasoning paths, aiding evaluators in aligning their assessments.
Visual comparisons of model performance revealed that both models achieved similar distribution patterns across the quality score scale. The highest scores were concentrated in cases where the correct diagnosis was explicitly listed or closely approximated. Yet, neither model demonstrated consistent superiority, and each had instances of both high and low performance.
Implications, Limitations and Future Directions
The demonstration that an open-source LLM can perform comparably to GPT-4 marks a significant development in the AI landscape for healthcare. It suggests that high-quality diagnostic support tools no longer require exclusive access to closed commercial models. Institutions could now consider deploying robust local systems based on open-source frameworks, offering greater control over data privacy, adaptability and cost-efficiency.
However, the study acknowledges several limitations. First, the lack of transparency around LLM training data constrains the interpretation of performance differences. Second, the evaluation scenario—based on structured case summaries—is not fully representative of real-world clinical environments, where data can be fragmented, incomplete or unstructured. Third, the sample size, especially for the newly published cases, limited the statistical power of the analysis. Future work should include evaluations across broader clinical contexts, such as electronic health records, and the assessment of longitudinal model reliability.
Additionally, specific examples illustrated areas of divergence. In one case, the open-source model correctly diagnosed Langerhans cell histiocytosis, which the closed-source model missed. In another, the proprietary model accurately identified neurosyphilis, while the open-source model did not. These cases reflect the models’ varying strengths depending on case characteristics, further reinforcing the need for nuanced evaluation and application.
Despite these limitations, the study makes a compelling case for the viability of open-source models in high-stakes medical applications. Their increasing parity with proprietary systems could accelerate adoption in resource-constrained settings, enhance transparency through open inspection of the model architecture and foster collaborative improvement across the medical AI research community.
As the capabilities of large language models continue to evolve, the line between proprietary and open-source performance in medical diagnostics is becoming less distinct. The study's findings suggest that advanced open-source models like Llama 3.1 can match or even slightly exceed the diagnostic capabilities of proprietary models such as GPT-4 in specific scenarios. While further validation is needed, especially in more varied clinical environments, the implications for scalability, accessibility and innovation are substantial. The availability of high-performing open-source tools may empower healthcare institutions to harness AI more widely, securely and affordably in the pursuit of better patient outcomes.
Source: JAMA Health Forum
Image Credit: Freepik