Which LLM would you trust with a spine case? A 200-patient head-to-head says it matters which one you pick

If you have spent any time using ChatGPT in a clinical context, you may have noticed that it often gives an answer that sounds authoritative but falls short in ways that are hard to pin down. The reasoning is plausible. It would pass a viva with a generous examiner. It would not pass one with an exacting one.

A paper published in the European Journal of Radiology in February 2026 tried to quantify exactly that — not whether LLMs can answer spine surgery questions, but how they reason, how they handle uncertainty, and whether the gap between models is clinically meaningful.


The study

Four models were tested: Gemini 2.5 Flash, DeepSeek, GPT-4o, and GPT-4o-mini. Each was deployed within a retrieval-augmented generation (RAG) framework — meaning responses were grounded in an indexed evidence base rather than generated from training data alone. Two hundred real-world spinal cases were used, covering diagnostic, therapeutic, and follow-up reasoning across a range of presentations. Five spine surgeons evaluated outputs using an 11-dimension rubric scored on a 5-point Likert scale, covering accuracy, comprehensiveness, clinical reasoning, evidence citation, humanistic care, and follow-up planning.


What the evaluation found

Gemini 2.5 Flash achieved the highest overall score — 49.25 out of 55 — outperforming all three comparators across most dimensions. It was particularly strong on humanistic care and follow-up planning. DeepSeek led on differential completeness, generating the most comprehensive differentials, but trailed Gemini on overall score. GPT-4o sat in the middle tier. GPT-4o-mini performed worst in core clinical reasoning dimensions despite stable surface-level technical metrics.

That last finding is the most important one. A model that scores well on technical structure while reasoning poorly is exactly the kind of tool that produces confident-sounding output that a non-expert struggles to critique. If you don’t know what good spine surgery reasoning looks like, GPT-4o-mini’s responses would not obviously signal their weakness.


What RAG adds — and what it doesn’t fix

The RAG framework here grounded each model’s responses in indexed source material before generating an answer. This reduces hallucination and anchors outputs to specific evidence. The performance figures in this study reflect ceiling-end deployment for current LLM clinical decision support, not typical use. The same models queried without that grounding would perform considerably worse.

Even within the RAG framework, inter-model variability was substantial. The infrastructure alone does not determine output quality. The model matters.


Caveats worth stating

This study was published in February 2026. The LLM landscape shifts fast enough that model rankings can change within months. What this study demonstrates is the principle — that meaningful performance differences exist between models on real clinical cases scored by practising surgeons — not a permanent hierarchy. A new model or significant update could shift the ordering.

Two hundred cases evaluated by five surgeons is rigorous by LLM research standards. It is still a limited sample for conclusions about diverse real-world spine practice.


What to take from it

If you or your department are considering trialling an LLM for clinical decision support, model selection is not cosmetic. Defaulting to the most prominent name is not a safe assumption about performance. Gemini 2.5 Flash outperformed GPT-4o by a statistically significant margin on real spine cases scored by practising surgeons. That gap should inform procurement conversations and institutional pilots.

The broader signal is that LLMs within a well-designed RAG framework can produce clinically reasonable outputs on complex spinal cases — assessed by surgeons, not automated benchmarks. That is more meaningful than most papers in this space provide. It does not mean these tools are ready for unsupervised clinical use. It means the quality gap between models is large enough to matter when choosing one.

Know which model you are using. Know what it was evaluated on. “I tested it on ChatGPT and it gave a sensible answer” is not the same standard as five spine surgeons systematically scoring 200 real cases.


References

  1. Li Y, et al. Explainable and evidence-linked recommendations for spine surgery via a retrieval-augmented LLM agent. Eur J Radiol. 2026;197:112734. PMID 41722332. https://doi.org/10.1016/j.ejrad.2026.112734

Leave a comment