Which LLM would you trust with a spine case? A 200-patient head-to-head says it matters which one you pick

A 200-case evaluation of four LLMs on real spine surgery scenarios, scored by five spine surgeons, found significant performance differences between models. Gemini 2.5 Flash outperformed GPT-4o. Model choice is not cosmetic.

The fracture that got lost in translation: LLMs and the limits of classification from text

LLMs can apply OTA/AO classification to fracture radiology reports reliably at the broad level — but subgroup accuracy fails where it matters, and hallucinations occur in documented, specific ways. Here is what that means in practice.