What your patient got from ChatGPT before their clinic appointment

Your patient with degenerative cervical myelopathy has read their MRI report. They have done what most patients with a significant new diagnosis now do before clinic: they asked an AI. They arrive with questions — and with some answers they already believe they have.

A study published in Clinical Spine Surgery in April 2026 tested what ChatGPT-5 actually produces when queried about the AO Spine evidence-based recommendations for DCM management — once framed for a clinician, and once framed explicitly for a patient at a sixth-grade reading level.


What the study found

Three fellowship-trained spine surgeons independently rated the outputs across five DCM clinical scenarios covering the range from severe myelopathy to non-myelopathic cord compression.

Clinician-facing outputs were highly accurate. Mean score 2.93 out of 3, no outputs rated inaccurate, near-perfect inter-rater agreement. ChatGPT-5 reproduced the key AO Spine recommendations correctly for each scenario.

Patient-facing outputs were different. Mean accuracy fell to 2.33 out of 3, with a range from 1 to 3 and no full inter-rater agreement. The specific problems identified were clinically important: the mJOA threshold distinguishing surgical from conservative management was frequently omitted; the boundary between surgical and conservative care was not consistently drawn; explanations were sometimes accurate in general tone but missing the precise detail a patient would need to give properly informed consent.

The readability analysis compounded this. Despite explicit instruction to produce output at a sixth-grade reading level — the standard recommended by the AMA and NIH for patient education materials — outputs came back at a mean SMOG grade of 10.4 and Flesch-Kincaid Grade Level of 6.9. Above the recommended threshold. Most patients with DCM are elderly adults, many dealing with cognitive effects of myelopathy itself. Content pitched at seventh-to-tenth-grade reading level is not equivalent to what a patient with moderate cervical myelopathy can reliably process.


Why the mJOA threshold matters

The mJOA scale distinguishes mild, moderate, and severe myelopathy. The threshold at which surgery is recommended is not an optional detail — it is the clinical hinge on which the recommendation turns. An AI output that explains DCM surgery in broadly appropriate terms while omitting this threshold is not a harmless simplification. It is clinically incomplete information that a patient might rely on when deciding whether to proceed with a major cervical operation.

For a patient scoring in the moderate range who has been told by an AI that surgery is generally recommended for DCM without any mention of grading, the gap between what they were told and what they need to know for informed consent is significant.


Caveats

Five scenarios, three surgeons, one AI model, one pilot study. A thin evidence base. ChatGPT-5 is a recent model and likely better than its predecessors at both accuracy and readability. These findings should not be generalised without replication in a larger study and across other LLMs.

But the pattern — clinician-facing accurate, patient-facing systematically worse — is consistent with findings from other LLM evaluations across orthopaedic and non-orthopaedic specialties. That consistency is harder to dismiss as a single-study artefact.


What this changes about the consent conversation

Assume your patients with significant diagnoses are querying AI tools before, during, and after their clinic appointments. The outputs they receive will be broadly accurate in tone and structurally reasonable. They will sometimes omit critical clinical details. They will regularly be pitched above the reading level of the people they are meant to inform.

Building awareness of this into the consultation — asking what the patient has already read, clarifying what might be missing or misframed — is becoming a standard part of responsible practice in high-stakes decisions. Not hostile to AI use by patients. Just realistic about what current tools reliably produce and where the gaps are.

ChatGPT-5 is a good clinician’s assistant for guideline recall. It is not yet a reliable patient educator for major surgical decisions. The difference between those two roles is the one that matters most in clinic.


References

  1. Khan RR, et al. Can large language models translate spine surgery guidelines for patients? A pilot validation study using AO Spine recommendations for degenerative cervical myelopathy. Clin Spine Surg. 2026. PMID 42084953. https://doi.org/10.1097/BSD.0000000000002083

Leave a comment