An expert-led benchmark using common patient questions: evaluating large language models for adolescent idiopathic scoliosis education
 
Yazarlar (5)
Cemre Aydin
Doç. Dr. Aslı Beril KARAKAŞ TANIR Kastamonu Üniversitesi, Türkiye
Anil Murat Ozturk
Figen Govsa
Mehmet Asim Ozer
Makale Türü Açık Erişim Özgün Makale (SSCI, AHCI, SCI, SCI-Exp dergilerinde yayınlanan tam makale)
Dergi Adı EUROPEAN SPINE JOURNAL (Q1)
Dergi ISSN 0940-6719 Dergi Bilgileri (2026)
Dergi Tarandığı Indeksler SCI-Expanded
Makale Dili İngilizce Basım Tarihi 06-2026
Kabul Tarihi 23-05-2026 Yayınlanma Tarihi 10-06-2026
Cilt / Sayı / Sayfa – / 0 / – DOI 10.1007/s00586-026-10046-8
Makale Linki https://doi.org/10.1007/s00586-026-10046-8
UAK Araştırma Alanları
Anatomi
Özet
MethodsA cross-sectional comparative design was used with 100 high-frequency patient questions covering ten clinical domains. Responses generated by both models using standardized zero-shot prompts were independently assessed by expert clinicians: factual accuracy by three raters (two orthopedic spine surgeons and one senior pediatric physiotherapist), and clarity and conceptual coverage by two raters (one surgeon and the physiotherapist). A structured evaluation framework examined three dichotomous dimensions relevant to patient education: factual accuracy, clarity and understandability, and conceptual coverage. Model performances were compared using McNemar’s test, and inter-model agreement was assessed with Krippendorff’s alpha.ResultsBoth models demonstrated equally high factual accuracy (91%). However, clarity was limited, with only one-third of responses rated as sufficiently …
Anahtar Kelimeler
Large language models | Patient education | Adolescent idiopathic scoliosis | Health literacy | Artificial intelligence | Clinical communication