Artificial intelligence in thoracic surgery consultations: evaluating the concordance between a large language model and expert clinical decisions

dc.contributor.authorDéniz Armengol, Carlos
dc.contributor.authorMarcè, Judith
dc.contributor.authorMacía, Ivan
dc.contributor.authorRivas Doyague, Francisco
dc.contributor.authorMuñoz, Ana
dc.contributor.authorParadela, Marina
dc.contributor.authorGarcía, Sonia
dc.contributor.authorMoreno, Camilo
dc.contributor.authorSerratosa, Inés
dc.contributor.authorGarcía, Marta
dc.contributor.authorRodríguez-Martos, Tania
dc.contributor.authorOjanguren, Amaia
dc.date.accessioned2026-07-17T11:17:57Z
dc.date.available2026-07-17T11:17:57Z
dc.date.issued2025-11-17
dc.date.updated2026-07-17T11:17:57Z
dc.description.abstractBackground: Artificial intelligence (AI) and large language models (LLMs) are increasingly used in clinical workflows, but their real-world application in thoracic surgery decision-making remains underexplored. Methods: This retrospective observational study assessed the concordance between diagnostic and therapeutic recommendations generated by Scholar GPT (based on GPT-4) and decisions made by board-certified thoracic surgeons. All outpatient consultations over one week in a tertiary care hospital were included. Each case was evaluated using a 6-point concordance scale (0–5), developed to quantify agreement in diagnosis and treatment planning. This was a retrospective observational, single-centre analysis; two independent thoracic surgeons assigned the concordance score. We report descriptive statistics and used t-tests/ANOVA for continuous variables and chi-square tests for categorical variables. Given the exploratory design, no a priori sample-size calculation or power analysis was performed. Results: A total of 81 consultations were analysed. The mean concordance score was 3.67 ± 1.17. High concordance (scores 4–5) occurred in 56.8% of cases, particularly in oncological diagnoses such as mediastinal and pleural tumours. Lower concordance was observed in complex or functional conditions like metastatic lung disease and thoracic outlet syndrome. No significant differences were found between consultation modalities or visit types. Conclusion: Scholar GPT demonstrated promising alignment with surgeon decisions in structured oncologic cases but showed variability in complex scenarios. While AI may assist in streamlining outpatient workflows, its use should remain complementary to expert clinical judgment. These findings are exploratory and should be interpreted with caution given the small sample size and single-centre, one-week design.
dc.format.extent7 p.
dc.format.mimetypeapplication/pdf
dc.identifier.idgrec771059
dc.identifier.issn2673-253X
dc.identifier.pmid41333105
dc.identifier.urihttps://hdl.handle.net/2445/230794
dc.language.isoeng
dc.publisherFrontiers Media
dc.relation.isformatofReproducció del document publicat a: https://doi.org/10.3389/fdgth.2025.1633278
dc.relation.ispartofFrontiers in Digital Health, 2025, vol. 7, p. 1-7
dc.relation.urihttps://doi.org/10.3389/fdgth.2025.1633278
dc.rightscc-by (c) Déniz Armengol, Carlos et al., 2025
dc.rights.accessRightsinfo:eu-repo/semantics/openAccess
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.sourceArticles publicats en revistes (Ciències Clíniques)
dc.subject.classificationIntel·ligència artificial en medicina
dc.subject.classificationCirurgia toràcica
dc.subject.classificationPresa de decisions
dc.subject.otherMedical artificial intelligence
dc.subject.otherThoracic surgery
dc.subject.otherDecision making
dc.titleArtificial intelligence in thoracic surgery consultations: evaluating the concordance between a large language model and expert clinical decisions
dc.typeinfo:eu-repo/semantics/article
dc.typeinfo:eu-repo/semantics/publishedVersion

Fitxers

Paquet original

Mostrant 1 - 1 de 1
Carregant...
Miniatura
Nom:
940391.pdf
Mida:
410.94 KB
Format:
Adobe Portable Document Format