Name: Artificial Intelligence in Surgical Coding: Evaluating Large Language Models for Current Procedural Terminology Accuracy in Hand Surgery
Keywords: AI in surgery, Current Procedural Terminology, ChatGPT, RD1-811, CPT coding, Hand surgery efficiency, Surgery, Large language models, Original Research

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type 01 Mar 2025 English Publisher:Elsevier BVJournal:Journal of Hand Surgery Global Online, volume 7, pages 181-185 (issn: 2589-5141,

Authors: Emily L. Isch; Jamie Lee; D. Mitchell Self; Abhijeet Sambangi; Theodore E. Habarth-Morales; John Vaile; EJ Caterson;

doi: 10.1016/j.jhsg.2024.11.013

pmid: 40182863

pmc: PMC11963066

Artificial Intelligence in Surgical Coding: Evaluating Large Language Models for Current Procedural Terminology Accuracy in Hand Surgery

- Summary
- Subjects
- Metrics

Abstract

The advent of large language models (LLMs) like ChatGPT has introduced notable advancements in various surgical disciplines. These developments have led to an increased interest in the use of LLMs for Current Procedural Terminology (CPT) coding in surgery. With CPT coding being a complex and time-consuming process, often exacerbated by the scarcity of professional coders, there is a pressing need for innovative solutions to enhance coding efficiency and accuracy.This observational study evaluated the effectiveness of five publicly available large language models-Perplexity.AI, Bard, BingAI, ChatGPT 3.5, and ChatGPT 4.0-in accurately identifying CPT codes for hand surgery procedures. A consistent query format was employed to test each model, ensuring the inclusion of detailed procedure components where necessary. The responses were classified as correct, partially correct, or incorrect based on their alignment with established CPT coding for the specified procedures.In the evaluation of artificial intelligence (AI) model performance on simple procedures, Perplexity.AI achieved the highest number of correct outcomes (15), followed by Bard and Bing AI (14 each). ChatGPT 4 and ChatGPT 3.5 yielded 8 and 7 correct outcomes, respectively. For complex procedures, Perplexity.AI and Bard each had three correct outcomes, whereas ChatGPT models had none. Bing AI had the highest number of partially correct outcomes (5). There were significant associations between AI models and performance outcomes for both simple and complex procedures.This study highlights the feasibility and potential benefits of integrating LLMs into the CPT coding process for hand surgery. The findings advocate for further refinement and training of AI models to improve their accuracy and practicality, suggesting a future where AI-assisted coding could become a standard component of surgical workflows, aligning with the ongoing digital transformation in health care.Observational, IIIb.

Related Organizations

Thomas Jefferson University
United States
Thomas Jefferson University Hospital
United States
Drexel University
United States
Jefferson Hospital for Neuroscience
United States

Keywords

AI in surgery, Current Procedural Terminology, ChatGPT, RD1-811, CPT coding, Hand surgery efficiency, Surgery, Large language models, Original Research

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	3
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Top 10%

Average

Green

gold