Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ Diagnosticsarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
Diagnostics
Article . 2025 . Peer-reviewed
License: CC BY
Data sources: Crossref
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
Diagnostics
Article . 2025
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
PubMed Central
Article . 2025
License: CC BY
Data sources: PubMed Central
versions View all 3 versions
addClaim

Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC

Authors: Ahmet Yıldırım; Orhan Cicek;

Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC

Abstract

Background/Objectives: The aim of this study was to evaluate the accuracy of aesthetic assessments performed by artificial intelligence (AI)-based large language models (LLMs) using the Aesthetic Component of the Index of Orthodontic Treatment Need (IOTN-AC), which is widely applied to determine the need for orthodontic treatment. Methods: A total of 150 frontal intraoral photographs from patients in the permanent dentition, scored from 1 to 10 on the IOTN-AC, were assessed by two AI-based LLMs (ChatGPT-5 and ChatGPT-5 Pro). Two experienced clinicians independently scored all photographs, with one evaluator’s scores used as the reference (κ = 0.91, ICC = 0.88). Model performance was analyzed by comparing IOTN-AC scores and treatment need classifications. In addition, performance parameters such as accuracy, precision, specificity, and sensitivity were evaluated. Statistical analyses included Spearman correlation, Cohen’s Kappa, ICC, Mean Absolute Error (MAE), Wilcoxon signed-rank test, and Bland–Altman analysis. Results: Both models demonstrated positive and significant correlations with the reference values for scoring and classification (p < 0.001). Compared to GPT-5 Pro, the GPT-5 model exhibited superior performance, with a lower error rate (MAE = 1.47) and higher classification accuracy (66.7%). Bland–Altman analysis showed that most predictions fell within the 99% confidence interval, and regression analysis revealed no systematic bias (p > 0.05). Conversely, the models failed to achieve consistently high performance in each of the performance parameters. Conclusions: The findings revealed that although AI-based LLMs are promising, statistical accuracy alone is insufficient for safe clinical use, and they should demonstrate consistently high performance across all parameters.

Related Organizations
Keywords

Article

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    6
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Top 10%
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Top 10%
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
6
Top 10%
Average
Top 10%
Green
gold