
A living review of automated data extraction methods for clinical systematic reviews identified 76 papers up to the end of 2022 describing unique extraction algorithms and assessed their methods and quality of reporting. The current in-progress update for 2024 observed a rapid increase in papers; it now includes 117 papers, of which 17 employ Large Language Models (LLMs) to automate data extraction from randomized controlled trials. In this commentary, we describe findings from the analysis of these LLM automation papers and discuss parallels to the field of automated data extraction in toxicology. We describe the evaluation strategies and reporting observed in LLM publications and focus on inconsistencies and potential pitfalls that may bias how the performance of LLMs is perceived by those likely to apply them to support systematic reviews. We then discuss potential applications of LLMs in evidence mapping and good practice in the reporting of LLM automation methods – based on a checklist and guideline developed during the 2023 Evidence Synthesis hackathon in Newcastle (UK).
Large Language Model, Automation, Automated data extraction
Large Language Model, Automation, Automated data extraction
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
