
arXiv: 2502.00562
ABSTRACT Large language models (LLMs) such as OpenAI's ChatGPT hold potential for automating engineering analysis, yet their reliability in solving multi‐step statics problems remains uncertain. This study evaluates the performance of GPT‐4o and o1‐preview on foundational statics tasks, from simple calculations to beam and truss analyses and compares their results to first‐year engineering students on a typical statics exam. To enhance accuracy, we developed a Custom GPT, embedding refined prompts directly into its instructions. This optimized model achieved performances scores ranging from 82% to 86% depending on the exam version, surpassing the 75% student average, demonstrating the impact of tailored guidance. Despite these improvements, LLMs continued to exhibit errors in nuanced or open‐ended problems, such as misidentifying tension and compression in truss members. These findings highlight both the promise and current limitations of AI in structural analysis, emphasizing the need for improved reasoning, multimodal capabilities, and targeted training data for future AI‐driven automation in civil and mechanical engineering.
Computational Engineering, Finance, and Science (cs.CE), FOS: Computer and information sciences, Computer Science - Computational Engineering, Finance, and Science
Computational Engineering, Finance, and Science (cs.CE), FOS: Computer and information sciences, Computer Science - Computational Engineering, Finance, and Science
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 2 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
