
The transition to AI-powered automation has become a pivotal focus in recent years, with large language models (LLMs) revolutionizing various domains. This study explores the utilization of LLM fine-tuning for non-language downstream tasks in material science and investigates diverse training strategies to enhance performance. We propose a novel approach that incorporates structured and unstructured scientific data from public datasets and literature into open-source models. The Scientific Question Answering Generation (SciQAG) model automates the generation of instructions from scientific texts, efficiently extracting knowledge without relying on manual extraction or domain-specific knowledge graphs. Additionally, we investigate multi-task training strategies that leverage the interdisciplinary nature of materials science, demonstrating superior predictive performance compared to single-task training. Extensive experiments on 23 scientific tasks relevant to materials science, including semiconductors, polymers, metal-organic-framework and so on, show that our LLMs achieve state-of-the-art performance, surpassing existing baselines. By relying on open-source models, our approach promotes transparency and reproducibility in the scientific community. The implications of this research extend beyond materials science, as the methodology can be adapted to other scientific domains. We aim to inspire further research and development in AI for science, enabling researchers to leverage LLMs to tackle complex scientific challenges and drive innovation in materials science and beyond.
Material Science; Large Language Models, Open Access, Deep Learning, AI, Semantics
Material Science; Large Language Models, Open Access, Deep Learning, AI, Semantics
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
