
arXiv: 2302.04723
handle: 20.500.14243/436389
Context: Requirements engineering researchers have been experimenting with machine learning and deep learning approaches for a range of RE tasks, such as requirements classification, requirements tracing, ambiguity detection, and modelling. However, most of today's ML/DL approaches are based on supervised learning techniques, meaning that they need to be trained using a large amount of task-specific labelled training data. This constraint poses an enormous challenge to RE researchers, as the lack of labelled data makes it difficult for them to fully exploit the benefit of advanced ML/DL technologies. Objective: This paper addresses this problem by showing how a zero-shot learning approach can be used for requirements classification without using any labelled training data. We focus on the classification task because many RE tasks can be framed as classification problems. Method: The ZSL approach used in our study employs contextual word-embeddings and transformer-based language models. We demonstrate this approach through a series of experiments to perform three classification tasks: (1)FR/NFR: classification functional requirements vs non-functional requirements; (2)NFR: identification of NFR classes; (3)Security: classification of security vs non-security requirements. Results: The study shows that the ZSL approach achieves an F1 score of 0.66 for the FR/NFR task. For the NFR task, the approach yields F1~0.72-0.80, considering the most frequent classes. For the Security task, F1~0.66. All of the aforementioned F1 scores are achieved with zero-training efforts. Conclusion: This study demonstrates the potential of ZSL for requirements classification. An important implication is that it is possible to have very little or no training data to perform classification tasks. The proposed approach thus contributes to the solution of the long-standing problem of data shortage in RE.
60 pages, 22 tables, 1 figure
FOS: Computer and information sciences, D.2.1, Requirements Engineering, AI for software engineering, Computer Science - Artificial Intelligence, 68T50, deep learning, Software Engineering, contextual word-embeddings, requirements classification, transfer learning, unsupervised learning, Requirements classification, Software Engineering (cs.SE), Computer Science - Software Engineering, Artificial Intelligence (cs.AI), Empirical studies, language models, requirements engineering, zero-shot text classification, AI for requirements engineering, multi-label classification, Zero-shot learning
FOS: Computer and information sciences, D.2.1, Requirements Engineering, AI for software engineering, Computer Science - Artificial Intelligence, 68T50, deep learning, Software Engineering, contextual word-embeddings, requirements classification, transfer learning, unsupervised learning, Requirements classification, Software Engineering (cs.SE), Computer Science - Software Engineering, Artificial Intelligence (cs.AI), Empirical studies, language models, requirements engineering, zero-shot text classification, AI for requirements engineering, multi-label classification, Zero-shot learning
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 90 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 1% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 1% |
