
doi: 10.1145/3485846
Automatic classification of electronic records is necessary to address the brewing crisis in the recordkeeping discipline, caused by escalating data volumes and digital rights legislation. Current solutions usually employ expert systems that classify records based on their metadata, but this approach is becoming unfeasible due to the increased variety of records and a growing lack of metadata. Text classification is a promising alternative now that the records themselves are machine readable. In this study, the performance of traditional text classification techniques was compared to newer natural language processing technologies in a series of experiments using authentic records data. While the latest Transformer language models showed superior classification skill, traditional methods still perform well. These results were discussed by a focus group of record managers, who believe that text classification can help them manage risk and meet compliance obligations. This is a first step toward aspirations of being able to synthesize narrative from a corpus of records.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 10 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
