
Abstract−In the digital era, learning videos are increasingly being used, however, they often contain irrelevant information, making it difficult to comprehend the content. This study proposes an approach based on the Whisper and T5 models to generate text summaries from YouTube educational video transcripts. Whisper is used for speech-to-text transcription, focusing on model variants that offer a low Word Error Rate (WER) and time efficiency. Subsequently, the T5 model is fine-tuned to produce accurate text summaries, with a strategy of segmenting the transcript to address input length limitations. Text preprocessing is not applied as it resulted in better evaluation quality. The results show that the combination of Whisper Turbo and the optimized T5 model provides the best performance, with F1-Scores on the ROUGE metrics of 39.23 (ROUGE-1), 13.17 (ROUGE-2), and 23.84 (ROUGE-L). This approach successfully generates more relevant and comprehensive text summaries, enhancing the effectiveness of video-based learning. Therefore, this research makes a significant contribution to the development of text summarization technology for learning videos.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 2 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
