
handle: 10576/13149
Scene-to-Speech (STS) is the process of recognizing visual objects in a picture or a video to say aloud a descriptive text that represents the scene. The recent advancement in convolution neural network (CNN), a deep learning feed-forward artificial neural network, enables us to recognize objects on mobile handled devices in real-time. Several applications have been developed to recognize objects in scenes and speak loud their relevant descriptions. However, the Arabic language is not fully supported. In this paper, we propose a bilingual mobile based application that captures video scenes and processes their content to recognize objects in real-time. The mobile application will then speak loud, in English or Arabic language, the description of the captured scene. The mobile application can be extended to further support eLearning technologies and edutainment games. People with visual impairments (VI), such as people with low vision and totally blind people, can benefit from the application to know about their surroundings. We conducted an elementary study about the usage of the mobile application with people with VI and they expressed their interest to use it in their daily lives.
Object Recognition, Assistive Technology, Deep Learning, Scene-To-Speech, Speech Synthesizer
Object Recognition, Assistive Technology, Deep Learning, Scene-To-Speech, Speech Synthesizer
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 7 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
