The recent digitization of education is creating new opportunities to provide an enhanced learning experience for different types of students. However, students with visual impairments encounter obstacles in engaging with online classes due to the ina...
The recent digitization of education is creating new opportunities to provide an enhanced learning experience for different types of students. However, students with visual impairments encounter obstacles in engaging with online classes due to the inaccessibility of visual information. At present, a variety of assistive technologies are employed in the context of online education. However, the reality is that the support systems available are not sufficient. To address these issues, this thesis proposes an automatic lecture video commentary system for visually impaired students.
The objective of this research is to enhance the learning experience of visually impaired students by analyzing visual content and generating commentary for lecture videos containing visuals. The system has been constructed as a cross-platform moblie application, with the objective of enabling users to operate the system and listen to lectures through voice commands. Once the user has selected the lecture video they wish to listen to, the system categorizes the various visual materials provided in the lecture video into text, images, tables, and diagrams. It then generates a commentary for each type of contents, converts it to speech, and provides it to visually impaired students. Naver Clova OCR API and Google Cloud Vision API are utilized to effectively analyze text and image information, and table and diagram commentary generation algorithms are developed to clearly understand the relationship between visual materials to generate more accurate commentary.
To evaluate the performance of this system, this thesis compares lecture videos with automatically generated commentary to lecture videos without commentary. The results showed that the videos with commentary showed significant improvements in comprehension and satisfaction compared to the videos without commentary, and the generation speed and readability of the commentary were also positively evaluated.
The automatic commentary system proposed in this thesis focuses on improving the accessibility of material within lecture videos including diagrams. To accomplish this, the thesis introduces a function which is capable of structurally conveying diagram information to the learner. This system recognizes and analyzes diagram images in lecture materials, extracts key information, and explains the information in a voice that can be easily understood by visually impaired students. To see the effectiveness of this system, this thesis conducted a usability evaluation for the most effective explanation method for visually impaired students, and based on the results, we established a diagram explanation method and implemented a diagram analysis algorithm.
The diagram analysis algorithm proposed in this thesis detects arrows in the diagram and analyzes the flow to generate sequentially-proper commentaries. The usability evaluation of the algorithm shows that it outperforms the image captioning feature provided by Google in terms of user satisfaction, appropriateness, and accuracy, indicating that it functions effectively for diagram explanation.
This research is expected to contribute to improving the accessibility of digital educational content by providing a more independent and inclusive learning environment for visually impaired students. It also raises the need for future research and development of automated commentary for other learning materials, such as charts and graphs, which can be used to build more advanced learning support systems.