
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
웹 환경에서 장애인의 디지털 기기 접근성 향상을 위한 멀티모달 인터페이스 연구
In order to improve the quality of life for people with disabilities, economic self-sufficiency is the most important issue. Participation Rate in Economic Activities of Persons with Disabilities with a university degree of higher is the highest, and the employment rate shows the same phenomenon, so the education and employment rate of the disabled are closely related. This means that going to university with a disability is an important process that increases the possibility of participation in future economic activities. Therefore, there is a need for an environment where people with disabilities can learn for themselves at school and at home. Although learning methods using various assistive devices and smart devices exist for the independent learning of the disabled, there are still inconveniences in using the devices on their own because current technology is very limited and does not embrace different types of disabilities people have. In the existing smart device and web environment, general users use Graphic User Interface (GUI) environment to record, retrieve, and reuse information. In order to control the device and the contents in the existing GUI environment, it is most important for the disabled to control the pointer using the mouse. However, it is difficult to use it, especially depending on the type of disability. Therefore, there is a need to solve a problem that occurs when various types of persons with disabilities control web interaction using a GUI in a web and in a smart device environment. The purpose of this study is to confirm the accessibility of digital devices and web contents by providing differentiated multi-modal interface according to the characteristics of each type of disabilities so that people with different disabilities can access digital devices and First, this study classifies new types of disabilities based on device accessibility features, and presents a specific technique to solve specific problems that occur when performing web interactions. Second, a multi-modal interface for each type of disabilities was developed. This enables users to control contents and menus based on the characteristics of their different disabilities and the types of interaction required. To begin with, the derivation of specific technology for each type of disabilities was fulfilled by analyzing the essential interaction types in the web environment and by newly classifying the disability types in terms of device accessibility. The classified types of disabilities specify the technology users need so that device and content control can be realized according to their disability characteristics. Then, the technology of the accessibility analysis was combined with the derived technology to finally derive the specific technology for each type of disabilities, and the validity was verified through interviews with expert and user groups. As a result, pointer movement and execution interactions were derived as essential interaction types in the web environment. Disability types were newly classified according to hand accessibility and visual perception, and specific technology for each type of disability were presented by applying voice and touch techniques which were derived from the accessibility analysis: ① Blind person, but hand available or partly available: Simple operation, assistive devices, voice input and output ② Low vision, but hand available or partly available: Content expansion and voice output ③ Normal vision , but upper limb impaired: eye tracking and voice command. To continue, problems when controlling content and menu using essential web interaction in digital device environment were discovered through our multimodal interface research. Thus, multimodal interfaces applying specific technology according to the types of disabilities is developed, and its usability and effects have been verified through our research. One of the multimodal interfaces developed in our research is Android-based mobile memo application which enables free voice memo and control for blind people. The application instantly recognizes menus with voice input and applies Bluetooth remote control and voice functions to freely navigate through multi-level menus or folders. Another multimodal interface for persons with low vision is a voice browser in a mobile environment based on Android. The selective focusing technique that selects only the desired area within the web content and expands it in the selected order or outputs it by voice is applied. Likewise, an eye tracking and voice command interface is developed, and usability verification was performed for GUI operation of the persons with upper limb impairment. Eye tracking technology is used to move the pointer, and voice commands technology is used to perform pointer clicks instantly. Furthermore, object enlargement function that automatically expands only the objects that can be clicked on the path of the user's eyes is implemented as well. This study presents a multi-modal interface applying specific technology for each type of disabilities to improve accessibility to digital devices and contents for users in the web environment. This study confirmed that the usability and accessibility were improved to make it easier for persons with different types of disabilities to control digital devices. This study enables independent learning through digital devices and the web for the disabled, and ultimately expects economic stability for the disabled through their economic self-sufficiency. 장애인의 삶의 질 향상을 위해서는 자립을 통한 경제적 보장이 가장 중요한 문제이다. 대졸 이상의 학력에서 경제 활동 참가율이 가장 높으며, 고율율도 같은 현상을 보이고 있어 장애인들의 학력과 고용률은 매우 밀접한 관계를 형성한다. 이는 장애인의 대학 진학이 미래 경제 활동 참가의 가능성을 높이는 중요한 과정이라 할 수 있다. 따라서 장애인이 학교와 가정에서 모두 스스로 학습이 가능한 환경이 필요하다. 장애인의 독립적인 학습을 위해 다양한 보조기기 및 스마트 기기를 사용한 학습 방법이 지원 되고 있지만, 다른 사람의 도움 없이 스스로 기기를 사용하기에는 여전히 불편함이 존재한다. 기존의 스마트 웹 환경에서 일반 사용자들은 정보의 기록, 검색 및 재사용하기 위해 WIMP (Window, Icon, Menu, Pointer)가 적용된 Graphic User Interface(GUI) 환경을 사용한다. GUI 환경에서 장애인이 기기 및 콘텐츠를 제어하기 위해서는 마우스를 이용한 포인터 제어가 가장 중요하지만, 장애 유형에 따라 사용의 어려움이 발생한다. 이를 위해 장애인이 GUI를 사용할 수 있도록 기술 개발되고 있지만, 장애 유형 별 일관된 기술 제공은 다양한 유형의 장애인이 사용하기에는 한계가 존재한다. 따라서 다양한 유형의 장애인이 웹 및 스마트 기기 환경에서 웹 인터랙션을 제어할 때 발생되는 문제점 해결이 필요하다. 본 연구는 다양한 유형의 장애인이 더 쉽게 디지털 기기와 콘텐츠에 접근할 수 있도록 장애 유형 별 특성에 따른 차별화 된 멀티모달 인터페이스를 제공하여 디지털 기기 및 웹 콘텐츠의 접근성 향상을 확인하고자 한다. 첫 번째로, 기기 접근성 기능을 기준으로 새롭게 장애 유형을 분류하고, 장애 유형 별 웹 인터랙션을 수행할 때 발생되는 문제점을 해결할 수 있도록 특성 기술을 제시 한다. 두 번째, 도출된 장애 유형 별 특성 기술과 필수 인터랙션 유형을 기반으로 콘텐츠 및 메뉴를 제어할 수 있는 장애 유형 별 멀티모달 인터페이스를 개발한다. 먼저 장애 유형 별 특성 기술 도출 연구는 웹 환경에서의 필수 인터랙션 유형을 분석하고 기기 접근성 관점에서 장애 유형을 새롭게 분류하였다. 분류 된 장애 유형은 각각의 특징에 따라 기기 및 콘텐츠 제어가 실현될 수 있도록 기술을 도출하였다. 이후 접근성 기술 분석에서 도출된 기술을 적용하여 최종적으로 장애 유형 별 특성 기술을 도출하였으며, 전문가와 사용자 집단의 인터뷰를 통해 타당성을 확인하였다. 웹 환경에서의 필수 인터랙션 유형으로 포인터의 이동과 실행 인터랙션을 도출하고, 장애 유형은 손 접근성 및 시각적 인식에 따라 새롭게 분류되었으며, 필수 인터랙션을 수행 할 수 있도록 접근성 분석에서 파생된 음성 및 터치 기술을 적용하여 각 장애 유형에 대한 특정 기술을 제시하였다: ① 전맹인, 손 사용 가능 및 일부 가능: 간단 조작 가능한 보조기기 및 음성 입출력 ② 저시력인, 손 사용 가능 및 일부 가능 : 콘텐츠 확대 및 음성 출력 ③정상 시력, 상지장애인 : 시선 추적 및 음성 명령. 다음으로 장애 유형 별 멀티모달 인터페이스 연구는 디지털 기기 환경에서 필수 웹 인터랙션을 이용하여 콘텐츠 및 메뉴를 제어할 때의 문제점을 찾고, 도출된 장애 유형 별 특성 기술을 적용한 멀티모달 인터페이스를 개발하고 사용성 및 효과성을 검증하였다. 전맹인을 위해 메뉴를 즉각적으로 인지하고, 다중 단계로 구성된 메뉴 또는 폴더를 자유롭게 탐색 할 수 있도록 블루투스 리모컨과 음성 기능을 적용하여 음성 메모 및 제어가 가능한 모바일 음성 메모 어플리케이션 개발하였다. 스마트폰 환경에서 메뉴 조작에 관한 효율성 및 정확성 비교 실험을 진행하여 정상인이 화면을 보면서 사용한 결과에 비해 전맹인이 리모컨을 사용하면서 음성 메모 것이 시간은 약간 더 소요되지만 정확도는 차이가 없음을 검증하였다. 저시력인을 위해 웹 콘텐츠 내에서 원하는 영역만을 골라 선택한 순서대로 확대 및 음성 출력해주는 선택적 포커싱 인터페이스를 적용한 모바일 보이스 웹 브라우저를 구현하였다. 톡백 서비스와 비교하여 선택적 포커싱 기능의 효과 검증 실험을 진행하여, 보이스브라우저가 원하는 영역만을 빠르고 정확하게 선택하여 읽을 수 있음을 확인하였다. 또한 만족도 평가를 통해 선택적 포커싱 기능이 저시력인의 사용에 효과적이고 사용자의 피로도가 감소함을 확인하였다. 상지장애인의 GUI 조작을 위한 시선 추적 및 음성 명령 인터페이스를 개발하고 사용성 검증을 진행하였다. 시선 추적 기술로 포인터의 움직임을 실행하고, 음성 명령으로 포인터 실행을 즉각적으로 수행 하도록 하였으며, 사용자의 시선이 이동하는 경로에 클릭이 가능한 객체만을 추출하여 자동 확대 시켜주는 웹 확장프로그램(Eye-Voice)을 개발하였다. 포인터 실행의 오작동 감소 실험과 적정 확대 비율 확인 실험을 통해, Eye-Voice가 포인터 실행의 오작동률 감소에 효과가 있었으며, 확대 기능의 효과와 140% 확대 비율이 사용자가 가장 편하고 정확하게 웹 브라우저를 사용할 수 있는 적정 비율임을 확인하였다. 본 논문은 웹 환경에서의 디지털 기기 및 콘텐츠 접근성을 향상시키기 위해 장애 유형 별 특성 기술을 적용한 멀티모달 인터페이스를 제시하였다. 이를 통해 장애 유형 별 더 쉽게 디지털 기기를 제어 할 수 있도록 사용성 및 접근성의 향상되었음을 확인하였으며, 장애인의 디지털 기기와 웹을 통한 독립적인 학습이 가능하여 궁극적으로 장애인의 자립을 통한 경제적 보장을 기대한다.
Efficient Deformable Modeling Network for Multi-View 3D Object Detection
이한림 숙명여자대학교 대학원 2024 국내석사
The 3-dimension (3D) object detection is an important task in autonomous driving and robot vision. Recently, multi-view 3D object detection has become crucial for better understanding the surroundings. In multi-view 3D object detection, there are two types of approaches: a Bird's-Eyes-View (BEV)-based and a sparse query-based approaches. The BEV-based method can aggregate spatial information across multiple images. However, the BEV-based method lost the height information since BEV feature is represented in 2-dimension (2D) space. Additionally, this method takes longer training and inference times than the sparse query-based due to the view-transformation between 2D and 3D space. For these reasons, the sparse query-based approach has gained much attention for its efficiency. This method has been explored to enhance connectivity between 3D sparse queries and surrounded 2-dimension (2D) images. It is important to detect an object that appears across two or more images. In this thesis, a new sparse query-based method is proposed. The method introduces a 4D query designed to clearly define the purpose of an object query to enhance the connectivity between the difference dimensions. The 4D query is defined that includes all target parameters: center, scale, orientation and velocity. However, it may lead unstable training due to its high-dimensional information. To alleviate the complexity of a 4D query, this thesis proposes a training strategy called 4D query denoising. This strategy aims to enhance the training stability and accelerate convergence. Also, the distance-wise feature sampling, a method that considers the relative size of objects based on distance, is introduced. This approach ensures a precise alignment between the position of sparse query and the key from image features. Finally, the proposed network extracts the 2D guidance to generate an initial position of the query by utilizing an auxiliary 2D detector. With promising queries, the query position can be updated easier than existing works. The proposed method demonstrates a remarkable performance improvement on the nuScenes dataset. In comparison to the StreamPETR as the latest state-of-the-art (SOTA) approach, we achieve the increase of 0.9% in mean Average Precision (mAP), 0.4% of nuScenes Detection Score (NDS), 0.6% of mean Average Translation Error (mATE), 0.2% of mean Average Scale Error (mASE), and 1.6% of mean Average Attribute Error (mAAE). In addition, we prove the proposed method has a faster convergence at least 2 times than the StreamPETR. 3차원 객체 탐지는 자율주행이나 로봇 비전에서 중요한 분야이다. 주변환경을 더 잘 이해하기 위해 최근에는 다중 뷰 3차원 객체 탐지의 중요성이 대두되고 있다. 다중 뷰 3D 객체 탐지 모델은 크게 조감도 (BEV) 기반과 희소 쿼리 기반 방식 두 가지로 나눌 수 있다. BEV 기반 방식은 여러 이미지들에 걸친 공간적인 정보를 결합하여 BEV 특징으로 변환한다. 그러나 BEV 특징은 2차원 공간으로 표현되기 때문에, 높이 정보가 소실된다. 또한, 이 방식은 2차원 이미지 특징을 조감도 특징으로 뷰 변환이 이루어져야 하기 때문에 희소 쿼리 기반 방식보다 학습과 추론에 오랜 시간이 걸린다. 한편, 희소 쿼리 방식은 효율성 측면에서 더욱 주목을 받고 있다. 이 방식은 3D 희소 쿼리와 2D 이미지 간의 연결성을 강화하는 방법에 대해 연구되고 있다. 본 연구에서는 새로운 희소 쿼리 방식을 제안한다. 이 방식에서는 다른 차원 간의 연결성을 강화하기 위해 객체 쿼리의 목적을 명확하게 설계한 4차원 쿼리를 정의한다. 4차원 쿼리 목표 파라미터인 중심점, 크기, 방향과 속도 모두를 포함하여 정의된다. 그러나 고차원의 정보를 사용하기 때문에 불안정한 학습을 야기할 수 있다. 이 문제를 해결하기 위해 본 논문은 4차원 쿼리 잡음 제거라고 하는 학습 방식을 제안한다. 이 방식은 학습을 안정화하고 학습 수렴이 빨라지도록 하는 것이 목표이다. 또한, 거리에 따른 특징 추출 기법을 제안한다. 이는 거리에 따른 객체의 상대적인 크기를 고려한다. 이 기법은 희소 쿼리의 위치와 이미지 특징으로부터 얻어지는 키 간의 정렬이 정확할 수 있도록 한다. 마지막으로 제안된 구조는 2차원 탐지 모델을 추가적으로 활용하여 2차원 객체 탐지 결과를 기반으로 4차원 쿼리의 초기 위치를 생성할 수 있도록 한다. 객체가 있을 확률이 높은 쿼리를 활용함으로써 기존 방식들 보다는 더 빠르게 쿼리가 학습될 수 있다. 제안한 기법의 효과를 입증하기 위해 nuScenes 데이터에서 성능 평가 실험을 진행하였다. 기존 네트워크인 StreamPETR 과 비교했을 때, 제안된 네트워크는 mAP 에서 0.9%, NDS 에서 0.4%, mATE 에서 0.6%, mASE 에서 0.2%, mAAE 에서 1.6% 향상된다. 또한, 제안된 네트워크가 StreamPETR 보다 2 배 이상 빠른 학습 수렴 속도를 보임을 입증하였다.
확장 현실 기반 멀티 디스플레이 환경에서 다중 콘텐츠 배치를 위한 저작 시스템
이다솜 숙명여자대학교 대학원 2018 국내석사
멀티 디스플레이 콘텐츠를 배치하는 것은 복잡한 저작 도구를 사용하거나 프로그래밍 과정을 거쳐야하기 때문에 전문가가 아닌 저작자에게는 어려움이 따른다. 선행연구에서는 멀티 디스플레이의 배치를 인식하고 이에 편집 없이 단순 콘텐츠의 전송만이 가능하였다. 또한 확장 현실 개념의 본격화에 따라 전시 환경에서도 가상 디스플레이에 대한 수요가 증가하였지만, 이를 제작하기 위해서는 전문가용 저작 도구를 이용해야 하는 번거로움이 있다. 본 연구에서는 우선 선행 연구를 기반으로 인식된 디스플레이에 효율적인 콘텐츠 배치를 위한 배치 유형을 체계적으로 분류 및 모델링한다. 이후 배치 유형에 따라 콘텐츠를 편집 및 재생하고, 가상 디스플레이를 추가하여 혼합된 디스플레이를 활용한 전시 콘텐츠를 저작할 수 있는 콘텐츠 저작 시스템을 설계한다. 실제 디스플레이의 인식, 콘텐츠 배치 유형 모델링 및 편집, 가상 디스플레이 배치 기능으로 구분하여 설계를 진행하고 이를 바탕으로 저작 인터페이스와 시스템을 구현한다. 구현 시스템을 바탕으로 성능 및 만족도 실험을 진행하여 저작 시스템의 활용성과 효과성 평가를 수행하고 결과를 분석하였다. 본 연구를 통해 첫째, 복잡한 프로그래밍 과정 없이 멀티 디스플레이에 콘텐츠를 배치하여 저작 시간을 단축시킬 수 있다. 둘째, 비용 ․ 물리적 제한이 없는 가상 디스플레이를 쉽게 추가하여 다양한 전시 콘텐츠의 구성이 가능하다. 셋째, 저작자가 전시 공간에서 실시간으로 디스플레이의 콘텐츠를 임의로 수정 및 편집할 수 있으므로 유동적인 전시가 가능하다. 이에 따라, 본 연구에서 제안하는 콘텐츠 저작 시스템을 활용하면 다양한 환경에서 누구나 쉽게 디스플레이를 활용한 콘텐츠 전시를 기획하고 저작할 수 있을 것으로 기대한다. Collocating multi-display content can be difficult for authors, not experts, because they require complex authoring tools or programming. In previous studies, it was possible to transmit simple content without editing and recognizing the arrangement of multi display. In addition, the demand for virtual display has increased in the display environment due to the expansion of the concept of extended reality, but it is troublesome to use a professional authoring tool to produce it. In this study, we first systematically classify and model layout types for efficient content placement on recognized displays based on previous studies. We then design a content authoring system that can edit and play content on the basis of the deployment type, and author virtual content and display content using mixed displays. It is divided into real display recognition, content layout type modeling and editing, virtual display layout function, and the authoring interface and system are implemented based on the design. On the basis of the implementation system, performance and satisfaction tests were conducted to evaluate the utilization and effectiveness of the authoring system, and the results were analyzed. Through this study, it is possible to shorten authoring time by arranging contents on multi display without complicated programming process. Second, cost. Easily add virtual displays without any physical limitations, enabling the configuration of various display contents. Third, because authors can arbitrarily modify and edit the contents of display in real time in the exhibition space, flexible exhibition is possible. Therefore, it is expected that anyone use the content authoring system proposed in this study will be able to plan and author the contents’ exhibition using display easily in various environments.
AV-TFN : 피아노 연주 자세 분류를 위한 오디오-비쥬얼 결합 텐서 네트워크
올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다. 올바른 자세로 악기를 연주할 수 있도록 피아노 연주와 IT기술을 접목한 연구들이 진행되었다. 기존의 피아노 연주 자세 분석 연구는 다음과 같은 한계점이 있다. 첫째, 사람마다 연주 관련 근 골격 계 질환을 일으키는 자세가 다르기 때문에 특정 자세를 부상 자세로 정의하고 분류하는 연구가 부상 방지에 도움이 되기 어렵다는 한계점이 있다. 둘째, 기존 연구들은 연주자의 자세 정보를 얻기 위하여 연주자의 센서를 부착하여 연주자가 연주하는데 방해가 된다. 셋째, 기계 학습을 이용한 자세 분석 연구는 자세를 분류하는 특징을 수작업으로 선택해야 하고 데이터 양이 늘어날 경우 다시 특징을 선택해야 하는 번거로움이 있다. 또한, 피아노 연주 자세는 소리와 관련성이 높은데 기존의 연구는 청각 정보를 고려하지 않고 시각 정보만 활용하였다는 한계점이 있다. 본 논문에서는 시청각 정보를 활용한 딥러닝 기반의 피아노 연주 자세 분류 연구를 진행한다. 첫째, 시각 정보만 활용하거나 기계학습방법을 사용하여 피아노 연주 자세를 분류하였던 과거의 연구들과 달리 제안하는 방법은 연주 자세와 관련성이 깊은 청각 정보를 활용하여 연주 자세를 분석하는 모델을 제안한다 둘째, 본 논문에서는 시청각 정보를 하나의 데이터 구조로 표현하는 방식인 오디오-비쥬얼 결합 방법을 제안한다. 제안하는 오디오-비쥬얼 결합 방법은 이전 관련연구들과 달리 청각 정보를 컬러 스케일로 표현하고 시각정보를 흑백 스케일로 표현하고 이 둘의 색상을 섞어 하나의 데이터 구조에 표현하는 방식으로 하나의 데이터 구조에 시청각 정보를 모두 표현 할 수 있는 데이터 표현 방법이다. 셋째, 제안하는 AV-TFN과 관련 연구인 VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network)과의 성능을 비교하고 제안하는 방법의 우수함을 보였다. 제안하는 AV-TFN은 속도는 유지하면서 F1점수가 평균 약 6.16 퍼센트포인트 향상되는 결과를 보였다. 넷째, 피아노 연주 자세 데이터 셋인 C3Pap (Classic piano performance posture version amateur & pro)를 제안한다. 제안하는 C3Pap는 타 관련 연구와 비교하여 피아니스트 총 명수, 프로페셔널 피아니스트와 아마추어 피아니스트의 비율, 각 클래스 별 프로페셔널 피아니스트의 영상과 아마추어 피아니스트의 영상 비율, 곡의 총 개수, 프로페셔널 피아니스트의 영상의 경우 세계 거장 피아니스트의 영상으로만 수집 등 다양한 방면에서 우수하다. Playing the instrument in the correct position is important because the correct position helps to produce good sounds and prevents injuries. Many studies have been conducted in the field of piano performance posture recognition that combine various piano playing and IT techniques. However, considering that different postures cause different playing-related musculoskeletal disorders (PRMDs), studies that define and recognize specific postures as injured postures are not efficient in preventing injuries. In addition, existing studies have the following limitations: (1) existing studies that use wearable devices interfere with playing due to attachment of sensors in the body of pianist; (2) existing studies that use machine learning has a disadvantage in that a feature for classifying a pose must be manually selected and features must be extracted again as the number of data increases; (3) existing studies have only used visual information for posture classification without considering audio information. In this paper, we propose an audio-visual tensor fusion network (simply, AV-TFN) for piano performance posture classification. Unlike existing studies that used only visual information or classifying piano postures using machine learning methods, the proposed method uses audio information to improve the accuracy in classifying the postures of professional and amateur pianists. For this, we first propose a dataset called C3Pap (Classic piano performance posture version amateur & pro). Unlike existing studies, where it was difficult to collect professional data, this study collects data from various professional and amateur pianists from the YouTube platform. In the collected dataset, the ratio of professional and amateur pianist data is even and has the advantage of having actual performance videos rather than data collected in a specific environment. Further, we propose a data structure that represents audio-visual information. The proposed data structure represents audio information in color scale and visual information on the black and white scale. We call this data structure as an audio-video tensor. Finally, we compare the proposed method with its variants: VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network). The experiment results demonstrate that AV-TFN outperforms existing studies and thus, can be effectively used in the classification of piano performance postures.
고속 철도 MEC 시스템 환경에서 심층 강화학습 기반 태스크 파티셔닝 기법 연구
구설원 숙명여자대학교 대학원 2025 국내박사
고속 철도 시스템(High-Speed Railway, HSR)은 높은 이동성, 독특한 채널 특성, 그리고 이기종 QoS(Quality of Service) 요구사항으로 인해 기존 네트워크 환경과는 다른 복잡한 문제를 야기한다. 특히, 승객들의 QoS를 유지하면서 MEC(Mobile Edge Computing) 자원을 효율적으로 활용하기 위해서는 태스크 파티셔닝(Task Partitioning) 전략이 필수적이다. 그러나 HSR MEC 환경에서는 시간에 따라 변화하는 고속 열차(High-Speed Train, HST)의 개수, 사용자 밀집도, 예측하기 어려운 태스크 발생 빈도, 그리고 MEC 서버의 부하 상태와 같은 동적 요소로 인해 기존의 휴리스틱, 메타휴리스틱과 같은 접근법만으로는 효과적인 문제 해결이 어렵다. 이러한 동적인 환경에서 심층 강화학습(Deep Reinforcement Learning, DRL)은 복잡한 자원 관리 문제를 해결하기 위한 유망한 대안으로 주목받고 있다. 본 연구는 DRL을 기반으로 한 태스크 파티셔닝 기법을 제안한다. 안정적이고 효율적인 학습을 보장하기 위해 MDP의 행동 공간을 줄이는 세 가지 단계로 알고리즘을 설계하였다. 첫 번째 단계에서는 태스크를 오프로딩할 target MEC 서버를 선택하고, 두 번째 단계에서는 본 연구에서 제안하는 선택 확률 기반 방식을 통해 태스크를 처리할 serving MEC 서버들을 선택한다. 마지막으로, 세 번째 단계에서는 DRL 알고리즘을 이용하여 최적의 태스크 파티셔닝 비율을 동적으로 결정한다. 심층 강화학습 알고리즘으로는 TD3(Twin Delayed Deep Deterministic Policy Gradient)를 활용하여 HSR MEC 환경의 특성을 반영한 태스크 파티셔닝 기법을 제안한다. TD3는 Clipped Double Q-learning, Delayed Policy Updates, Target Policy Smoothing Regularization과 같은 주요 특징을 통해 학습 안정성과 성능을 동시에 보장한다. 이를 통해 태스크 처리량을 최대화하고, 태스크 처리 시간과 MEC 서버 간 부하 편차를 최소화하는 것을 목표로 한다. 제안된 기법은 다양한 고속 철도 시나리오를 모델링하여 성능을 평가하였으며, 실험 결과 기존 기법과 비교하여 태스크 처리 성능을 효율적으로 개선하고 QoS를 보장함을 확인하였다. 향후 연구는 본 연구를 확장하여 HST의 전력 자원 최적화 문제를 결합하고, 더욱 효율적인 HSR MEC 시스템 구축을 목표로 할 예정이다. The High-Speed Railway (HSR) system presents complex challenges due to its high mobility, unique channel characteristics, and heterogeneous Quality of Service (QoS) requirements, which differentiate it from conventional network environments. In particular, task partitioning strategies are essential for efficiently utilizing Mobile Edge Computing (MEC) resources while maintaining passenger QoS. However, in the HSR MEC environment, dynamic factors of the environment such as the changing number of High-Speed Trains (HSTs) over time, varying user density, unpredictable task arrival rates, and fluctuating MEC server load states make it difficult to effectively solve these problems using the existing approaches such as traditional heuristic, meta-heuristic. In such dynamic environments, Deep Reinforcement Learning (DRL) has emerged as a promising alternative for addressing complex resource management issues. This study proposes a DRL-based task partitioning method. To ensure stable and efficient learning, the algorithm is designed in three stages to reduce the action space of the Markov Decision Process (MDP). In the first stage, the target MEC server for task offloading is selected. In the second stage, serving MEC servers to process the tasks are chosen based on the selection probability proposed in this study. Finally, in the third stage, a DRL algorithm dynamically determines the optimal task partitioning ratio. The proposed task partitioning method utilizes the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to reflect the characteristics of the HSR MEC environment. TD3 ensures stability and performance in learning through key features such as Clipped Double Q-learning, Delayed Policy Updates, and Target Policy Smoothing Regularization. The method aims to maximize task throughput, minimize task processing delay and the load variance among MEC servers. The performance of the proposed method was evaluated by modeling various HSR scenarios, and experimental results confirmed that it effectively improves task processing performance and ensures QoS compared to existing methods. Future research will extend this study to incorporate HST power resource allocation optimization and aim to build a more efficient HSR MEC system.
한글 획요소 정보를 이용한 CNN 기반 유사 폰트 추천 기법
전자연 숙명여자대학교 대학원 2021 국내석사
With the rise of new media, various forms of communication have emerged, and as a result, online communication has diversified. In order to improve the efficiency and competitiveness of their contents, people’s interest in typeface design is rapidly increasing, and the number of fonts and the designs are diversifying. Although the influence of typefaces is rapidly increasing, the Hangul font system stores font information only by font names or font manufacturing company names. There are limits when trying to identify the fonts with just the shape of the characters. If the font used is not saved in the computer, the user needs to replace the font with a font similar to the one used in the file, so finding the font can be difficult with the current system since checking all the fonts one by one is time consuming. Also, if the users design a font, they need to find if there are existing ones that are similar to the one they designed which is an exhausting job currently. It is practically impossible to identify and use thousands of Hangul fonts. Therefore, the influence and importance of typefaces has increased rapidly, but the lack of the Hangul font system has caused inconvenience for both users of fonts and designers who create fonts, making it difficult to use them efficiently. Recently, an increasing number of attempts have been made to solve the typeface-related problem using deep learning technology, but Hangul is a combination of characters, and unlike Roman characters, its structure is complicated, making it difficult to analyze the characters. The structure of Hangul consists of Stroke Element, Skeleton, and Spacing. Therefore, we believe that analyzing fonts according to the characteristics of Hangul will produce better results. In this study, we aim to analyze the characteristics of Hangul fonts based on the information from Stroke Elements of Hangul fonts which is different from the previous attempts to analyze by looking at the font as a whole. Furthermore, we propose a method to automatically recommend similar fonts based on the information. 뉴미디어가 등장하면서 다양한 형태의 커뮤니케이션이 등장하였고 그로 인해 온라인 커뮤니케이션도 다각화 되었다. 따라서 콘텐츠의 효율성과 경쟁력을 높이기 위해서 글꼴 디자인에 대한 관심도가 급격하게 증가하였으며 폰트의 수와 디자인의 형태 또한 다양해지고 있다. 한글 폰트 시스템은 단순히 폰트명 혹은 폰트 제작사명으로만 글꼴 정보를 저장한다. 따라서 일부 글자의 모양만을 보고 해당 폰트를 식별하고자 할 경우나 파일 내에 사용된 폰트가 시스템 PC에 설치되어 있지 않아 이를 대체할 비슷한 모양의 폰트를 찾는 경우 혹은 제작된 폰트의 디자인이 기존의 폰트와 유사한 디자인인지 판단해야하는 경우 등의 여러 문제가 발생하였을 때 이들을 일일이 확인하며 수동적으로 이용해야만 하는 한계가 존재한다. 수천 종이 넘는 한글 폰트를 모두 확인하고 사용하기에는 현실적으로 불가능하다. 따라서 글꼴의 영향력과 중요성은 급증되었으나 한글 폰트 시스템의 미비로 글꼴을 이용하는 사용자와 글꼴을 제작하는 디자이너 모두 불편함을 느끼며 더 나아가 한글 폰트를 효율적으로 사용하는 데 어려움을 겪고 있다. 최근 딥러닝 기술을 이용하여 이러한 글꼴 관련 문제를 해결하려는 시도가 증가하고 있으나 한글은 조합형 글자로 로마자와 달리 구조가 복잡하여 글자를 분석하는데 비교적 어려움이 많이 발생한다. 한글의 구조는 모양을 나타내는 획요소(Stroke Element)와 골격 정보(Skeleton) 그리고 공간 정보(Spacing)로 이루어져 있다. 따라서 한글의 특성에 맞게 글꼴을 분석한다면 더 나은 결과가 도출될 것이라 판단하여 본 연구에서는 글자 전체를 기반으로 하는 기존의 방법 대신 한글의 획요소 정보를 기반으로 글꼴의 특성을 분석하고자 하며 더 나아가 해당 정보를 통해 유사한 폰트를 자동으로 추천하는 기법을 제안하고자 한다. 첫 번째로, 한글 글꼴의 모양을 주제로 연구된 논문들을 기반으로 한글 글꼴의 모양에 대한 속성을 수집하고 전문가를 대상 으로 중요도 평가를 진행하여 글꼴을 분석하는데 사용할 대표 획요소 특성을 선정하였다. 두 번째로, 선정한 대표 획요소 특성으로 글꼴 간 유사도를 분석하기 위해 글자에서 획요소 정보를 자동으로 검출하는 시스템을 딥러닝 학습을 통해 구현하였다. 딥러닝 객체 모델 Faster R-CNN Inception V2를 파인 튜닝(Fine-Tuning)하였으며 mAP 기반 검증 결과 99.99%의 검출 정확도가 도출되었다. 추가로 학습하지 않은 글자와 글꼴에서 획요소 검출을 진행했을 때도 평균 91.4%의 정확도가 도출되었다. 마지막으로 선정한 대표 획요소 특성을 모두 조합(Combination)하여 각각의 조합으로 도출된 유사 폰트 추천 결과와 사람이 평가한 Ground-Truth를 비교하였다. 글꼴 간의 유사도는 CNN 모델을 기반으로 이미지 임베딩을 진행하고 Cosine Similarity를 통해 분석하였다. 그 결과, 획요소 특성의 조합 개수는 4-5개정도 조합하였을 때 가장 최적의 추천 결과를 보였으며 ‘상투’, ‘꼭지점’, ‘가지’의 획요소가 필수적으로 포함되어야 한다는 것을 확인할 수 있었다. 최적의 획요소 조합과 CNN 모델로 추천하였을 때 0.955의 코사인 유사도 값을 얻었으며, 기존의 방식보다 매우 정확한 결과가 도출된다는 것을 통계 분석을 통해 확인하였다. 따라서 본 연구에서는 한글의 획요소 정보를 이용해 글꼴을 분석하여 CNN 기반 유사 폰트 추천 기법을 제안하였고 유의미한 획요소 정보의 조합을 도출하였다. 이를 통해 한글 글꼴 자동화가 가능해질 것이라 판단하며 더 발전하여 손글씨 등의 비정형범위까지 적용하게 된다면 필적 감정 등의 다양한 분야에서 응용되어 사용할 수 있을 것이다.
Chae-Lin Kim 숙명여자대학교 대학원 2023 국내석사
Facial expression recognition (FER) is one of the essential tasks in both computer vision and human-computer interaction (HCI) fields. It has been widely used in applications such as autonomous driving, robotics, and e-learning enhancement by recognizing emotion through facial expressions. Though its practicality, Convolution Neural Network (CNN) -based FER have fallen into the overfitting problem due to the few numbers of samples available in the FER dataset. To address this issue, we propose to a few-shot learning (FSL) method for FER. FSL is a training mechanism that can predict new categories of samples with only a few data. It learns the relation between data by similarity learning and inference test data by way of learning. In this way, FSL can help to solve the overfitting problem in FER. This thesis proposes a method using the relationNet, which learns relation similarity among datasets. Based on the relationNet, we design a channel selection module and additional spatial data construction. To effectively exploits the best from a few datasets, we make a representative feature as an averaged feature of sample features. Then this representative feature of each channel is compared with each channel information of sample features to find which sample channel feature is the most similar channel information. By comparing channel information, the channel from a selected sample is extracted as an optimal channel of the corresponding sample feature. Therefore, one reconstructed feature is composed of each sample's channel information by the designed module. Focusing on fine-grained features, we figure out that facial expressions have significant information on eyes and lip area. We generate eyes and lip image patches and set this additional data as support and query sets. We prove that the selected optimal feature and additional spatial information can improve the generalization performance. Comparing to the existing method, the average performances on RAFDB, FER2013, SFEW, and AFEW datasets are increased by 3.5%, 3.68%, 5.58%, and 2.31% of accuracy, respectively. 얼굴 감정인식 (FER) 은 컴퓨터 비전 및 인간-컴퓨터 상호작용 (HCI) 분야에서 중요한 작업 중 하나이다. FER은 얼굴 표정을 통해 감정을 인식하면서 자율 주행, 로봇공학 및 e-러닝 향상과 같은 응용프로그램들에서 널리 사용되고 있다. 하지만 그 실용성에도 불구하고, 기존의 컨볼루션 신경망 (CNN) 기반 FER은 FER 데이터셋에서 제한된 수의 샘플로 인해 과적합 문제에 직면하고 있다. 이 문제를 해결하기 위해 우리는 FER에 대한 퓨샷러닝 (FSL) 방법을 제안했다. FSL은 단 몇 개의 데이터만으로도 새로운 범주의 샘플을 예측할 수 있는 학습 메커니즘이다. FSL은 유사도 학습을 통해 데이터 간의 관계를 학습하고 테스트 데이터 또한 학습을 하는 방식과 같이 추론하며 작동한다. 이렇게 함으로써, FSL은 FER의 과적합 문제 해결에 도움을 줄 수 있다. 본 연구에서는 데이터셋 간의 관계 유사성을 학습하는 relationNet을 사용하는 방법을 제안한다. RelationNet을 기반으로 채널 선택 모듈과 추가적인 공간 데이터 구성을 설계했다. 몇 개의 데이터셋으로부터 최적의 정보를 효과적으로 활용하기 위해 샘플 피쳐들의 평균 피쳐를 대표 피쳐로 선정하였다. 다음으로는 각 채널의 대표 피쳐를 샘플 피쳐의 각 채널 정보와 비교하여 어떤 샘플의 채널 피쳐가 가장 유사한 채널 정보를 갖는지 찾게된다. 채널 정보를 비교함으로써 선택된 샘플에서의 채널이 해당 샘플 피쳐의 최적의 채널로 간주되어 추출된다. 따라서 설계된 모듈에 의해 하나의 재구성된 피쳐는 각 샘플의 채널 정보들로 구성되었다. 또한, 세밀하다는 특징에 중점을 두어, 얼굴 표정이 눈과 입술 영역에서 중요한 정보를 가지고 있다는 것을 알아내었다. 따라서 눈과 입술 이미지 패치를 생성하고 이 추가 데이터를 서포트와 쿼리 셋으로 설정했다. 선택된 최적의 피쳐와 추가적인 공간 정보가 일반화 성능을 향상시킬 수 있다는 것을 입증했다. 기존 방법과 비교하였을 때, RAFDB, FER2013, SFEW 및 AFEW 데이터셋에서의 평균 성능은 각각 3.5%, 3.68%, 5.58%, 2.31%의 정확도가 향상된다.
콘텐츠 감정과 글꼴 감정 키워드 매칭을 통한 글꼴 추천 시스템
지영서 숙명여자대학교 대학원 2023 국내석사
디지털 콘텐츠의 양이 급격히 증가하고 콘텐츠에 다채로운 글꼴이 적용되고 있다. 콘텐츠와 어울리는 글꼴을 사용하는 것은 디자인, 가독성, 의미 전달 효과 모두에 영향을 끼치지만, 일반 사용자는 글꼴을 선택하는 과정에 어려움을 느끼고 잘못된 글꼴을 선택하는 경우가 존재한다. 따라서, 본 연구에서는 입력한 콘텐츠의 감정에 따라 어울리는 글꼴을 추천해주는 시스템을 설계하고 구현했다. 콘텐츠 감정에 어울리는 글꼴을 추천하기 위해 글꼴의 감정을 분석하고자 했고 키워드를 통해 글꼴을 분석했다. 한글 글꼴을 잘 나타내고 효과적으로 분석 가능한 키워드를 선정해 글꼴 사이의 연관성을 평가하는 실험을 진행했다. 기존의 평가 인터페이스는 글꼴 간의 비교가 불가능하고 작업량이 너무 많다는 문제가 있었고 이를 해결하기 위해 본 평가에 효과적인 인터페이스를 설계해 평가했다. 키워드에 따라 관련성이 높을 글꼴을 상, 중, 하 세 단계로 분류할 수 있고 모든 글꼴을 한 화면에 표현해 상호 비교를 통해 단계별 범주화를 할 수 있도록 구현했다. 구현한 인터페이스를 이용해 구한 글꼴의 키워드 응답자 비율에 가중치를 곱해 하나의 키워드 속성 값으로 변환했다. 변환한 속성 값의 효용성을 검증하기 위해 키워드 선택 기반 글꼴 추천 프로토타입 시스템을 설계해 사용성 평가를 진행했다. 프로토타입 사용성 평가 결과 기존 글꼴 검색 시스템과 비교해 높은 만족도를 보였으나, 모호한 상황에서 원하는 키워드를 선택하는 것에 어려움을 느끼는 사용자가 존재했다. 최종적으로 콘텐츠의 감정에 따라 어울리는 글꼴 추천 시스템을 구현했다. 콘텐츠의 감정과 글꼴 키워드는 서로 다른 분류 기준이기 때문에 비교를 위해서는 서로 다른 감정 분류 기준을 하나의 공간으로 이동해 비교하는 매핑 모델이 필요했고 모든 감정을 하나의 공간에 표현 가능한PAD(pleasure, arousal, dominant) 모형을 활용해 콘텐츠와 글꼴을 PAD 값을 변환해 비교하는 매핑 방식과 분류 기준 사이의 상관 관계 분석을 통해 비교하는 매핑 방식 두 가지를 설계해 비교했다. 콘텐츠와 각 계산 모델을 적용한 글꼴 추천 리스트를 제시하고 더 적절하다고 판단되는 모델을 평가받았고 상관 관계 분석을 통한 매핑 방식이 효과적이라는 것을 확인했다. 상관 관계 분석을 통한 매핑 방식을 적용한 글꼴 추천 인터페이스를 구현하여 사용성 평가를 진행하여 시스템의 효용성을 확인했다. 본 연구는 콘텐츠 감정에 따라 어울리는 글꼴을 추천해주는 시스템을 제안하였고, 이를 통해 디자인적 경험이 부족한 사용자도 상황에 어울리는 글꼴을 빠르고 정확하게 추천받을 수 있도록 했다. 또한 본 연구에서 제안한 방식을 적용해 영상 제작, 전자 출판, 소셜 네트워크 등 다양한 디지털 콘텐츠 분야에서 활용할 수 있을 것이다. The amount of digital content is rapidly increasing, and colorful fonts are being applied to content. Using a font that matches the content affects all of the design, readability, and message conveying effect, but general users find it difficult to select the font and sometimes choose the wrong font. Therefore, in this study, we designed and implemented a system that recommends suitable fonts according to the emotions of the input content. In order to recommend a font that matches the emotion of the content, the emotion of the font was analyzed, and the font was analyzed through keywords. An experiment was conducted to evaluate the relationship between fonts by selecting keywords that represent Korean fonts well and can be effectively analyzed. Existing evaluation interfaces had problems that comparison between fonts was impossible and the amount of work was too much. To solve this problem, an effective interface for this evaluation was designed and evaluated. Depending on the keyword, fonts with high relevance can be classified into three levels, upper, middle, and lower, and all fonts are displayed on one screen to enable step-by-step categorization through mutual comparison. Using the implemented interface, the percentage of keyword respondents in the obtained font was multiplied by the weight and converted into a single keyword attribute value. In order to verify the effectiveness of the converted attribute values, we designed a font recommendation prototype system based on keyword selection and conducted usability evaluation. As a result of the usability test of the prototype, users showed high satisfaction compared to the existing font search system, but there were users who found it difficult to select the desired keyword in an ambiguous situation. Finally, we implemented a font recommendation system that matches the emotion of the content. Since emotion in contents and font keywords are different classification criteria, a mapping model that compares different emotion classification criteria in one space was needed for comparison. There are two mapping methods: mapping method that compares content and fonts by converting PAD values, and a mapping method that compares content and fonts through correlation analysis between classification criteria, using a PAD (pleasure, arousal, dominant) model that can express all emotions in one space. A recommended list of fonts applying each calculation model was presented, and a model judged to be more appropriate was evaluated, and it was confirmed that the mapping method through correlation analysis was effective. By implementing a font recommendation interface to which a mapping method through correlation analysis was applied and conducting a usability test, it was found that the system was effective. This study proposed a system that recommends suitable fonts according to the emotion of the content, and through this, even users who lack design experience can quickly and accurately recommend fonts suitable for the situation. In addition, by applying the method proposed in this study, it can be used in various digital content fields such as video production, electronic publishing, and social networks
시각장애 학습자를 위한 강의 동영상 자동 해설 시스템 구현
은나현 숙명여자대학교 대학원 2025 국내석사
The recent digitization of education is creating new opportunities to provide an enhanced learning experience for different types of students. However, students with visual impairments encounter obstacles in engaging with online classes due to the inaccessibility of visual information. At present, a variety of assistive technologies are employed in the context of online education. However, the reality is that the support systems available are not sufficient. To address these issues, this thesis proposes an automatic lecture video commentary system for visually impaired students. The objective of this research is to enhance the learning experience of visually impaired students by analyzing visual content and generating commentary for lecture videos containing visuals. The system has been constructed as a cross-platform moblie application, with the objective of enabling users to operate the system and listen to lectures through voice commands. Once the user has selected the lecture video they wish to listen to, the system categorizes the various visual materials provided in the lecture video into text, images, tables, and diagrams. It then generates a commentary for each type of contents, converts it to speech, and provides it to visually impaired students. Naver Clova OCR API and Google Cloud Vision API are utilized to effectively analyze text and image information, and table and diagram commentary generation algorithms are developed to clearly understand the relationship between visual materials to generate more accurate commentary. To evaluate the performance of this system, this thesis compares lecture videos with automatically generated commentary to lecture videos without commentary. The results showed that the videos with commentary showed significant improvements in comprehension and satisfaction compared to the videos without commentary, and the generation speed and readability of the commentary were also positively evaluated. The automatic commentary system proposed in this thesis focuses on improving the accessibility of material within lecture videos including diagrams. To accomplish this, the thesis introduces a function which is capable of structurally conveying diagram information to the learner. This system recognizes and analyzes diagram images in lecture materials, extracts key information, and explains the information in a voice that can be easily understood by visually impaired students. To see the effectiveness of this system, this thesis conducted a usability evaluation for the most effective explanation method for visually impaired students, and based on the results, we established a diagram explanation method and implemented a diagram analysis algorithm. The diagram analysis algorithm proposed in this thesis detects arrows in the diagram and analyzes the flow to generate sequentially-proper commentaries. The usability evaluation of the algorithm shows that it outperforms the image captioning feature provided by Google in terms of user satisfaction, appropriateness, and accuracy, indicating that it functions effectively for diagram explanation. This research is expected to contribute to improving the accessibility of digital educational content by providing a more independent and inclusive learning environment for visually impaired students. It also raises the need for future research and development of automated commentary for other learning materials, such as charts and graphs, which can be used to build more advanced learning support systems. 최근 교육 방식의 디지털화는 다양한 유형의 학생들에게 더 나은 학습 경험을 제공하기 위한 새로운 기회를 열어주고 있다. 그러나 시각장애 학생들은 시각적 정보에 대한 접근성 문제로 인해 온라인 수업 참여에 어려움을 겪고 있다. 현재 온라인 교육을 위해 다양한 보조 기술이 사용되고 있으나 현실적으로 적절한 지원 시스템이 충분히 제공되지 않고 있다. 이러한 문제를 해결하기 위해, 본 연구는 시각장애 학생들을 위한 강의 동영상 해설 자동 생성 시스템을 제안한다. 본 연구의 목표는 시각적 자료가 포함된 강의 동영상에 대해 시각적 자료를 분석하고 이에 대한 해설을 생성하여 제공함으로써 시각장애 학생들의 학습 환경을 개선하고자 한다. 본 시스템은 크로스 플랫폼 어플리케이션 형태로 제작되었으며 음성을 통해 시스템을 조작하고 강의 내용을 청취할 수 있도록 설계되었다. 사용자가 청취하길 원하는 강의 동영상을 선택하면, 강의 동영상 내에서 제공되는 다양한 시각적 자료를 텍스트, 이미지, 표, 다이어그램으로 분류하여 종류별 해설을 생성하고, 이를 음성으로 변환해 시각장애 학생들에게 제공한다. 해설 생성은 다음과 같은 단계로 이루어 진다. 첫 째로, 텍스트는 Naver Clova OCR API를 이용하여 장면 이미지 내의 텍스트들을 모두 추출한다. 단, 표와 다이어그램과 같은 텍스트가 들어간 이미지 내의 텍스트와 구분하여 추출한다. 장면 이미지 내에 그림이 들어가 있는 경우, Google Cloud Vision API를 활용하여 이미지 내용을 해설로 추출한다. 다음으로 텍스트가 포함된 그림 중 화살표가 없는 그림은 표로 인식하여 표 맞춤형 구조적으로 해설 생성 알고리즘을 사용한다. 다이어그램은 텍스트 객체와 화살표를 유기적으로 연결하여 해설을 생성하는 알고리즘을 개발하여 구현하였다. 이 시스템의 성능을 평가하기 위해, 본 연구는 자동 생성된 해설을 포함한 강의 동영상을 시각장애 학생들을 대상으로 테스트하였다. 테스트 결과, 해설이 제공된 동영상은 해설이 제공되지 않은 동영상에 비해 이해도와 만족도 측면에서 유의미한 향상을 보였으며, 해설 생성 속도와 가독성에서도 긍정적인 평가를 받았다. 이러한 결과는 본 시스템이 시각장애 학생들이 디지털 학습 환경에서 시각적 정보를 보다 쉽게 접근하고 이해할 수 있도록 돕는 효과적인 도구임을 시사한다. 본 연구에서 제안하는 다이어그램 분석 알고리즘은 다이어그램 내에서 화살표를 탐지하고 흐름을 분석하여 해설을 생성한다. 본 알고리즘에 대한 사용성 평가를 진행한 결과 사용자 만족도, 적절성, 정확도 측면에서 구글에서 제공하는 이미지 캡션 기능보다 우수한 성능을 나타내었으며 이를 통해 다이어그램 해설에 효과적으로 기능하는 것을 알 수 있다. 본 연구는 시각장애 학생들에게 보다 독립적이고 포괄적인 학습 환경을 제공함으로써, 디지털 교육 콘텐츠의 접근성을 높이는 데 기여할 수 있을 것으로 기대된다. 또한, 향후 다양한 학습 자료에 대한 해설 자동화 연구 및 개발의 필요성을 제기하며, 이를 바탕으로 보다 진보된 학습 지원 시스템을 구축할 수 있을 것이다.
전산모사를 이용한 NdFeB계 희토류 자석의 열처리 거동 연구
오늘날 희토류는 IT첨단산업분야에서 없어서는 안 될 매우 중요한 소재이다. 반도체, 스마트폰, TV, 의료기기, 전기자동차 등 기계, 전자, IT 제품 외에도 태양광, 전기자동차, 풍력발전 터빈 등의 청정에너지 산업에서도 핵심적인 재료로 활용되고 있어 수요가 나날히 증가하고 있다. 그러나, 희토류의 지각내 매장량이 적고, 지리적 편중성이 높아 전량 수입에 의존하고 있어 수급이 원활하지 못하다. 이를 해결하고자, 희토류의 저감 및 재활용 관련 기술이 연구되고 있으며 그 중 NdFeB계 영구자석의 재활용 기술이 가장 활발히 연구되고 있다. 그러나, 대부분 습식공정으로, 공정이 복잡하고 사용하는 약품과 발생되는 폐수로 인한 환경오염 문제와 공정 중 소요되는 에너지를 고려하여 더 간단한 재사용 기술에 대한 연구가 필요하다 판단된다. 폐 NdFeB 영구자석은 모든 공정에 앞서 취급의 용이성을 위해 자성을 완전히 제거하는 탈자 공정을 거치는데 통상 큐리온도 이상의 온도를 가하여 탈자한다. 이 때, Nd 자석은 고온의 열처리로 표면 및 자석 내부가 산화되어 영구적으로 자기적 특성이 손상이 발생하여 재사용이 어렵다. 따라서, 본 연구에서는 실제 열처리와 전산모사를 병행하여 NdFeB 자석의 손상없는 탈자 열처리 조건을 연구하고, 탈자된 Nd 자석을 다시 착자하여 Nd 자석의 재사용 가능성 여부를 확인하고자 한다. Today, rare earth elements are indispensable and very important materials in the high-tech IT industry. In addition to machinery, electronics, and IT products such as semiconductors, smartphones, TVs, medical devices, and electric vehicles, it is also used as a key material in the clean energy industry such as solar power, electric vehicles, and wind turbines, and demand is increasing day by day. However, due to the low amount of rare earth reserves in the earth's crust and high geographical concentration, the entire amount is dependent on imports, making supply and demand difficult. In order to solve this problem, research is being conducted on technologies related to the reduction and recycling of rare earth elements, and among them, the recycling technology of NdFeB-based permanent magnets is being studied most actively. However, most of them are wet processes, and the process is complicated, and considering the environmental pollution problem caused by the chemicals used and wastewater generated, and the energy required during the process, it is judged that research on simpler reuse technologies is necessary. Waste NdFeB permanent magnets undergo a demagnetization process that completely removes magnetism for ease of handling prior to all processes, and is usually demagnetized by applying a temperature higher than the Curie temperature. At this time, the surface and inside of the Nd magnet are oxidized due to high-temperature heat treatment, which permanently damages magnetic properties, making it difficult to reuse. Therefore, in this study, the conditions for demagnetization heat treatment without damage of NdFeB magnets are studied through actual heat treatment and computer simulation, and the demagnetized Nd magnet is remagnetized to confirm the reusability of the Nd magnet.