RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 학위유형
      • 주제분류
        펼치기
      • 수여기관
        펼치기
      • 발행연도
        펼치기
      • 작성언어
        펼치기
      • 지도교수
        펼치기

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • Intelligent Data Acquisition for Predictive Modeling in Manufacturing Systems

      심재웅 서울대학교 대학원 2021 국내박사

      RANK : 2943

      Predictive modeling is a type of supervised learning to find the functional relationship between the input variables and the output variable. Predictive modeling is used in various aspects in manufacturing systems, such as automation of visual inspection, prediction of faulty products, and result estimation of expensive inspection. To build a high-performance predictive model, it is essential to secure high quality data. However, in manufacturing systems, it is practically impossible to acquire enough data of all kinds that are needed for the predictive modeling. There are three main difficulties in the data acquisition in manufacturing systems. First, labeled data always comes with a cost. In many problems, labeling must be done by experienced engineers, which is costly. Second, due to the inspection cost, not all inspections can be performed on all products. Because of time and monetary constraints in the manufacturing system, it is impossible to obtain all the desired inspection results. Third, changes in the manufacturing environment make data acquisition difficult. A change in the manufacturing environment causes a change in the distribution of generated data, making it impossible to obtain enough consistent data. Then, the model have to be trained with a small amount of data. In this dissertation, we overcome this difficulties in data acquisition through active learning, active feature-value acquisition, and domain adaptation. First, we propose an active learning framework to solve the high labeling cost of the wafer map pattern classification. This makes it possible to achieve higher performance with a lower labeling cost. Moreover, the cost efficiency is further improved by incorporating the cluster-level annotation into active learning. For the inspection cost for fault prediction problem, we propose a active inspection framework. By selecting products to undergo high-cost inspection with the novel uncertainty estimation method, high performance can be obtained with low inspection cost. To solve the recipe transition problem that frequently occurs in faulty wafer prediction in semiconductor manufacturing, a domain adaptation methods are used. Through sequential application of unsupervised domain adaptation and semi-supervised domain adaptation, performance degradation due to recipe transition is minimized. Through experiments on real-world data, it was demonstrated that the proposed methodologies can overcome the data acquisition problems in the manufacturing systems and improve the performance of the predictive models. 예측 모델링은 지도 학습의 일종으로, 학습 데이터를 통해 입력 변수와 출력 변수 간의 함수적 관계를 찾는 과정이다. 이런 예측 모델링은 육안 검사 자동화, 불량 제품 사전 탐지, 고비용 검사 결과 추정 등 제조 시스템 전반에 걸쳐 활용된다. 높은 성능의 예측 모델을 달성하기 위해서는 양질의 데이터가 필수적이다. 하지만 제조 시스템에서 원하는 종류의 데이터를 원하는 만큼 획득하는 것은 현실적으로 거의 불가능하다. 데이터 획득의 어려움은 크게 세가지 원인에 의해 발생한다. 첫번째로, 라벨링이 된 데이터는 항상 비용을 수반한다는 점이다. 많은 문제에서, 라벨링은 숙련된 엔지니어에 의해 수행되어야 하고, 이는 큰 비용을 발생시킨다. 두번째로, 검사 비용 때문에 모든 검사가 모든 제품에 대해 수행될 수 없다. 제조 시스템에는 시간적, 금전적 제약이 존재하기 때문에, 원하는 모든 검사 결과값을 획득하는 것이 어렵다. 세번째로, 제조 환경의 변화가 데이터 획득을 어렵게 만든다. 제조 환경의 변화는 생성되는 데이터의 분포를 변형시켜, 일관성 있는 데이터를 충분히 획득하지 못하게 한다. 이로 인해 적은 양의 데이터만으로 모델을 재학습시켜야 하는 상황이 빈번하게 발생한다. 본 논문에서는 이런 데이터 획득의 어려움을 극복하기 위해 능동 학습, 능동 피쳐값 획득, 도메인 적응 방법을 활용한다. 먼저, 웨이퍼 맵 패턴 분류 문제의 높은 라벨링 비용을 해결하기 위해 능동학습 프레임워크를 제안한다. 이를 통해 적은 라벨링 비용으로 높은 성능의 분류 모델을 구축할 수 있다. 나아가, 군집 단위의 라벨링 방법을 능동학습에 접목하여 비용 효율성을 한차례 더 개선한다. 제품 불량 예측에 활용되는 검사 비용 문제를 해결하기 위해서는 능동 검사 방법을 제안한다. 제안하는 새로운 불확실성 추정 방법을 통해 고비용 검사 대상 제품을 선택함으로써 적은 검사 비용으로 높은 성능을 얻을 수 있다. 반도체 제조의 웨이퍼 불량 예측에서 빈번하게 발생하는 레시피 변경 문제를 해결하기 위해서는 도메인 적응 방법을 활용한다. 비교사 도메인 적응과 반교사 도메인 적응의 순차적인 적용을 통해 레시피 변경에 의한 성능 저하를 최소화한다. 본 논문에서는 실제 데이터에 대한 실험을 통해 제안된 방법론들이 제조시스템의 데이터 획득 문제를 극복하고 예측 모델의 성능을 높일 수 있음을 확인하였다.

    • Robust Adaptation Strategies of Data-Driven Models for Chemical Process System

      이경민 서울대학교 대학원 2024 국내박사

      RANK : 2943

      The modeling of chemical processes is an important step in the simulation, optimization, and control of systems, and various studies have employed both first-principles and data-driven methods. However, these methods face several challenges when modeling industrial processes. First-principles modeling can construct reliable models if the mechanism of the process is accurately understood. However, the increasing complexity of modern processes poses challenges to accurately capturing their mechanics. Data-driven modeling does not require prior knowledge of the process, making it applicable to complex systems. However, in industrial processes, data is often limited to normal operating conditions, which can lead to overfitting of the data-driven model. In addition, processes are subject to rectifications and changes in operating conditions, posing the challenge of acquiring sufficient data for the modified process. Therefore, to apply data-driven modeling in industrial chemical processes, adaptation strategies are needed to address the aforementioned challenges. In this regard, this thesis proposes effective and robust adaptive strategies to address the challenges faced during data-driven modeling of industrial chemical processes. First, the application of envelope Bayesian optimization is proposed to minimize the data needed for optimizing the catalyst packing ratio in the Fischer-Tropsch microchannel reactor. The envelope Bayesian optimization was modified to accommodate the different optimization dimensions of single-channel and 4-channel reactors. The proposed optimizer not only reduces the data required to reach the optimal point compared to generic Bayesian optimization but also demonstrates robust optimization performance even when the optimization range varies. Second, a hybrid modeling approach is proposed for predicting concentration, integrating prior process knowledge with data-driven modeling methods. This approach leverages a first-principles model of the process to enhance the accuracy and robustness of concentration prediction. The proposed hybrid model improves generalization performance and exhibits better prediction accuracy compared to the data-driven model when encountering process changes. Moreover, the hybrid model is used to construct a refractive index fault detection model by augmenting the monomer concentration of the stream which has longer sampling frequency intervals. This fault detection model demonstrates higher performance compared to models relying solely on existing process data. Proposed robust adaptation strategies to overcome the challenges of implementing data-driven modeling in industrial chemical processes, encompassing issues such as acquiring new process data, handling data imbalance, and addressing differences in sampling frequency. These adaptation strategies were then applied to various chemical processes, demonstrating improved performance compared to conventional modeling methods. 화학 공정의 모델링은 시스템의 시뮬레이션과 최적화 그리고 제어를 하는 데 있어 필수적인 과정이다. 따라서 제일 원리 또는 데이터 기반 방법론을 이용하여 모델링을 하는 연구들이 진행해 왔다. 그러나 실제 산업 공정들을 모델링 하기에 기존의 방법들은 몇 가지 문제점들이 존재한다. 먼저 제일원리 기반 모델링의 경우 프로세스의 메커니즘이 정확히 파악될 경우 높은 신뢰도의 모델을 얻을 수 있지만, 화학 산업 공정의 경우 복잡한 시스템을 가지고 있어 신뢰도 있는 모델을 구성하는 데 비용과 노력이 많이 든다. 데이터 기반 모델링은 공정에 대한 사전 지식이 필요하지 않아 복잡한 시스템에 적용이 용이하다. 그러나 실제 산업 공정의 경우 잘 제어된 시스템이기 때문에 정상운전 범위 내의 데이터가 주를 이루기 때문에 모델의 과적합 문제를 발생시키게 된다. 그리고 산업 현장에서는 공정의 수정이나 운전 시스템의 변화가 발생하는데, 이 경우 변화된 프로세스의 데이터를 충분히 얻어야 한다는 문제가 있다. 따라서 데이터 기반 모델링을 실제 화학 산업 공정에 적용하기 위해서는 앞서 언급한 문제점들을 해결할 수 있는 적응 전략이 필요하다. 이러한 관점을 기반으로 본 논문은 실제 화학 산업 공정에 데이터 기반 모델링 과정에서 발생하는 문제점들을 해결하기 위한 효과적이고 강건한 적응전략을 제시한다. 먼저 피셔-트롭쉬 마이크로채널 반응기의 촉매 충전 비율을 최적화하는 데 필요한 데이터의 수를 줄이기 위해 엔벨로프 베이지안 최적화를 적용하는 방법을 제시하였다. 이 때, 단일 채널 반응기와 4-채널 반응기의 최적화 차원이 달라지는 것을 해결하기 위해 엔벨로프 베이지안 최적화 방법론을 수정하여 적용하였다. 제안한 최적화 방법론은 기존 베이지안 최적화에 비해 최적점에 도달하는 데 필요한 데이터 수가 적을 뿐 아니라 최적화 범위가 달라지더라도 강건한 최적화 성능을 보였다. 둘째로, 제일 원리 모델링과 데이터 기반 모델링을 결합한 하이브리드 모델링 기법을 제안하여 산업 공정의 데이터 불균형 및 샘플링 빈도 차이로 발생하는 문제를 해결하였다. LSTM과 제일 원리 모델을 결합하여 프로세스 내의 단량체 조성을 예측 진행하였다. 제안한 하이브리드 모델을 공정에 대한 사전 지식을 데이터 기반 모델링에 결합함으로써 모델의 일반화 성능을 높여 데이터 기반 모델링 보다 공정의 변화가 생겼을 때의 뛰어난 예측 성능을 보였다. 개발한 하이브리드 모델을 통해 샘플링 빈도 주기가 긴 스트림 내 단량체의 조성 값을 대체하여 굴절률 이상 감지 모델을 개발하였다. 개발한 이상 감지 모델은 기존 데이터만을 쓴 모델에 비해 높은 예측 정확도를 나타내었다. 본 논문은 실제 화학 산업 공정에 데이터 기반 모델링 적용시 발생하는 문제들인 새로운 프로세스에 대한 데이터 필요성과 데이터 불균형 및 샘플링 빈도 차이를 해결할 수 있는 강건한 적응 전략을 제시하였다. 제시한 적응 전략들을 여러 화학 공정에 적용하여 기존 방법론들 대비 제안한 방법론의 개선된 성능을 보여 주었다.

    • 개인정보 전송요구권과 데이터 경제 : 쟁점과 실행 방향

      권정 서울대학교 대학원 2024 국내박사

      RANK : 2943

      우리나라는 인공지능의 본격적인 활용과 데이터 경제 시대의 도래에 발맞추어 정보주체의 의사에 따라 데이터를 통합할 수 있는 데이터 전송의 법적 근거를 마련하였다. 「신용정보의 이용 및 보호에 관한 법률」이 2020년 2월 4일 개정되면서 금융 마이데이터 사업인 본인신용정보관리업과 개인신용정보의 전송요구권을 규율하였다. 이어서 2023년 3월 14일 개정된 「개인정보 보호법」에 전 분야에서의 개인정보 전송요구권에 대한 법적 근거가 마련되었다. 기존 문헌들은 개인정보 전송요구권을 정보주체의 개인정보 자기결정권 강화에 초점을 맞추어 법리적으로 분석해 왔다. 그러나 이미 개인정보 전송요구권이 법제에 도입된 현 시점에서는 개인정보 전송요구권 실행의 활성화가 직면 과제라 할 수 있다. 이에 따라 본 논문은 정보주체 뿐 아니라 개인정보를 제공하거나 제공받는 기업들까지 고려하여 개인정보 전송요구권에 대해 공리주의적, 법경제학적 관점에서 분석한다. 이를 통해 최적의 개인정보 전송요구권 실행 방안을 검토하고, 이를 실현하기 위한 법제도의 개선 방향을 모색한다. 개인정보 전송요구권은 경제적인 관점에서 접근하여도 결국 정보주체의 효용을 늘리는 결과를 가져온다. 개인정보 전송요구권의 도입은 데이터의 유통량을 늘리고 가격을 낮추는 효과를 가져온다. 이에 따라 데이터를 활용하여 최종생산품을 생산하는 생산자는 생산요소로서의 데이터를 낮은 가격에 확보할 수 있게 되어 생산비용을 절감할 수 있다. 그 결과 소비자들은 더 낮은 가격에 더 많은 최종생산품을 소비할 수 있게 된다. 특히 정보주체의 개인정보를 활용한 맞춤형 상품들은 제품차별화로 이어져 해당 정보주체의 효용을 늘려줄 수 있다. 개인정보 전송요구권의 도입은 데이터 최종생산품 시장의 경쟁을 활성화하는 효과도 가져온다. 소비자는 데이터 전송이 가능하면 전환비용을 절감할 수 있어 새로운 데이터 서비스 플랫폼으로의 전환이 간편해진다. 나아가 데이터 전송은 하나의 기업이 여러 서비스를 결합하여 판매하던 상품의 언번들링을 용이하게 하여 소비자의 선택권을 확대한다. 또한 개인정보 전송요구권은 데이터 피드백 루프, 네트워크 효과, 데이터 접근 배타성을 약화하여 데이터 최종생산품 시장의 경쟁도 촉진할 수 있다. 정보주체의 권리에만 주목하는 관점은 단기의 데이터 유통을 넘어 장기의 데이터의 생산까지 고려했을 때 한계가 있다. 데이터 전송에 대한 적절한 대가가 주어지지 않으면 결국 데이터 생산의 유인이 낮아지는 공유지의 비극이 발생할 수 있다. 다만 개인정보 전송요구권 도입 초기에 무리하게 전송 대가를 요구하면 개인정보의 유통이 저하될 수밖에 없으며 개인정보 전송요구권의 실효성도 떨어진다. 따라서 전송 대가가 없어도 무방한 분야에서 우선 시행할 필요가 있다. 이러한 산업은 정보 제공의 대칭성이 유지될 수 있는 산업으로, 특히 동일한 계위의 기업들이 병존하는 산업이 이에 해당한다. 개인정보 전송요구권이 경쟁질서에 미치는 순기능을 고려하면 개인정보 전송요구권은 독점적 산업에 우선적으로 도입하는 것이 효율적이라고 여겨질 수 있다. 하지만 정보 제공의 대칭성까지 고려하면 같은 계위의 시장 참여자들이 다수 존재하는 과점산업이 무정산 데이터 전송을 하기에 적합한 산업일 것이다. 대표적인 예는 대형병원 중심의 보건의료 산업과 통신·인터넷 산업이다. 이들 산업에 우선적으로 전송요구권을 도입하는 것이 적합한 추가적인 사유들 또한 살펴본다. 다만 데이터 전송의 대칭성은 계위가 다른 기업과는 유지될 수 없다. 계위가 다른 기업 간의 전송 대가에 대해서는 각 산업별 특수성을 고려하여 개별적으로 정해야할 것이다. 나아가 개인정보 전송요구권이 실효성을 거두려면 거래비용을 낮추는 제도적 기반이 구축되어야 한다. 데이터 전송 과정에서의 동의 절차를 간명화하여야 하며, 데이터가 다른 기관으로 전송되어도 호환될 수 있는 형태로 표준화될 필요가 있다. 또한 마이데이터 산업의 성장을 도모하기 위해 진입규제가 데이터 보안과 안보의 유지까지 종합적으로 고려한 합리적인 수준으로 완화되어야 한다. 마지막으로 2024년 5월 1일에 입법예고 된 개인정보 전송요구권 관련 사항을 구체화 하는 개인정보 보호법 시행령 일부개정령(안)과 관련된 정책적 현안에 대해서 검토한다. 중소기업의 정보전송자로서의 부담, 민감정보 해외 유출 위험, 개인정보 제3자 제공과의 운영상 유사성 등 현재 실제로 문제되는 사항을 살펴보고 개선 방향을 제안한다. 개인정보 전송요구권은 정보주체의 권리 강화에서 나아가 데이터 경제 활성화라는 목표도 달성할 수 있어야 한다. 이를 위해서는 개인정보 전송요구권에 대하여 공리주의적 관점에서 접근할 필요가 있다. 한국은 개인정보 전송요구권의 도입으로 목표한 바를 달성하기 위해서는 장기적으로 개인정보 전송요구권이 데이터 경제 활성화에 기여할 수 있는 방향으로 정책을 마련하여야 한다. 이에 관한 본 연구가 개인정보 전송요구권에 대한 법경제학적 연구의 배경이 되고, 나아가 한국이 데이터 중심의 4차 산업혁명 시대를 선도하는데 이바지하기를 기대한다. The Republic of Korea has established a legal framework for data transfer, allowing data integration according to the data subject's intent, in light of the rise of artificial intelligence and the emergence of the data economy. The amendment of the Credit Information Use and Protection Act on February 4, 2020 introduced the financial MyData business and the right to data portability of personal credit information. Additionally, the amendment of the Personal Information Protection Act on March 14, 2023 established a legal basis for the right to data portability across all sectors. Existing studies on the right to data portability have mainly focused on the legal analysis of the right, emphasizing the enhancement of the data subject’s right to informational self-determination. As the right to data portability is now implemented, however, the pressing concern is activating the exercise of the right. Accordingly, this paper approaches the right to data portability from a utilitarian and law and economics point of view, taking into consideration the benefits and costs of data transfer for all relevant parties: the data subjects, the companies providing data, and the companies receiving data. In doing so, the paper seeks the optimal execution plan of the right to data portability and suggests directions for legal system improvements. The right to data portability ultimately increases utility for the data subject even though it is approached from economic perspective. The right enhances data circulation and lowers its price, allowing producers of data final products who use data as a production factor to acquire it at a lower cost. Consequently, consumers can enjoy more final products at lower prices. Customized products that utilize the personal information of data subjects can particularly increase utility through product differentiation. The introduction of the right to data portability also stimulates competition in the market for final data products. Data portability reduces switching costs making it easier for consumers to switch to new data service platforms. Furthermore, data transfer facilitates the unbundling of products that were previously sold by a single company, thereby expanding consumer choices. In addition, the right to data portability can weaken data feedback loops, network effects, and excludability in data access, promoting market competition. Considering not only short-term data circulation but also long-term production of data, it is insufficient to solely focus on the aspect of the rights of the data subject regarding the right to data portability. Without appropriate compensation for data transfer, the incentive to produce data reduces and a tragedy of the commons may occur. If, however, compensation for data transfer is adopted prematurely, the circulation of data will decline, and the effectiveness of the right will diminish. Therefore, it is reasonable to implement the right to data portability in industries where compensation for transfer is not necessary. This is particularly relevant in industries that can maintain symmetry of information provision, such as those where companies of the same tier coexist. Given the positive effects of the right to data portability on market competition, it might be readily regarded efficient to prioritize the implementation of the right in monopolistic industries. Nevertheless, considering the symmetry of information provision, oligopolistic industries with many market participants at the same tier is more suitable for settlement-free data transfer. Typical examples include the healthcare industry centered around major hospitals and the telecommunications and internet industries. Additional reasons for prioritizing the introduction of the right to data portability in these industries are also examined. Such symmetry, however, cannot be maintained between companies of different tiers. compensation for data transfer should be addressed on an individual basis, given the unique characteristics of each industry. For the right to data portability to be effective, an institutional foundation that lowers transaction costs should be established. Consent procedures in the data transfer process should be simplified, and data should be standardized to ensure interoperability when transferred to other institutions. Also, in order to promote growth of the MyData industry, entry regulations should be appropriately eased, considering data security and safety comprehensively. Last, this paper addresses policy issues related to the proposed amendment to the Enforcement Decree of the Personal Information Protection Act, announced on May 1, 2024, which specifies matters related to the right to data portability. It examines current issues, such as the burden of small and medium- sized enterprises as data transmitters, the risk of overseas leakage of sensitive information, and operational similarities of the right to data portability and provision of personal information to third parties, and proposes directions for improvement. The right to data portability should not only strengthen the rights of data subjects but also achieve the goal of revitalizing the data economy. To this end, it is necessary to approach the right to data portability from a utilitarian perspective. For the Republic of Korea to achieve its objectives through the implementation of this right, policies must be developed to support the long-term enhancement of the data economy. This study seeks to provide a legal and economic background for research on the right to data portability and to assist the Republic of Korea in leading the data-driven Fourth Industrial Revolution.

    • 집합과 계층구조 개념을 이용한 데이터스트림에서의 유사 시퀀스 매칭 기법

      여은지 연세대학교 일반대학원 2016 국내석사

      RANK : 2943

      본 논문에서는 집합과 계층구조 개념을 이용하여 데이터스트림에서 새로운 유사 시퀀스 매칭 기법인 SHS(Set and Hierarchy-based Similar sequence matching)를 제안하였다. 데이터스트림이란 시간의 흐름에 따라서 계속해서 순차적으로 무한히 생성되는 데이터를 말한다. 이러한 데이터스트림의 환경은 대량의 데이터가 실시간으로 빠르고 무한하게 들어오는 특징을 갖는다. 인터넷과 이동통신의 발달에 힘입어 네트워크 패킷, 이동통신 통화기록, 주식시세 변화 데이터 등 다양한 데이터스트림들이 사용되고 있는데, 본 논문에서는 사용자별로 시간에 따라 선호한 영화의 평점 데이터를 주요 대상 데이터스트림으로 하였다. 최근 디지털 기기의 보급과 소셜 네트워크 서비스의 발전으로 인해 영화에 대한 평점 데이터를 사용자들이 언제 어디서나 실시간으로 계속해서 생성할 수 있게 되면서, 예전에는 정적인 데이터로 보았던 영화 평점 데이터를 동적인 데이터스트림으로 모델링 할 수 있게 되었다. 그리고 데이터스트림 처리 응용으로는 영화 평점을 사용한 추천 시스템(Recommendation Systems)에 초점을 맞춰 연구를 수행하였다. 추천 시스템은 수많은 아이템 중에서 사용자가 선호할만한 아이템을 추천해주는 시스템이다. 추천 시스템의 가장 널리 쓰이는 알고리즘은 협업 필터링(collaborative filtering)이다. 협업 필터링은 추천 서비스를 받을 액티브 사용자(active user)와 유사한 다른 사용자를 찾고, 이렇게 찾은 유사 사용자가 선호한 아이템 중에서 액티브 사용자가 선호한 적이 없는 다른 아이템을 추천하는 방식이다. 협업 필터링은 전적으로 유사한 사용자를 기반으로 추천을 수행하므로 유사 사용자를 정확하게 찾는 것이 무엇보다 중요하다. 본 논문에서는 협업 필터링의 핵심 요소인 유사 사용자 매칭 방법을 보다 정확하게 수행하기 위해 해당 문제를 데이터스트림에서의 유사 시퀀스 매칭으로 변환하여 해결하였다. SHS는 다음의 세 가지 특징을 갖는다. 첫 번째로 시간의 흐름에 따른 사용자의 선호에 집합 개념을 도입한 “선호 아이템 집합 시퀀스” 구조를 제안하였다. 사용자의 선호에 단순히 시간 정보만을 추가하여 유사 사용자 매칭을 수행하면 두 사용자가 공통으로 선호한 아이템이 정확히 같은 순서에 존재해야 유사한 사용자로 선정되고 조금이라도 순서가 다르면 유사하지 않다고 판단되는 문제가 있다. 본 논문에서는 이러한 문제를 해결하기 위해서 시간의 흐름에 따른 사용자의 선호를 일정 시간 간격씩 모아서 집합으로 묶음으로써 해당 일정 시간 간격 안에서는 사용자의 선호 아이템의 시간 정보가 정확히 일치하지 않아도 유사한 사용자를 찾아낼 수 있는 방법을 제시하였다. 이때 아이템 집합 시퀀스 간의 유사도를 측정하기 위해 유클리디안 거리를 집합으로 확장한 유클리디안 집합 거리를 제안하였다. 두 번째로 사용자의 선호에 아이템의 계층구조를 고려한 매칭 방법을 제안하였다. 기존의 협업 필터링은 단순히 두 사용자 간의 공통되는 아이템의 정보만을 통해 시간에 따라서 더 많은 공통 선호 아이템이 존재할 경우 유사한 사용자로 선정하였다. 이러한 기존 방법은 두 사용자가 실제로는 선호도가 유사할지라도 공통적으로 선호도를 표시한 아이템의 수가 적으면 유사 사용자로 선정되지 않는 문제가 있었다. 이러한 문제를 선호 데이터 희소 문제라고 하며, 유사 사용자 매칭의 성능을 저하시키는 대표적인 이유이다. 본 논문에서는 아이템 그 자체만을 비교하는 것이 아니라, 계층구조를 갖는 아이템 속성까지도 유사도 판단에 고려함으로써 이러한 문제를 해결하는 방법을 제시하였다. 세 번째로 유사 사용자 매칭 문제를 유사 시퀀스 매칭 문제로 변환하여 사용자의 선호를 최근의 시점에서만 검색하는 것이 아니라 과거의 모든 시점에 대해서도 검색이 가능하도록 하였다. 사용자의 선호는 변화하므로 현재 액티브 사용자의 최근 선호와 유사한 선호를 과거에 가졌었던 다른 사용자가 존재할 수 있다. 본 논문에서는 과거 시점의 유사 사용자를 찾을 수 있도록 하기 위해서 액티브 사용자의 최근 선호와 다른 사용자들의 현재 시점뿐만 아니라 과거 시점까지 검색하는 서브 시퀀스 매칭 방법을 제안하였다. 또한 제안한 유사 시퀀스 매칭을 수행할 때에 실제로 유사하지만 유사하지 않다고 판단되는 착오기각이 발생하지 않음을 증명하였다. 실험 결과, 제안하는 SHS가 실제 영화 평점 데이터에서 유사 시퀀스 매칭을 수행하여 기존의 방법보다 유사한 사용자를 보다 정확히 찾아내는 것을 보였다. 이러한 결과로 볼 때 본 논문에서 제안하는 SHS가 데이터스트림 환경의 추천 시스템에서 보다 정확한 추천을 가능하게 하는데 유용하게 활용될 수 있을 것으로 판단된다. In this thesis, we have proposed a new technique for similar sequence matching in data streams which we call it as Set and Hierarchy-based Similar Sequence Matching(SHS). A data stream is a sequence of data entries that continuously arrive in a sequential order. Due to advances in Internet and mobile communication technologies, there are many types of data streams in real applications such as network packets, mobile phone call logs, and stock quotes data. As a target data stream, we have focused on the movie rating data which reflect the preference of users over time. In the past, the movie rating data is considered as static data, but now it can be modeled as dynamic data streams because users can rate movies in real time, anytime, and anywhere due to the recent spread of digital communication devices and social network services, As a target data stream application, we have focused on Recommendation Systems(RSs) using the movie rating data. RSs produce a list of recommended items that the active user may have an interest. The most widely used algorithm for RSs is collaborative filtering. Collaborative filtering first find similar users who have similar interests with the active user, and then recommend items that preferred by the similar users but not rated by the active user. The most important step for collaborative filtering is to exactly find similar users because recommended items are completely gleaned from the similar user. In order to improve accuracy of finding similar users, we have transformed the similar user matching problem into the similar sequence matching problem in data streams. Our SHS has the following three characteristics. First, we have proposed the preferred item set sequence which reflects the time concept of user preference. A preferred item set sequence is an ordered list of sets where each set collects preferred items within in a specific time interval. We also have proposed a similarity measurement between item set sequences, the Euclidean set distance, which extends Euclidian distance. Second, we have exploited the item hierarchy concept of attributes for similar sequence matching. The existing collaborative filtering method identifies similar users based on idea that the more common preferred items two users have, the more similar they are. However, this method has a problem that two users who have actually similar interests are not matched as similar users when they have only a few number of common preferred items. This is a data sparsity problem and considered as a major reason reducing the accuracy of the similar user matching. To solve the data sparsity problem, we have proposed a method for the similar user matching using not only preferred items themselves but also the attribute hierarchies of the items. Third, in order to find the similar user not only in current time but also in past time, we have transformed the similar user matching problem into the similar subsequence matching problem. We also have proved that the method does not incur false dismissals which are actually similar to the active user but discarded in the results of the similar sequence matching. Through experiments with real data sets, we have shown that SHS provides higher accuracy than the existing method for finding similar users. From the experiment results, we consider that our SHS is a practical and useful method for improving accuracy of RSs in data stream.

    • Data-driven disability persona augmentation for service journey inference : focusing on daily living support needs of people with developmental disabilities

      Lee, Sangyun Sungkyunkwan university 2025 국내석사

      RANK : 2943

      장애인을 위한 서비스디자인은 그 필요성과 잠재적 가치에도 불구하고, 소수성과 시장성 부족이라는 이중의 제약으로 인해 학문적·실무적 관심에서 상대적으로 저평가되어 왔다. 특히 발달장애인의 경우, 스펙트럼성으로 인해 개인 간 차이가 극단적으로 다양하게 분포함에 따라, 일반화와 대표성 확보가 어려워 설계 및 평가 과정에서의 정량적 분석이 난해하다. 본 연구는 이와 같은 발달장애인의 소수성, 스펙트럼성, 산발성이라는 구조적 문제를 해결하기 위한 대안으로, 데이터 증폭 기법과 알고리즘적 클러스터링을 활용하여 대표 페르소나를 도출하고, 이를 일상 시나리오에 매핑하여 일상생활지원서비스 니즈를 정량적으로 진단하는 Augmented Persona Mapping Framework를 제안한다. 본 연구는 두 가지 목표를 중심으로 설계되었다. 첫 번째 목표는 제한된 수의 실제 발달장애인 데이터를 증폭하여 스펙트럼 다양성을 확보함으로써, 일반화 가능한 대표 페르소나를 정의할 수 있는 알고리즘을 개발하는 것이다. 두 번째 목표는 도출된 대표 페르소나를 일상생활지원 시나리오에 대입하여, 활동 단위별로 필요한 기능과 제약 요인을 구조화하고, 서비스 니즈 수준을 수치화하는 정량적 분석 체계를 마련하는 것이다. 이를 위해 다음과 같은 연구가 진행되었다. 우선, 실제 발달장애인 6명의 인터뷰 데이터를 K-Vineland-II 알고리즘에 기반하여 총 2,406명의 가상 데이터로 증폭하였다. 이후 주성분 분석 및 계층적 클러스터링을 통해 기능 특성에 따라 5개의 대표 페르소나를 도출하였다. 이후 대표 페르소나를 일상생활 기반 시나리오(e.g., 학교 생활, 교통 이용, 식사 준비, 쇼핑 등)에 대입하여, 각 시나리오 수행을 위한 활동 단위별로 필요한 기능 및 적응행동 요소에 대한 수행 제약 요인을 대형 언어 모델을 통해 해석하였다. 수치화하여 분석된 서비스 니즈 수준 및 세부 내용은 고객 여정 지도 형태로 시각화하였다. 본 프레임워크의 유효성은 실제 발달장애인 사례의 수행 데이터를 기준으로 검증하였으며, 그 결과 대표 페르소나는 실제 발달장애인의 장애 수준에 대해 75% 이상의, 이에 기반한 저니맵은 실제 발달장애인의 일상 수행 경향과 86% 이상 수준의 높은 정합도를 나타냈다. 이를 통해 해당 프레임워크가 활동 제한의 원인과 서비스 개입 필요 수준을 구체적으로 제시할 수 있는 분석 도구로 기능함을 확인하였다. 이는 장애인의 일상생활 내 서비스 니즈를 직관이 아닌 구조적·반복 가능한 방식으로 해석할 수 있음을 시사한다. 본 연구는 장애인의 행동 데이터를 구조화·증폭하여 페르소나 및 저니맵 설계에 활용할 수 있는 가능성을 실증함으로써, 정성 중심의 기존 장애 연구에 정량 기반 접근을 도입하고, 실증성과 반복 가능성을 동시에 확보하는 분석 프레임워크를 제시하였다는 점에서 의의가 있다. 향후에는 이를 디지털 트윈 기반 UX 설계 및 AI 기반 서비스디자인과 연계함으로써, 사용자 반응 예측 및 시뮬레이션 기반 서비스 전략 수립으로 확장할 수 있을 것으로 기대된다. Despite the clear necessity and potential value of service design for people with disabilities, it remains underrepresented in both academic and practical domains due to the dual limitations of population scarcity and low market viability. In particular, the extreme behavioral diversity among people with developmental disabilities—stemming from spectrum-based variability—makes generalization and representativeness especially difficult, posing fundamental challenges to quantitative analysis in service design and evaluation. To address these structural issues—scarcity, spectrum diversity, and behavioral fragmentation—this study proposes the Augmented Persona Mapping Framework. The framework integrates data augmentation and algorithmic clustering to derive representative personas and maps them onto daily life scenarios to quantitatively assess support needs in daily living. This research was designed around two core objectives. The first is to develop an algorithm that generates representative personas by augmenting a small set of real-world cases to reflect spectrum-level diversity. The second is to apply these personas to common daily living scenarios, structurally identifying activity-specific barriers and required functions, and visualizing support needs through a quantitative analytical framework. This study expanded interview data from six individuals with developmental disabilities into 2,406 synthetic cases using a K-Vineland-II–based data augmentation process. Principal component analysis and hierarchical clustering analysis were then used to extract five representative personas categorized by functional characteristics. These personas were applied to real-life scenarios—such as school routines, transportation use, meal preparation, and shopping—and large language models were employed to interpret both functional and adaptive behavioral limitations and quantify the required level of support. The results were visualized as structured journey map. Validation was conducted by comparing the framework’s outputs to real performance data from two people with developmental disabilities not included in the training data. The representative personas demonstrated over 75% alignment with actual disability levels, and the journey maps showed more than 86% similarity in activity patterns, confirming the framework’s effectiveness as an analytical tool for identifying service intervention points. This study provides empirical evidence for structuring and amplifying behavioral data for use in persona modeling and service need mapping. It introduces a replicable and data-driven alternative to traditional qualitative methods in disability-focused design. Furthermore, the proposed framework holds future potential for simulation-based service strategy development, particularly when integrated with digital twin environments and AI-driven service design.

    • Data Mining Method for Offshore Structures based on Big Data Technology

      박성우 서울대학교 대학원 2019 국내석사

      RANK : 2943

      As many products as ships and offshore structures are constructed in the shipyard, and various data are generated and stored in the design or construction stage. Big data technology needs to be applied to process data of large size quickly, obtain meaningful results and use it for decision making. In this paper, we propose a solution to two of the problems that may occur in the shipyard. One of the two problems which can arise in the shipyard has mainly happened in the design stage. Engineers can make the mistake of choosing the wrong material in the design process, and the wrong material selection in the design process can directly lead to a design error. Another problem may arise during the procurement and purchase process. In the absence of additional information such as lead time of material or inventory at the time of procurement, additional time is required to retrieve the data. Both problems arise predominantly from the unskilled. Therefore, the purpose of this study is to establish a 8 system that can inform the engineers about the relationships between materials which can be obtained by association analysis and material requirements which can be obtained by regression analysis. This kind of system can help the engineers to reduce design errors and time consuming due to the procurement process. The information of piping materials used in an offshore structure can be regarded as ‘big data’ because of their various types and size, and the data mining algorithms based on the big data technology are applied to data related to the offshore structures. To analyze the relationship between materials for design, ‘frequent pattern growth algorithm’ was used. For material requirement analysis, big data technology-based regression analysis was used to generate a regression model, respectively. Finally, the proposed method was used to check the relationship between materials, and to predict material requirement, and verified the effectiveness of the proposed method by comparing each result with actual cases. 조선소에 많은 선박과 해양 구조물들이 건조되면서 다양한 데이터들이 설계 및 건조 과정에서 생성되고 누적된다. 누적된 데이터를 빠르게 처리하여 의사설정에 이용하려는 필요성에 발생함에 따라 빅데이터 기술의 도입 필요성도 함께 커지고 있다. 본 연구에서는 조선소에서 발생할 수 있는 두 가지 사례에 대하여 빅데이터 기반 데이터 마이닝 방법을 통한 해결책을 제안하고자 한다. 첫 번째 문제점은 설계 단계에서 발생할 수 있는 문제로, 설계 과정에서 적절하지 못한 자재를 선정하여 그것이 오작으로 이어지는 경우이다. 또 다른 한가지는 구매 및 조달 과정에서 발생할 수 있는 문제로, 조달 과정을 관리하기 위한 자재 관련 추가적인 정보를 검색하는데 추가적인 시수가 소요된다는 점이다. 두 가지 문제 모두 미숙련자에게서 주로 발생하며, 본 연구에서는 연관성 분석을 이용한 자재 추천과 회귀 분석을 이용한 소요량 예측이라는 90

    • Semi-supervised Representation Learning with Decomposition-based Data Augmentation for Time Series Analysis

      김도균 서울대학교 대학원 2024 국내박사

      RANK : 2943

      The rapid advancements in data collection methods and storage technologies have dramatically increased the variety and volume of time series data available. Although this data is pivotal for numerous industrial applications and decision-making processes, a significant challenge arises due to the labor-intensive and time-consuming nature of data labeling. This challenge is compounded by the continuous accumulation of data, which leads to a scenario where unlabeled data far outnumbers the labeled data. In response to these challenges, this thesis employs deep learning-based representation learning techniques within a semi-supervised framework for time series analysis. These techniques facilitate automated labeling for subsets of data, aiding decision-making processes in scenarios with limited labeled data. By integrating deep learning into representation learning, our method effectively addresses the imbalance between labeled and unlabeled data, extracting valuable insights even from sparsely labeled datasets. This thesis initially proposes a deep learning-based representation model, termed NNCLR-TS, designed to extract features from univariate time series data using a novel single-step, semi-supervised contrastive learning approach. This model comprises an encoder for representation extraction, and a memory structure known as the support set, which aids in pseudo-labeling and facilitates nearest neighbor operations. Within the encoder, two convolutional networks analyze the data from both temporal and frequency perspectives, allowing the model to learn a diverse range of features. Furthermore, the Support set, a dedicated memory structure, stores representations extracted by the encoder in a latent space. This arrangement aids in pseudo-labeling and the selection of training pairs via nearest-neighbor operations. Appropriate augmentation techniques are essential for contrastive learning. We introduce a novel time series decomposition-based data augmentation technique based on STL decomposition. Unlike jittering and scaling, which may compromise intrinsic time series characteristics such as periodicity, our proposed augmentation technique preserves these features, resulting in more natural augmented data. We also propose new loss functions that utilize label information, enhancing the learning performance beyond traditional contrastive learning loss functions. These include loss functions considering the similarity within a batch and between the nearest neighbors of given data. This novel approach not only improves the model's accuracy but also ensures its applicability in various real-world scenarios. The proposed model is applied to various time series classification datasets to validate its performance in univariate time series classification. We investigate performance improvements achieved by utilizing label information, even in scenarios with minimal labeled data. Finally, we adapt the proposed model for anomaly detection tasks within a self-supervised framework, applying it to various anomaly detection datasets. We assess the model's performance using metrics like precision and recall and explore the potential for performance enhancement through transfer learning. Our experimental results demonstrate that, in both time series classification and anomaly detection tasks, the proposed model outperforms existing semi-supervised and self-supervised representation learning models. 데이터 수집 수단 및 저장 기술의 발전에 따라 활용할 수 있는 시계열 데이터의 종류 및 양이 증가하고 있다. 이러한 데이터는 다양한 산업 현장에서 필수적인 역할을 하며, 그 데이터가 갖고 있는 의미를 파악함으로써 의사결정에 도움을 받을 수 있게 된다. 이를 위해 일반적으로 전문가의 레이블링 작업이 필수적으로 요구된다. 그러나, 지속적으로 수집되는 대량의 데이터를 전문가가 일일이 레이블하는 것은 비효율적이며 시간과 비용이 많이 든다는 문제가 있다. 이에, 본 논문은 준지도 학습 기법을 이용하여 시계열 분석을 진행한다. 이 기법은 일부 데이터만 레이블링이 되어 있는 상황에서 나머지 데이터에 대한 자동화된 레이블링을 가능하게 하여, 사용자의 의사결정에 도움을 줄 수 있다. 본 논문은 먼저 단변량 시계열 데이터로부터 대조적 학습을 통해 표현을 추출하는 모델을 제안한다. 인코더 내부의 두 개의 합성곱 연산 기반의 네트워크는, 데이터를 시간적 관점 뿐만 아니라 주파수적 관점에서도 접근하여 다양한 특징을 모델이 학습할 수 있도록 한다. 또한 메모리 구조의 차용을 통해 인코더로부터 추출된 표현을 잠재 공간 내에 저장해두고 이를 수도 레이블링 및 최근접이웃 연산을 통한 학습쌍 선정을 할 수 잇도록 한다. 대조적 학습에는 적절한 증강 기법이 필수적으로 요구된다. 본 논문에서는 STL기법을 기반으로 하는 새로운 시계열 분해 기반의 데이터 증강 기법을 제안한다. 이 때 분해된 각 요소 중 일부 요소에 대해서 샘플링 기반의 변형을 가함으로써 기반 데이터의 분포를 따르는 증강된 데이터를 생성할 수 있도록 한다. 지터링 및 스케일링 등은 시계열의 주기성 등의 특징을 해칠 수 있는 위험성이 존재하는 반면, 제안된 증강 기법은 시계열의 특징을 따르도록 하여 보다 더 자연스러운 증강된 데이터를 생성할 수 있도록 한다. 본 논문에서는 레이블 정보를 활용할 수 있는 새로운 손실함수를 제안한다. 배치내 데이터간의 유사도를 고려하는 손실함수와, 주어진 데이터의 최근접이웃간의 유사도를 고려하는 손실함수를 새롭게 제안함으로써 기존 대조적 학습 손실함수에서 활용할 수 없었던 레이블 정보를 활용하여 학습 성능을 높일 수 있도록 한다. 제안된 모델을 다양한 시계열 분류 데이터셋에 적용하여 단변량 시계열 데이터 분류 문제에서의 성능을 검증한다. 레이블이 극히 일부만 존재하는 상황에서 레이블 정보를 활용했을 때 성능이 향상될 수 있는지 탐구한다. 마지막으로 레이블 정보를 활용할 수 없는 이상 탐지 문제를 위해 제안 모델을 자기지도 학습 상황에 맞춰 모델을 수정한다. 수정된 모델을 여러 이상 탐지 데이터셋에 적용하여 정밀도 및 재현율 등의 지표를 통해 모델의 성능을 검증한다. 또한 전이학습을 통한 모델의 성능 향상 가능성을 탐구한다. 실험결과 분석을 통해 시계열 분류 문제 및 이상 탐지 문제에서 제안모델이 기존 준지도 및 자가지도 표현 학습 모델에 비해 더 뛰어난 성능을 보임을 확인한다.

    • Data-Driven Approach for Turbulence Modeling in Rotating Flows and Stratified Flows

      Huang, Xinyi The Pennsylvania State University ProQuest Dissert 2023 해외박사(DDOD)

      RANK : 2943

      Turbulence modeling, including wall models in large-eddy simulations (LESs) and RANS models in Reynolds-averaged Navier-Stokes (RANS) simulations, is usually considered for traditional canonical flows, such as a channel flow or a shear layer flow. Consequently, the simulations with turbulence modeling have difficulty in handling unconventional flows, including rotating flows and stratified flows. Some of the main difficulties lie in the fact that these flows have multiple flow controlling parameters (FCPs), and thus, the flow behavior is hard to explore, let alone get accurate modeling.The data-driven approach is considered a possible solution to this. The increasing computational resources and shared turbulence data allow another way to utilize the data other than pure human analyses of the physics. However, pure data-driven methods are often criticized for their weak interpretability and generalizability.In this dissertation, multiple data-driven techniques are applied to some persistent problems in turbulence modeling under the circumstances of rotating flows and stratified flows. The problems include not only the accurate modeling of the flow but also the efficient FCP space exploration, model selection, and uncertainty quantification, etc. Both the dataset and existing knowledge of physics are utilized, and then data-driven approach shows the interpretability and generalizability. They show how these traditionally difficult problems can be tackled through physics-informed data-driven approach, which significantly saves human labor.In chapter 1, a detailed literature review of physical problems, difficulties in turbulence modeling, and data-driven approach provide a brief overview of the current research and the objectives of this dissertation.In chapters 2 and 3, wall-modeled large-eddy simulations (WMLESs) are explored. For a spanwise rotating channel, the mean flow shows a linear profile and wall models can be developed in both a physics-based approach and a data-driven approach. The data-driven approach shows better accuracy and capability to generalize, which makes it a more appealing choice to save human labor in developing wall models. For an arbitrarily directional rotating channel, the mean flow does not have a known profile so it is currently impossible to find a wall model through physical understanding. To handle a large number of FCPs, the FCP space is explored by Bayesian optimization and a wall model is developed through a surrogate model, namely Gaussian progress regression. In summary, the capability of wall modeling is extended to flows with rotation.In chapters 3, 4 and 5, RANS simulations are explored. For flows controlled by different physical processes, a recommender system is developed to automate the process of model selection. Meanwhile, the feature vectors from the recommender system align with existing experiences of the dominating physical processes in a quantity of interest (QoI) and the ability of a RANS model. Therefore, the prediction from a recommender system can be physically interpreted since it is consistent with human experiences. For a stratified wake, the multi-stage behavior of the flow requires a switch of modeling as the flow develops. A linear logistic regression classifies the flow into weakly stratified turbulence (WST) regime and strongly stratified turbulence (SST) regime accurately when the dataset of only one flow condition is fed into the classifier during training. The classifier will find the dominating physical processes which is shared among different flow conditions, which is also consistent with how the regimes are physically identified. The applicable range of a RANS model is then identified through a global epistemic uncertainty quantification (UQ) method for a stratified shear layer. This method allows the exploration of dominating terms in a RANS model and determining a priori if a calibration can generalize. They are quantified through effectiveness and inconsistency, which are factors that calibration will consider.In general, a data-driven approach has been used for multiple applications in turbulence modeling. The involvement of data shows power in improving prediction accuracy and saving human labor, and the consideration of the underlying physics enables its interpretability and generalizability. More work can be done in the future for multiple aspects of turbulence modeling to realize accurate prediction in real-world flow conditions.

    • Output stationary NPU를 위한 데이터 레이아웃 최적화 및 벡터 연산 유닛 설계

      김보열 서울대학교 대학원 2024 국내박사

      RANK : 2943

      To handle the high computational demands and memory bandwidth requirements of deep neural networks (DNNs), a new type of processor called the neural processing unit (NPU) has been proposed. NPUs commonly consist of a matrix unit that supports matrix multiplication and convolution operations, and a vector processing unit (VPU) responsible for vector operations and general computations. While matrix units can be implemented in various ways, the systolic array is widely used in many NPUs. Systolic arrays can be categorized based on their dataflow. The weight stationary systolic array (WS-SA) employing weight stationary dataflow has been increasingly adopted in recent NPUs. However, large systolic arrays benefit from using output stationary dataflow in the output stationary systolic array (OS-SA), which can more easily enhance computation utilization. Nonetheless, the data layout—a method of how data is stored in the memory—introduces different constraints for OS-SA compared to WS-SA. This difference makes it inefficient to operate an OS-SA based NPU using the previous NPU organization and software. This study approaches the inefficiencies arising from the data layout in OSSA based NPUs in two ways. Firstly, we develop a software framework, known as the layout mapping framework, capable of representing the data layout of OS-SA. We then employs heuristics to optimize the selection of a data layout that reduces the overall execution time of DNNs among various data layouts available in OS-SA. Secondly, an instruction set is designed to efficiently handle the frequent data rearrangements that occur when using the OS-SA’s data layout in the vector processing unit (VPU). To select the optimal data layout for OS-SA, a layout mapping framework capable of representing the data layout of OS-SA was established. While similar frameworks targeting CPU, GPU, and WS-SA-based NPUs already exist, there is a limitation of existing frameworks which is their inability to adequately represent the data layout specific to OS-SA. To address this issue, a new data layout representation based on Graphene was adopted, tailored specifically for OS-SA NPUs within the layout mapping framework. The introduction of this new data layout representation alone demonstrated a performance improvement of up to 39 times faster in convolutional neural networks compared to using the previous data layout representations. Next, the study proposed and designed a layout mapping heuristic optimized for each layer within a DNN. Traditional layout mapping heuristics predominantly used a propagation-based approach, where the data layout is initially set for operations in the systolic array, and then propagated to ensure adjacent layers share the same layout. This approach, however, does not explore the variety of data layouts available for OS-SA, ignore the performance of non-systolic array operations, and does not optimize the endpoint of propagation. To overcome these limitations, this research introduces a new optimization method based on simulated annealing. It proposes two new state transition operations tailored to the layout mapping problem and enhances convergence stability by incorporating an additional propagation technique. This method demonstrated a performance increase of approximately 20 % in BERT-base models, showcasing its effectiveness. Finally, the instruction set for the VPU was tailored to accommodate the frequent data rearrangements used in OS-SA. The VPU often exchange data with the systolic array, thus it needs to process data stored in the OS-SA’s data layout or convert data into the OS-SA’s data layout before it is input into the OS-SA. Specifically, the VPU must adhere to the blocked data layout characteristic of OS-SA. Previously, the instruction set for VPUs in NPUs was designed with WS-SA in mind, necessitating multiple instructions for data rearrangement or including instructions that resulted in high hardware costs in the VPU. To address these issues, an instruction set aligned with OS-SA’s blocked data layout was proposed. This proposed instruction set includes operations such as block-broadcasting and block-rotate, which differentiate between operations inside and outside of blocks. When VPUs including this proposed instruction set were used in NPUs, they demonstrated an 8 % faster performance in BERT-base compared to NPUs with VPUs designed for WS-SA. Keywords: Data layout, NPU, Vector processing unit, Data layout mapping Student Number: 2017-23638 Deep neural network (DNN)의 높은 연산량과 memory bandwidth 요구량을 처리하기 위해 neural processing unit (NPU)이라는 새로운 형태의 프로세서를 고안되었다. NPU는 공통적으로 matrix multiplication과 convolution 연산을 지 원하는 matrix unit과 벡터 연산 및 범용 연산을 담당하는 vector processing unit (VPU)로 이루어져 있다. Matrix unit은 다양한 방법으로 구현되지만, systolic array가 많은 NPU에서 사용되고 있다. Systolic array는 dataflow에 따라서 종류를 나눌 수 있는데, 최근 NPU에서 많이 채택되는 systolic array는 weight stationary dataflow를 사용하는 weight stationary systolic array (WS-SA)이다. 하지만 큰 systolic array 에서는 output stationary dataflow를사용하는 output stationary systolic array (OS-SA)가쉽게 utilization을 높일 수 있다는 장점이 있다. 다만 data의 저장 방식인 data layout 에 대해 OS-SA는WS-SA와는다른 constraint 가 생기게 되고,이는 기존WS-SA 를 포함하는 NPU 구성 방식으로는 OS-SA가 포함된 NPU를 효율적으로 구동할 수 없게 만든다. 본 연구는 OS-SA 기반 NPU의 data layout으로 발생하는 비효율을 해결하 기 위해 2 가지 방법으로 접근한다. 첫번째로, OS-SA의 data layout을 표현할 수 있는 software framework 인 layout mapping framework를 구성하고, OS-SA의 여러 data layout 중 전체 DNN 수행 시간을 감소시키는 data layout을 선택하 는 문제를 heuristic으로 최적화하였다. 두번째는 OS-SA의 data layout을 사용할 경우 빈번하게 일어나는 data rearrangement를 VPU에서 효율적으로 처리할 수 있도록 instruction set을 구성하였다. 먼저 data layout을 선택하기 위하여 OS-SA의 data layout을 표현할 수 있는 layout mapping framework를 구성하였다. Layout mapping framework는 기존에 도 CPU, GPU, WS-SA의 NPU를 목표로 많이 구현되어 있고, framework 내부에 는 data layout을 표현하기 위한 data layout representation 과 mapping을 수행 하기 위한 layout mapping heuristic들이 구현되어 있다. 하지만 기존 framework 의경우 data layout representation이 OS-SA의 data layout을표현하지못한다는 한계점이 존재한다. 이 문제를 해결하기 위해 최근 제안된 Graphene 기반 data layout representation을 도입하여 OS-SA NPU를 목표로 하는 layout mapping framework를 구성하였다. 해당 data layout representation의 도입만으로도 기존 data layout representation을사용하는것보다 convolutional neural network에서 39 배 빠른 성능향상을 보여주었다. 다음으로, DNN 내의 각 layer 별 data layout을 최적화하는 layout mapping heuristic을 설계하고 제안하였다. 기존 layout mapping heuristic은 전파 방식의 mapping heuristic을 주로 사용하였다. 이 방식은 systolic array에서 수행되는 연 산에대해먼저 data layout을설정한뒤에인접 layer들이동일한 data layout을가 지도록 data layout을전파해나간다.이방식은 OS-SA가가지는여러 data layout 을 exploration해보지않고, systolic array에서수행되는 layer의 data layout은수 동적으로 정해지며, 전파가 종료되는 지점을 최적화하지 않는다는 한계점이 있다. 본 연구에서는 이 한계점 해결을 위해 simulated annealing 기반의 최적화 방법을 새로 제안한다. Layout mapping 문제에 맞춰 두 가지 state transition operation 을 새롭게 제안하였고, 이에 더불어 추가 전파 기법을 더해 수렴 안정성을 증가시 켰다. 이러한 방법으로 BERT-base에서 약 20 %의 성능 증가를 보일 수 있었다. 마지막으로, OS-SA에서 자주 사용되는 data rearrangement에 맞춰 VPU의 instruction set을 구성하였다. VPU의 경우 systolic array와 data를 주고받기 때 문에 OS-SA의 data layout으로 저장된 data를 입력으로 받아 연산을 수행하거나 OS-SA로 입력될 data를 OS-SA의 data layout으로 바꿔주는 작업을 수행해야 한 다. 특히 OS-SA의 사용하는 data layout 특징인 blocked data layout을 따라야 한다. 기존 NPU에서 VPU의 instruction set 은 WS-SA를 목표로 구성이 되어 data rearrangement를 위해 여러 instruction을 사용해야 하거나, 높은 하드웨어 비용을 만들어 내는 instruction이 포함이 되어있었다. 이 문제를 해결하기 위해 OS-SA의 blocked data layout에 맞춘 instruction set을 제안하였다. 제안하는 instruction set은 block 내부와 외부를 구분 지어 수행되는 block-broadcasting, block-rotate instruction을 포함하고 있다. 제안하는 instruction set을 포함하는 VPU를구성하였을때WS-SA을목표로한 VPU를포함한 NPU보다 BERT-base 에서 8 % 빠른 성능을 보여줄 수 있었다. 주요어: Data layout, NPU, Vector processing unit, Data layout mapping 학번: 2017-23638

    • 스마트시티 IoT 품질 지표 개발 및 우선순위 도출 : 정형 센서 데이터를 중심으로

      양현모 연세대학교 정보대학원 2021 국내석사

      RANK : 2943

      In the era of the Fourth Industrial Revolution, the importance of "Big Data" is increasing enough to be likened to "21st century crude oil." For smart city IoT data, more attention should be paid to quality control because the quality of data leads to the quality of public services. Although data quality has been presented by ISO/IEC agencies and by various domestic and foreign agencies through various perspectives, it has a limitation that it is limited to the 'user' center. Data quality indicators centered on 'supplier' are required to collect and deliver data produced by sensors. To overcome these limitations, this study derived three categories and 13 indicators of supplier-oriented smart city IoT data quality evaluation index based on FGI technique. The data subject to the quality assessment was limited to 'Structured sensor data' that is currently mainly loaded in the data hub of smart city project . Through AHP analysis, the priority of the index categories and data quality indicators were derived, and the one-sample T-test, one-way ANOVA, and reliability analysis were conducted using SPSS 25 to investigate the validity of each indicator based on the four indicator feasibility measurement items. As a result, priorities were derived for each category in the order of "Sensor Data Collection Phase, Data Convergence and Delivery Phase, and Overall Operational Phase" and the final ranking of the indicators was determined in the order of "Confidence, Completeness, Timeliness, Data Volume and Objectivity". In addition, one-sample T-test, one-way ANOVA, and reliability analysis confirmed that the validity of the indicators is guaranteed. This study is of academic significance in that it derived the smart city IoT sensor-type data quality index from the perspective of the data provider that has not been studied before. In addition, for individuals or entities performing the task of collecting, aggregating and transmitting sensor data, the indicators can contribute to improving sensor data quality by providing the basic requirements that the data should have. Also, data quality control can be carried out based on the index priority derived from the survey of experts in the IoT field to provide an improvement in the efficiency of quality control work. 4차산업혁명 시대에 ‘빅데이터’는 ‘21세기 원유’로 비유될 만큼 그 중요성이 증대되고 있다. 스마트시티 IoT 데이터의 경우 데이터의 품질이 공공서비스의 품질로 이어지기 때문에 더욱 품질관리에 주의를 기울여야 한다. 데이터 품질은 ISO/IEC 기관 및 국내/외 여러 기관에서 다양한 관점을 통해 지표가 제시되었지만, 이는 ‘사용자’ 중심에 한정되어 있다는 한계점을 지닌다. 센서에서 생산되는 데이터를 수집하고 전달하는 ‘공급자’ 중심의 데이터 품질 지표가 요구된다. 본 연구는 이러한 한계점을 극복하기 위하여 FGI 기법을 기반으로 공급자 중심의 스마트시티 IoT 데이터 품질 평가지표 3개의 카테고리와 13개의 지표를 도출하였다. 품질 평가의 대상이 되는 데이터는 현재 스마트시티 프로젝트 데이터 허브에 주로 적재되는 ‘정형 센서 데이터’로 한정하였다. AHP 분석을 통하여 지표의 카테고리와 데이터 품질 지표의 우선순위를 도출하였고 4개의 지표 타당성 측정 항목을 기반으로 각 지표의 타당성을 조사하기 위하여 SPSS 25를 활용한 일표본 T-검정과 일원배치 분산분석, 그리고 신뢰도 분석을 수행하였다. 그 결과 카테고리별로‘센서 데이터 수집 단계, 데이터 융합 및 전달 단계, 전체 운용 단계’ 순서로 우선순위가 도출되었고 지표의 최종 순위는 ‘신뢰성, 완전성, 적시성, 충분성, 객관성’ 순서로 결정되었다. 더불어 일표본 T-검정과 일원배치 분산분석, 신뢰도 분석을 수행한 결과 지표의 타당성이 보장됨을 확인하였다. 본 연구는 그동안 연구되지 않았던 데이터 공급자 관점에서 고려한 스마트시티 IoT 센서 정형 데이터 품질 지표를 도출하였다는 점에서 학문적 의의를 가진다. 더불어 센서 데이터를 수집하고 취합하여 전달하는 직무를 수행하는 개인 혹은 기업에게 해당 지표는 데이터가 지녀야 하는 기본적인 요건을 제시함으로써 센서 데이터 품질 향상에 기여할 수 있다. 또한 IoT 분야 전문가들의 설문을 통해 도출된 지표 우선순위를 기반으로 데이터 품질관리를 수행하여 품질관리 업무 효율의 향상을 제공할 수 있다.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼