RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 학위유형
      • 주제분류
      • 수여기관
        펼치기
      • 발행연도
        펼치기
      • 작성언어
      • 지도교수
        펼치기

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • 반복 가중치 L1-주성분 분석을 이용한 전자코 손실 데이터 복원

      전홍민 단국대학교 대학원 2017 국내석사

      RANK : 249711

      전자코 시스템은 인간의 후각의 단점을 보완하고 위험한 환경에서 사용 가능하다는 장점이 있어 많은 분야에서 데이터 측정 도구 등 다양한 용도로 사용될 수 있다. 좋은 분류 성능을 갖는 전자코를 설계하기 위해서는 효과적인 분류기의 선택도 중요하지만 분류기의 입력으로 사용되는 특징을 분류 목적에 적합하도록 추출하는 것 또한 매우 중요하다. 기체 분류에 유용한 특징을 추출하기 위해서는 먼저 센서를 통해 전자코 데이터를 안정적으로 취득하여야 한다. 그러나 전자코가 활용되는 실제 환경에서는 설치 환경이나 전기적인 문제에 의해서 센서 데이터에 이상치(outlier)나 노이즈가 발생할 수 있는데, 이는 효과적인 특징 추출을 어렵게 만들어 결과적으로 전자코 시스템의 분류 성능을 떨어뜨릴 수 있다. 본 논문에서는 센서를 통한 데이터 취득 과정에서 일부 데이터가 손실되거나 노이즈에 의해 손상되었을 경우, 통계적 학습을 기반으로 데이터를 복원하는 방법을 제안하였다. L2-PCA(Principal Component Analysis)는 센서 데이터와 같은 고차원 데이터에 대해 차원을 축소하고 노이즈 제거를 위한 복원 방법으로 사용된 바 있다. 그러나 L2-norm을 사용하는 L2-PCA의 경우 특징을 추출하는 과정에서 이상치 데이터 값이 과도하게 반영되어 실제 데이터의 분포를 제대로 나타내지 못하는 경우가 발생한다. 제안한 방법에서는 데이터 복원을 위해 L2-PCA 대신 이상치 데이터에 보다 강인한 L1-PCA를 사용하고 반복 가중치 적합법(Iteratively Reweighted Fitting, IRF)을 적용함으로써 데이터의 복원 에러를 감소시킬 수 있었다. 복원된 데이터와 무손실 데이터와의 RMS 에러 측정 결과, 제안한 방법에 의해 데이터가 효과적으로 복원되는 것을 확인 할 수 있었으며, 손상된 데이터와 복원 데이터에 대한 가스 분류 실험을 통해, 데이터가 손상 된 경우에도 제안한 방법으로 데이터를 복원함으로써 데이터 손상으로 인한 분류 성능의 감소 폭이 크게 줄어든 것을 확인 할 수 있었다. The electronic nose system can be used in a variety of applications such as data measurement tools in many fields because it has the advantages of complementing the shortcomings of human nose and being usable in a dangerous environment. In order to design an electronic nose with good classification performance, it is important to select an effective classifier, but it is also important to extract features that are used as inputs to the classifier to suit the classification purpose. In order to extract useful features for gas classification, electronic nose data must be acquired stably through the sensor. However, in an actual environment where an electronic nose is used, an outlier or noise may occur in the sensor data due to an installation environment or an electrical problem. This may make it difficult to extract an effective feature, and as a result, the classification performance of the electronic nose may deteriorate. In this paper, we propose a method to reconstruct data based on statistical learning when some data is lost or damaged by noise during data acquisition through sensor. L2-PCA (Principal Component Analysis) has been widely used as a reconstruction method for noise removal and dimension reduction for high dimensional data such as sensor data. However, in case of L2-PCA using L2-norm, outlier data values ​​are over-reflected in the process of extracting features, which may not reflect the distribution of original data. In the proposed method, it is possible to reduce data reconstruction error by using L1-PCA and Iteratively Reweighted Fitting (IRF) method instead of L2-PCA for data reconstruction. As a result of the RMS error measurement between the reconstructed data and the lossless data, it was confirmed that the data was effectively reconstructed by the proposed method. Even if the data was damaged through the gas classification experiment on the damaged data and the reconstructed data, It can be confirmed that the reduction of the classification performance is greatly reduced by reconstructing the data.

    • 불균형 정형데이터 문제 해결을 위한 SMOTE와 CycleGAN 기반 하이브리드 오버샘플링 기법 개발 및 적용 : 금융사기를 중심으로

      노정담 국민대학교 일반대학원 2020 국내석사

      RANK : 249711

      현대 사회는 사람의 행동 하나가 데이터가 되며 이는 곧 엄청난 데이터의 흐름을 만든다. 20년 전 인터넷 속 전체 데이터의 양이 현대 사회 속에서는 1초마다 저장된다. 이러한 추세는 앞으로 더욱 더 심화될 것이며 이러한 빅데이터를 활용하기에 따라서 엄청난 이점을 줄 수 있을 것으로 판단된다. 이러한 데이터의 분석을 위해서는 편향되지 않은 데이터가 필요한데 대부분의 빅데이터는 한쪽으로 편향인 불균형 상태며 이는 분석의 정확도를 떨어뜨리는 원인 중 하나이다. 또한 2종 오류의 비용이 큰 분야에서는 불균형 데이터를 사용한 분석을 믿을 수 없는 실정이기 때문에 이러한 문제점을 해결하는 것은 매우 중요하다. 정형 데이터 분야에서는 이러한 문제점을 해결하기 위해서 전통적인 방식의 오버샘플링이 발전해왔고 비정형 데이터에서는 딥러닝의 발전과 더불어 발전한 생성 모델이 불균형 문제의 해결책으로 떠올랐다. 기업에서 현재 상황을 이용하여 미래를 예측하는 일은 매우 중요하고 어려운 일이다. 이러한 문제를 해결하기 위해 대부분의 기업에서는 통계기반의 머신러닝을 이용하고 있다. 통계기반의 머신러닝을 이용한 예측에서 중요한 점은 편향되지 않은 양질의 정형데이터이다. 본 연구에서는 비정형 데이터에서 오버샘플링을 하기 위해 자주 사용하는 생성 모델인 CycleGAN을 정형 데이터에 맞게 변형시켜 TDOGAN을 만들었다. 오버샘플링과 TDOGAN을 이용해 편향되지 않은 양질의 정형 데이터를 만들어 냈다. PCA를 사용하여 개인정보를 가린 실제 금융사기 데이터에 본 논문에서 제안하는 하이브리드 오버샘플링 기법을 적용하였다. In modern society, one person's actions become data, which creates a huge flow of data. Twenty years ago, the total amount of data on the Internet is stored every second in modern society. This trend is expected to intensify in the future, and the use of such big data will bring enormous advantages. The analysis of such data requires unbiased data, and most of the big data is unbalanced, which is one of the reasons for lowering the accuracy of the analysis. It is also very important to solve this problem because the analysis using unbalanced data is unreliable in the field where the cost of type 2 error is high. In the field of structured data, traditional oversampling has evolved to address these problems, and in unstructured data, the generation model developed with the development of deep learning has emerged as the solution to the problem of imbalance. It is very important and difficult for companies to use the current situation to predict the future. To solve these problems, most companies are using statistics-based machine learning. An important aspect of forecasting using statistics-based machine learning is unbiased, quality structured data.In this study, the cycleGAN, a often used generation model for oversampling in unstructured data, was modified to fit the structured data to create TDOGAN. By using OVER-sampling and TDOGAN, it has produced high-quality structured data that is not biased. The hybrid oversampling technique proposed in this paper has been applied to actual financial fraud data covering personal information using PCA.

    • 트랜잭션 기반 머신러닝 문제의 자동화된 특성 추출을 위한 딥러닝 활용

      우덕채 국민대학교 일반대학원 2019 국내박사

      RANK : 249711

      머신러닝은 통찰력을 도출하거나 분류 및 예측을 하기 위해, 주어진 데이터를 수학적 모델에 적합시키는 방식으로 정보 기술의 발전과 다양한 스마트 기기의 등장으로 활용 가능한 데이터의 양이 기하급수적으로 증대된 빅데이터 시대에서 편향이 개입되지 않은 패턴 발견으로 높은 예측 성능을 보이고 있다. 이러한 머신러닝 수행 과정에서 해결하고자 하는 문제를 잘 설명할 수 있는 속성을 생성하는 특성공학은 머신러닝 성능에 큰 영향을 미쳐 그 중요성이 지속적으로 강조되어 오고 있다. 하지만, 이러한 중요성에도 불구하고 반복적인 검증 절차와 원천 데이터에 대한 이해 뿐만 아니라 도메인 특성에 대한 깊은 이해를 필요로 함에 따라 여전히 어려운 과업으로 여겨지고 있다. 따라서, 본 연구에서는 이러한 특성공학 과업 중 전문 지식을 요구하며 반복적으로 수행되어야 하는 특성 추출의 복잡성 및 어려움을 해결하고 머신러닝 모델의 성능을 높이기 위한 방법으로 딥러닝 기법의 적용을 제안한다. 다른 머신러닝 기법과 달리 복잡한 비정형 데이터 처리 분야에서 딥러닝 기법이 뛰어난 성능을 보이는 가장 대표적인 이유는 원천 데이터 자체로부터 특성 추출이 가능하다는 점이다. 이러한 딥러닝 기법의 장점을 비즈니스 문제 해결에 적용하기 위하여 본 연구에서는 트랜잭션 데이터로부터 자동적으로 특성을 추출하거나 직접 예측 및 분류가 가능한 딥러닝 기반의 방법들을 제안하고 데이터 특성에 따른 차이를 실험하였다. 특히, 트랜잭션 데이터와 텍스트 데이터의 구조적 유사성에 기반하여 기존의 텍스트 처리에 높은 성능을 보이고 있는 기법을 적용하였으며 트랜잭션 데이터의 특성에 따라 각 방법들의 적합성을 검증하였다. 본 연구를 통해 자동화된 특성추출의 가능성을 탐색할 수 있을 뿐만 아니라 특성 추출 과업 수행 전에 일정 수준 이상의 성능을 보이는 준거 모델의 확보가 가능할 것으로 판단된다. 또한, 해결하고자 하는 비즈니스 문제와 보유하고 있는 데이터 특성에 따라 적합한 딥러닝 모델 선택의 가이드라인을 제시할 수 있으리라 기대된다. Machine learning (ML) is a method of fitting given data to a mathematical model to derive insights or to predict. In the age of big data, where the amount of available data increases exponentially due to the development of information technology and smart devices, ML shows high prediction performance due to pattern detection without bias. The feature engineering that generates the features that can explain the problem to be solved in the ML process has a great influence on the performance and its importance is continuously emphasized. Despite this importance, however, it is still considered a difficult task as it requires a thorough understanding of the domain characteristics as well as an understanding of source data and the iterative procedure. Therefore, we propose methods to apply deep learning for solving the complexity and difficulty of feature extraction and improving the performance of ML model. Unlike other techniques, the most common reason for the superior performance of deep learning techniques in complex unstructured data processing is that it is possible to extract features from the source data itself. In order to apply these advantages to the business problems, we propose deep learning based methods that can automatically extract features from transaction data or directly predict and classify target variables. In particular, we applied techniques that show high performance in existing text processing based on the structural similarity between transaction data and text data. And we also verified the suitability of each method according to the characteristics of transaction data. Through our study, it is possible not only to search for the possibility of automated feature extraction but also to obtain a benchmark model that shows a certain level of performance before performing the feature extraction task by a human. In addition, it is expected that it will be able to provide guidelines for choosing a suitable deep learning model based on the business problem and the data characteristics.

    • 워드 임베딩 기법을 활용한 범주형 데이터의 군집분석 : 고차원 데이터를 중심으로

      조현 국민대학교 일반대학원 2019 국내석사

      RANK : 249711

      머신 러닝 분야의 대표적인 비지도 학습 방법 중 하나인 군집분석은 데이터를 서로 유사한 집단끼리 묶어주는 분석으로, 마케팅, 공학, 의학 등 다양한 분야에서 활용되고 있다(Wilson et al, 2011). 오늘날 군집 분석과 관련된 연구는 꾸준히 진행되고 있는데, 데이터가 수치형 일 때 활용될 수 있는 연구가 주를 이루고 있으며 범주형 데이터의 군집분석과 관련된 연구는 활발하게 진행되고 있지 않다(Mingoti et al, 2012). 이에 본 연구에서는 범주형 데이터의 군집분석 시, 텍스트 분석에서 주로 사용되고 있는 워드 임베딩 기법을 활용하여 데이터를 수치형으로 변환을 한 뒤 수치형 데이터에 대한 군집분석 방법을 적용하는 방법을 제시하고자 한다. 워드 임베딩은 현재 가장 많이 사용 되고 있는 기법인 Word2vec, FastText, Glove 기법을 각각 적용하였고, 각 기법 적용시 어떠한 성능의 차이를 보이는지를 비교분석 하였다. 또한 제시하는 모형의 성능을 기존의 범주형 데이터의 군집분석 모형과 비교해 보면서 모형의 우수성을 검증하였고, 이때 가장 많이 알려진 방법인 K-mode, ROCK 등 의 방법과 비교분석 하였다. 데이터의 구조에 따른 모형의 성능의 변화를 파악하기 위해 시뮬레이션을 통해 다양한 조건의 데이터를 생성한 뒤 각 데이터 조건별 모형의 성능을 비교 하였고, 나아가 실제 데이터에서도 모형이 잘 군집하는지를 평가하기 위하여 실제 데이터를 통해 모형의 성능을 평가 하였다. Clustering algorithms is technique for grouping similar data and have been used in a variety of fileds such as engineering, medicine, marketing, etc. There are lot of study about clustering analysis but, majority of studies are about algorithms for nimerical data. In this study, we propose a method that transform categorical data to numerical data using word embedding. We used three word embedding model(Skip-gram, FastText, Glove) and compared the performance with algorithms for categorical data(K-mode, ROCK). To determine the performance of the model depending on the structure of the data, we generated data with different conditions and furthermore, evaluated performance of the model through real data. we used Silhouette score and Adjusted Rand score for performance evaluation. By the Simulation, We Compared performance of the model by the number of categories and the number of data and as a result, embedding using the glove shown the best performance except where the number of categories is high and the number of data is low. We compared performance of the model through the real hospital care data and performance was good in order of K-means using glove embedding, K-mode, K-means using word2vec, K-means using FastText.

    • 효과적인 데이터 분석을 위한 창의적 문제해결에 대한 연구 : 디자인씽킹을 중심으로

      우수미 단국대학교 대학원 2018 국내석사

      RANK : 249711

      4차 산업 시대에 접어들면서 데이터는 폭발적으로 증가하며 데이터 중심 패러다임의 변화가 발생했다. 이로써 데이터의 관심과 중요성이 커지는 시대가 되었고 이러한 데이터를 다각적으로 분석하고 그 속에서 새로운 인사이트를 발견하는 데이터사이언티스트라는 직업이 생겨났다. 데이터사이언티스트를 통해 효과적으로 데이터를 분석하기 위한 역량을 알아본다. 데이터사이언티스트는 무의미한 데이터 속에서 의미 있는 인사이트를 찾아내는 기존의 직업과는 차별적인 필수 역량이 중요시되고 있다. 따라서 이러한 역할을 효과적으로 하기 위해서는 창의적 문제해결 능력은 필수조건으로 보고 있다. 창의적 문제해결 능력을 향상시키는 법을 알아보기 이전에 창의적 문제해결이란 무엇인지에 대해 창의성, 문제해결, 창의적 문제해결로 분해하여 각각의 선행 이론을 확인해 본다. 효과적으로 창의성을 향상 시킬 수 있는 방법으로 이 연구에서 택한 디자인씽킹에 대한 이론적 고찰을 해본다. 이러한 디자인씽킹이 창의성 향상에 얼마나 큰 효과가 있는지 교육을 통해 검증해 본다. 28명의 용인시 고등학교 2학년 학생을 대상으로 2일간의 디자인씽킹 교육을 받게 한 후, 사전/사후 한국교육 개발원에서 개발한 KEDI 창의적 인성 검증 테스트를 하였다. SPSS 18.0을 사용하였고 신뢰도 분석, 정규성 검증, 대응T검증을 실시하여 창의성의 변화를 측정했다. 측정 결과 창의성이 통계적으로 의의 있게 향상된 것으로 나타났다. 검증을 통하여 디자인씽킹 교육이 창의적 문제해결 능력을 향상 시키는 데 도움이 될 수 있다는 결론을 얻을 수 있었다. 따라서 창의적 문제해결 능력이 필요하고, 이 능력을 향상하기를 원하는 직업군의 사람들에게 디자인씽킹 교육이 도움 될 것으로 보았다. 차후 연구에서는 데이터사이언티스트를 대상으로 하는 디자인씽킹 교육을 통해 데이터 분석 능력이 창의적 문제해결에 얼마나 크게 영향을 미치고 작용하는지에 대한 연구, 검증이 필요할 것이다. 주제어 : 데이터사이언티스트, 창의성, 창의적 문제해결, 가추법, 디자인씽킹

    • 모바일 웹 기반의 데이터 정보 시각화에 대한 연구 : 공공 데이터 기반으로

      유진아 단국대학교 대학원 2019 국내석사

      RANK : 249711

      스마트폰 기술의 발전과 보급의 대중화는 모바일을 통해 필요한 정보를 찾는 이용자들이 증가하고 있다. 과학기술정보통신부의 2018 인터넷 이용 실태조사에 따르면 인터넷 접속기기로 스마트폰이 94.3%로 비중이 매우 높다. 이용자들이 증가함에 따라 정부부처에서도 모바일 웹 환경에서 정책 또는 관련 데이터들의 정보를 시각화하여 이용자에게 전달하는 추세이다. 정부부처 업무수행 및 홍보활동 중 정책 홍보활동에 대한 조사결과에 따르면 정부의 정책 정보나 홍보물을 접하는 경로가 스마트폰이 60%로 가장 높고 TV 55%, 컴퓨터 42% 순으로 나타났다. 스마트폰으로의 정보의 집중 양상이 유지되는 등 매체의 영향력이 압도적으로 유지되고 있다. 이에 따라 모바일로 정보를 습득하는 사용자가 많음으로써 효과적인 정보의 전달은 중요해지고, 정보의 의미를 정확하고 가장 빠르게 전달하는 효과적인 방법은 무의식적으로도 즉각적인 인지가 가능한 이미지 혹은 그림으로 제공하는 방법이다. 정보를 시각화하여 전달하는 정보와 사용자 간의 상호작용이 원활한 인포그래픽이 주목받고 있다. 이에 본 연구에서 유형 별 인포그래픽 대상으로 연구를 진행하였다. 본 연구는 선행연구 조사를 통해 공공데이터의 정보 시각화 역할 및 필요성을 파악하고, 효과적인 정보의 이해와 전달을 위한 정보 시각화의 표현 요소와 정보 시각화의 인지부하 실험도구를 도출하고, 사례조사 및 분석 통해 공공데이터의 인포그래픽 유형 별 사례를 선정하고, 공공데이터의 인포그래픽 유형 별 정보 시각화의 표현 사례를 분석하였다. 공공데이터의 효과적 전달을 위한 시각화에 있어 중요하게 고려해야 할 요인을 연구하기 위해 20~30대 청년세대를 대상으로 설문조사를 실시하였고, 연구목적을 달성하기 위해 설문조사 결과를 바탕으로 IPA를 수행하였다. 또한 모바일 웹 환경에서 인포그래픽의 유형 별 정보 시각화 정도가 공공데이터의 효과적인 전달을 목적으로 하는 학습과 인지에 어떠한 영향을 미치는지 그 결과를 도출하기 위해 인지부하 설문문항을 활용하여 실험조사를 진행하여 설문조사를 바탕으로 결과를 분석하였다. 본 연구는 인지부하를 고려한 공공데이터의 효과적인 정보 전달을 위한 정보시각화에 주목하였다. 이를 위해 공공데이터 기반의 인포그래픽을 유형 별 수집 및 분석하였다. 수집된 자료는 유형에 의해 정리된 후, 정보시각화 표현 기준 바탕으로 분석 되었다. 따라서 본 연구의 결과는 다음과 같다. 첫째, 공공데이터의 효과적인 전달을 위해 정보 시각화에서 명확한 데이터를 통해 식별이 용이한 화면의 구성이 우선적으로 필요하다. 적당한 그래픽 요소는 유형에 따라 활용하면 사용자가 더욱 쉽게 정보를 습득하고 이해할 수 있다. 공공데이터의 정보시각화는 인포그래픽으로 표현함으로써 텍스트로 이미지 등 정보유형에 따라 다양하게 표시할 수 있다. 뿐만 아니라 배경 이미지를 잘 활용하면 효과적인 정보 전달과 사용자의 흥미를 유발할 수 있다. 둘째, 공공데이터의 효과적인 전달을 위해 정보 시각화에서 유용한 콘텐츠의 제공이 필요하다. 공공데이터는 정부 또는 공공기관이 보유하는 있는 데이터로서 사용자의 전반적인 생활 편의성 확보와 이를 기반으로 한 정책참여 등 다양한 기능을 제공하고 있다. 따라서 교통, 기상, 의료, 경제, 환경, 여가 등에서 사용자의 일상생활에 직접적으로 관련이 될 수 있는 콘텐츠의 시각적 위계 설정과 시간적, 공간적, 분야별 변화를 나타냄으로써 현재의 상황을 용이하게 파악할 수 있는 유용한 콘텐츠 등의 제공이 필요하다. 셋째, 공공데이터의 효과적인 전달에서 필요한 사용자의 신체적, 정신적 부담을 줄이고, 데이터의 이해도를 증가하기 위해서는 정보에 대한 흥미유발과 메타포를 형성할 수 있는 정보의 시각화가 필요하다. 즉, 정보에 대한 이해를 돕기 위해 은유 또는 비유에 대한 시각화를 통해 사용자의 흥미를 유발할 수 있어야 한다. 이를 위해 주로 활용할 수 있는 인포그래픽을 통한 정보의 시각화 방안은 캐릭터 등의 만화적 요소를 활용하거나 일상생활과 관련된 행동, 심리 등을 활용한 정보, 두 가지 이상의 정보 유형이나 개념을 비교함으로써 사용자의 이해를 돕는 것이 필요하다. 넷째, 공공데이터의 효과적인 전달을 위해 사용자의 신체적, 정신적 측면에서 인지 부하를 감소시킬 수 있는 정보의 시각화가 필요하다. 이를 위해 일상생활이나 어떠한 행동이나 직업, 심리 등과 관련된 흥미성 자료를 기반으로 한 정보를 중심으로 캐릭터 등의 만화적 요소를 활용하여 정보를 전달하여 신체적, 정신적 측면에서 피곤하지 않고 직관적으로 흥미성을 주는 시각화가 필요하다. 다섯째, 공공데이터의 효과적인 전달을 위해 사용자의 이해가 용이한 내용 및 구성을 위한 정보의 시각화가 필요하다. 이를 위해 공공데이터에 대해 시간적 전개, 경로의 전개를 통해 전반적인 사항을 포괄적으로 이해할 수 있는 스토리텔링 방식의 정보 시각화가 필요하다. 또한 제품 또는 개념을 두 가지 이상 비교하는 방식인 비교분석형의 정보 시각화가 필요하다. As the development and popularization of smartphone technology become popular, users are looking for information through mobile. According to a survey by the Ministry of Science and Technology (MIC) on Internet usage by 2018, smartphones accounted for 94.3% of the total number of Internet access devices. As the number of users increases, government ministries also tend to visualize policy and related data in the mobile web environment and deliver it to users. According to the results of public relations activities conducted by government ministries and agencies, 60% of the respondents had access to government policy information or publicity materials, followed by 55% of TVs and 42% of computers. The influence of the media on smartphones has remained intact. As a result, there are many users who acquire information through mobile, so that effective information transfer becomes important, and an effective method of conveying the meaning of information accurately and fast is a method of providing images or pictures that can be unconsciously recognized immediately. Infographics, which are easy to interact with information and information that visualize and deliver information, are attracting attention. In this study, the study was carried out as an infographic for each type. The purpose of this study is to identify the role and necessity of information visualization of public data through previous studies and to derive cognitive load experiment tools of information visualization and information visualization for effective understanding and transmission of information, The case of infographic type of public data was selected and the case of information visualization by infographic type of public data was analyzed. In order to investigate the factors that should be considered important in visualization for effective transmission of public data, a questionnaire survey was conducted for young people in their 20s and 30s, and an IPA was conducted based on the survey results to achieve the research purpose . In addition, in order to derive the effect of information visualization level of infographic in mobile web environment on learning and cognition that is aimed at effective transmission of public data, we conducted experiment survey using cognitive load questionnaire The results were analyzed based on the survey. This study focuses on information visualization for effective information transmission of public data considering cognitive load. For this purpose, we collected and analyzed infographic of public data based on type. The collected data were analyzed by type and then based on information visualization expression standard. The results of this study are as follows. First, in order to efficiently transmit public data, it is necessary to construct a screen that can be easily identified through clear data in information visualization. Proper graphical elements can be more easily learned and understood by users if they are based on type. Information visualization of public data can be displayed variously according to the type of information such as text and images by representing it in infographic form. In addition, the use of background images can lead to effective information transmission and user interest. Second, it is necessary to provide useful contents in information visualization to efficiently transmit public data. Public data is the data held by the government or public institutions and provides various functions such as ensuring the user's overall life convenience and participating in policies based on the data. Therefore, visual hierarchy of contents that can be directly related to user's daily life in transportation, weather, medical, economic, environment, leisure, etc., and time, space, It is necessary to provide contents and the like. Third, in order to reduce the physical and mental burdens of users in order to efficiently transmit public data and to increase the understanding of data, it is necessary to visualize information that can induce interest in information and form a metaphor. In other words, to help understand the information, it is necessary to visualize the metaphor or metaphor to induce the user's interest. The information visualization method that can be used mainly for this purpose is to utilize comic elements such as characters, information using behavior related to daily life, psychological information, comparing two types of information or concepts, It is necessary to help. Fourth, in order to efficiently transmit public data, it is necessary to visualize information that can reduce the cognitive load in the physical and mental aspects of users. To do this, visualization is used to convey information by using comic elements such as characters, focusing on information based on interesting data related to everyday life or any behavior, occupation, psychology, etc., so that it is not tired and intuitively interesting in physical and mental aspects. Fifth, in order to efficiently transmit public data, it is necessary to visualize information for user's easy understanding and composition. For this, information visualization of storytelling method is needed to comprehensively understand the overall contents through temporal development and development of public data. In addition, information visualization of comparative analysis type which is a method of comparing two or more products or concepts is needed. visualization for effective information transmission of public data considering cognitive load. To do this, we collected and analyzed infographic of public data base. The collected data were analyzed by type and then based on information visualization expression standard. The results of this study are as follows. First, in the first half, the result that the visualization factor of information is important about clarity, contents, screen composition, and interest inducing factor is derived, so that it is based on information that is interesting to the user, In order to be easy to understand and use, it is necessary to provide an easy-to-understand format using visualization to make it easier to communicate to general users.

    • 정보전달을 위한 데이터시각화의 사용자경험 연구 : 스마트홈 실내공기 서비스를 중심으로

      이현진 단국대학교 대학원 2019 국내석사

      RANK : 249711

      사물의 인터랙션이 실시간으로 이루어지는 초연결 사회가 도래함에 따라, 방대한 데이터가 거의 전 분야에서 실시간으로 수집되고 있다. 데이터는 기하급수적으로 증가하고 있으나, 그 중 가치 있는 데이터를 추출해 의미 있는 데이터를 적극적으로 활용하는 전달방법에 대한 연구는 아직은 전문가 위주로 이루어져 있다고 할 수 있어 비전문가인 일반 사용자들을 위한 연구가 필요하다. 본 연구에서는 데이터 시각화 요소를 활용할 경우와 정보 내용만 전달할 경우 정보인지도가 높은 쪽을 밝히고, 매체풍요도 이론에 따라 색상이나 이미지를 사용했을 경우와 그렇지 않을 경우의 차이를 도출하고자 실험을 진행했다. 그 결과 데이터 시각화 요소의 활용이 사용자의 태스크 실행의도에 미치는 영향은 시각화 요소를 활용해 정보 전달을 할 경우가 그렇지 않을 경우보다 더 긍정적으로 나타남을 확인하였다. 추가로 정보인지도가 높을 수록 태스크 실행의도에 긍정적인 영향을 미치는 것으로 보아 향후 보다 많은 사용자들이 불편함 없이 가치있는 데이터를 활용할 수 있도록 다방면의 연구가 필요하다고 제안하는 바이다. Among the smart home service, fine dust and indoor air quality data closely related to our daily life. I conducted a research to determine whether the presence of data visualization elements in information delivery has a noticeable effect on information awareness and execution intent for a given task. For this, according to the medium richness theory, which is the basis of the study, and through the preliminary studies, the image and color of the data visualization elements were set as independent variables and the fine dust sensitivity which can affect the results were set as the control variables. After that, subjects were divided into 4 groups according to each independent variable and examined the degree of information awareness and the intention to perform 'ventilation' task. As a result, it is confirmed that the use of data visualization elements affect user 's intention of task execution more positively than without. Fortunately, data analysts field and global corporations are leading the way by researching and presenting 'accessibility' guides on IT technologies such as web and mobile, to build an improved user experience for everyone and to change overall social awareness. As a result, the problem of visualization of data and information is expected to be solved gradually, but it is still focused on 'visualization' and this is also one of the biggest limitations of this study. However, I hope that this study will contribute to bring more interests of given subject and therefore everyone should be able to freely use data and information regardless of their environment and physical condition.

    • 연관 분석 알고리즘의 분석 및 개선 : Apriori 알고리즘 개선을 중심으로

      허침 건국대학교 대학원 2023 국내석사

      RANK : 249711

      사회의 진보와 생산성이 지속적으로 향상됨에 따라 컴퓨터 기술도 지속적으로 혁신되고 대중화되었으며 기업의 정보 수집 능력도 크게 향상되었으며 데이터의 규모는 전례 없는 규모에 도달했다. 이러한 방대한 데이터 정보에 대해 기업의 원래 기술과 도구는 더 이상 요구를 충족시킬 수 없으므로 이러한 데이터를 처리하기 위한 새로운 기술을 연구하고 개발해야 한다. 이제 컴퓨터 기술과 수학 이론을 바탕으로 방대한 양의 데이터를 자동으로 처리하고 가치 있는 정보와 지식을 마이닝 및 분석하여 실제 의사결정에 도움을 주는 풍부하고 강력한 데이터 마이닝 기술이 개발되었다. 이런 배경에서 탄생한 데이터 마이닝 기술은 실제 업무와 떼려야 뗄 수 없는 관계다. 현재 데이터 마이닝 기술은 기업 운영의 효율성과 수익을 향상시키기 위해 다양한 산업에서 널리 사용되었다. 또한 다양한 비즈니스 요구를 충족시키기 위해 데이터 마이닝 기술은 클러스터링, 분류, 연관 분석 및 시계열 분석과 같은 다양한 방향으로 개발되었다. 상관 규칙 분석은 데이터 마이닝 기술에서 더 중요한 연구 방향이다. 이 논문에서는 관련 규칙 알고리즘 분석에 중점을 둔 데이터 마이닝 기술을 자세히 소개하고 실험적 연구와 고전적인 Apriori 알고리즘을 개선한다. Apriori 알고리즘의 경우 가장 핵심적인 문제는 빈번한 아이템 셋을 생성하는 것이다. 이에 데이터베이스 압축, 후보 아이템 셋 축소, 규칙에 맞지 않는 아이템 셋 사전 선별 등의 방향으로 개선해 클래식한 Apriori 알고리즘을 기반으로 알고리즘의 효율성을 더욱 향상시킨다. 이론적 분석이 완료된 후 파이썬을 통해 개선된 알고리즘의 유효성을 검증하기 위한 실험을 수행할 것이다. With the progress of society and the continuous improvement of productivity, computer technology is constantly innovating and popularizing, the information collection ability of enterprises has also been greatly improved, and the scale of data has reached an unprecedented scale. For such a huge amount of data information, the original technology and tools of the enterprise can no longer meet the needs, so it is necessary to research and develop new technologies for processing these data. Now, based on computer technology and mathematical theory, rich and powerful data mining technology has been developed, which can automatically process massive data, mine and analyze valuable information and knowledge, and provide help for practical decision-making. The data mining technology born in this context is bound to be inseparable from the actual business. At present, data mining technology has been widely used in all walks of life to improve the operational efficiency and revenue of enterprises. In addition, in order to adapt to different business needs, data mining technology has developed in different directions, such as clustering, classification, association analysis and time series analysis and so on. Association rule analysis is an important research direction in data mining technology. This paper will introduce the data mining technology in detail, focus on the analysis of the association rule algorithm, and combine the experimental research and improvement of the classic Apriori algorithm. For the Apriori algorithm, the core problem is to generate frequent itemsets. In this regard, it can be improved by compressing the database, reducing the candidate item set, and pre-screening the itemsets that do not meet the rules, so as to further improve the efficiency of the algorithm on the basis of the classic Apriori algorithm. After completing the theoretical analysis, experiments will be conducted through python to verify the effectiveness of the improved algorithm.

    • 빅데이터 기반 추천시스템 구현을 위한 다중 프로파일 앙상블 기법

      김민정 국민대학교 2016 국내석사

      RANK : 249711

      기존의 협업필터링 추천시스템 연구는 상품에 대한 고객의 평점(rating)이나 구매 여부 데이터로부터 하나의 프로파일을 생성하고 이를 기반으로 추천 성능을 향상시킬 수 있는 새로운 알고리즘을 개발하는 위주로 진행되어 왔다. 그러나 빅데이터 환경이 도래하면서 기업이 수집할 수 있는 고객 데이터가 풍부해지고 다양해짐에 따라, 보다 정확하게 고객의 선호도나 행태를 파악하는 것이 가능하게 되었고 이러한 데이터, 즉 퍼스널 빅데이터(personal big data)를 추천시스템에 활용하는 연구의 필요성이 대두되고 있다. 본 연구에서는 마케팅의 시장세분화 이론에 근거하여 퍼스널 빅데이터로부터 고객의 선호도나 행태를 다양한 관점에서 표현할 수 있는 5종의 다중 프로파일(multimodal profile)을 개발하고, 이를 활용하여 협업필터링 추천시스템의 성능을 개선하고자 한다. 제안하는 5종의 다중 프로파일은 프로파일 통합 유사도, 개별 프로파일 유사도 평균, 개별 프로파일 유사도 가중 평균이라는 세 가지 앙상블 기법을 통해 협업필터링의 이웃(neighborhood) 탐색과정에 적용된다. 실제 퍼스널 빅데이터에 본 연구에서 제안하는 방법론을 적용한 결과, 단일 프로파일을 사용하는 협업필터링 알고리즘보다 추천 성능이 개선되었으며 개별 프로파일 유사도 가중 평균 앙상블 기법이 가장 높은 추천 성능을 보여주었다. 본 연구는 빅데이터 환경에서 추천시스템을 개발하고자 할 때, 어떠한 정보를 이용하여 고객의 특성을 규명하는 프로파일을 만들고 이를 어떻게 결합하여 사용하는 것이 효과적인지 처음으로 제안하였다는 점에서 그 의의가 있다. The recommender system is a system which recommends products to the customers who are likely to be interested in. Based on automated information filtering technology, a variety of recommender systems have been developed. Collaborative filtering (CF), one of the most successful recommendation algorithms, has been applied in a number of different domains such as recommending Web pages, books, movies, music and products. However, it has been known that CF has a critical shortcoming. CF finds neighbors whose preferences are similar to those of the target customer and recommends products those customers have most liked. Therefore, CF works properly only when there’s a sufficient number of ratings on common product from customers. When there is a shortage of customer ratings, CF makes the formation of a neighborhood inaccurately, leading to poor recommendations. To improve the performance of CF based recommender systems, most related studies have been focused on the development of novel algorithms under the assumption of using a single profile, which is created from user's rating information for items, purchase transactions, or Web access logs. With the advent of big data, companies can collect more data and use a variety of information with big size. Therefore, many companies realize the importance of using big data because it makes companies to improve their competitiveness and to create new value. In particular, utilizing personal big data in the recommender system is one of the most critical issue. It is why personal big data facilitate more accurate identification of the preferences or behaviors of users. The proposed recommendation methodology is as follows: first, multimodal user profiles are created from personal big data in order to grasp the preferences and behavior of users from various viewpoints. We derive five user profiles based on the personal information such as rating, site preference, demographic, Internet usage, and topic in text. Next, the similarity between users is calculated based on the profiles and then neighbors of users are found from the results. One of three ensemble approaches is applied to calculate the similarity. Each ensemble approach uses the similarity of combined profile, the average similarity of each profile, and the weighted average similarity of each profile, respectively. Finally, the products that people among the neighborhood prefer most to are recommended to the target users. For the experiments, we used the demographic data and a very large volume of Web log transaction for 5,000 panel users of a company that is specialized to analyzing ranks of Web sites. R and SAS E-miner was used to implement the proposed recommender system and to conduct the topic analysis using the keyword search, respectively. To evaluate the recommendation performance, we used 60% of data for training and 40% of data for test. The 5-fold cross validation was also conducted to enhance the reliability of our experiments. A widely used combination metric called F1 metric that gives equal weight to both recall and precision was employed for our evaluation. As the results of evaluation, the proposed methodology achieved the significant improvement over the single profile based CF algorithm. In particular, the ensemble approach using weighted average similarity shows the highest performance. That is, the rate of improvement in F1 is 16.9 percent for the ensemble approach using weighted average similarity and 8.1 percent for the ensemble approach using average similarity of each profile. From these results, we conclude that the multimodal profile ensemble approach is a viable solution to the problems encountered when there is a shortage of customer ratings. This study has significance in suggesting what type of information we can use in order to create profile in the environment of big data and how we can combine and utilize them effectively. However, our methodology should be further studied to consider for its real-world application. We need to compare the differences in recommendation accuracy by applying the proposed method to different recommendation algorithms and then to identify which combination of them would show the best performance.

    • 베이지안 프레임워크를 활용한 암시적 피드백 데이터 기반 추천시스템 설계

      왕재준 국민대학교 일반대학원 2024 국내석사

      RANK : 249711

      The problem of individual utility maximization has never been free fromthe constraint of resource scarcity. While advancements in productiontechnologies and distribution processes have partially alleviated thisconstraint, a new one has emerged: the problem of adverse selection dueto information overload. In a context where users are confronted with anoverwhelming variety of options, identifying items that best match theirpersonal preferences becomes increasingly difficult. Recommendationsystems mitigate this challenge by selectively presenting items tailored toeach user's taste, thereby reducing information search costs and preventingadverse selection.Given that the core function of recommendation systems lies inidentifying items aligned with user preferences, the primary source oftraining data is the history of user-item interactions. Among these, explicitfeedback—where users voluntarily rate items—directly reveals preferences, but is difficult to obtain in real-world services. As a result, implicitfeedback data such as views, cart additions, and purchases are commonlyused as alternatives. However, implicit feedback signals are inherentlyambiguous and cannot be reliably interpreted as indicators of preference.This study interprets the ambiguity of implicit feedback signals as anepistemic uncertainty regarding user preferences and proposes a latentfactor model that addresses this issue through a Bayesian framework.Specifically, user behavior vectors learned from implicit feedback arerearranged via attention to the user's historical items, resulting in acontextually embedded representation of implicit preference. Likewise, itemrepresentations are personalized by conditioning them on the target user’shistory.In this model, attention scores are treated not as deterministic constants,but as probabilistic variables following a distribution, which are modeledwithin a Bayesian framework. This probabilistic attention mechanismenables the model to incorporate the epistemic uncertainty of userpreference signals into the user and item representations.Experimental results demonstrate that the proposed model effectivelycaptures the uncertainty inherent in implicit feedback data. Furthermore, itsperformance improvement over baseline models is more pronounced indatasets with sparse user-item interactions. 개인의 효용 극대화 문제는 언제나 자원의 희소성이라는 제약에서 자유롭지 못했다. 생산 기술과 유통 프로세스의 발전으로 인하여 이 제약은 일부 완화되었으나, 새로운 제약이 발생하였다. 정보과부하에 따른 역선택 문제이다. 헤아릴 수 없이 다양한 아이템을 선택할 수 있는 상황 속에서, 사용자는 오히려 어떤 아이템이 자신의 선호체계에 가장 적합한지 탐색하기 어렵게 되었다. 이 새로운 제약 상황에서 추천시스템은 사용자 개개인에게 그 기호에 알맞은 아이템을 선별하여 제안함으로써, 사용자의 정보 탐색 비용을 줄이고 역선택을 방지하는데 기여하고 있다. 이처럼 사용자의 기호에 알맞은 아이템을 선별하는 것이 핵심 기능이라는 점에서, 사용자가 아이템에 대하여 상호작용한 이력은 추천시스템 설계 시 추천 품질을 주요하게 결정하는 데이터이다. 이 중 사용자의 선호가 직접적으로 드러나는 이력은, 사용자가 아이템의 선호도를 자발로 수치화한 명시적 피드백 데이터이다. 하지만 실제 서비스 환경에서는 명시적 피드백 데이터를 획득하는 데 어려움이 있다. 때문에 조회, 장바구니, 구매 등 암시적 피드백 데이터가 그 대안으로 활용되고 있다. 그런데 이 암시적 피드백 데이터는 단순히 선호로 간주하기에는 그 신호가 모호하다는 한계점이 존재한다. 이에 본 연구는 암시적 피드백 데이터의 선호 신호 모호성 문제를 사용자 선호에 대한 인식적 불확실성 문제로 정의하고, 이를 베이지안 프레임워크로써 풀이하는 잠재요인 모형을 제안하였다. 구체적으로, 암시적 피드백 데이터로부터 학습되는 사용자의 행동 벡터를 그 사용자의 히스토리 아이템과의 어텐션을 적용하여 암묵적 선호 표현으로서 벡터 공간에 재배열하였다. 마찬가지로 아이템의 특성 벡터를 목표 사용자의 히스토리 맥락에서 재해석함으로써 해당 아이템의 특성을 개인화된 표현으로서 재구성하였다. 이때 어텐션 스코어를 확정적인 값을 취하는 상수에서 어떠한 확률 분포를 따르는 확률변수로 전환하고, 이를 베이지안 프레임워크로 모델링하였다. 이로써 암시적 피드백 데이터로 인해 발생하는 선호에 대한 인식적 불확실성을 사용자와 아이템의 벡터 표현에 반영하였다. 실험 결과, 제안 모형이 암시적 피드백 데이터의 선호 신호 모호성 문제를 효과적으로 다룰 수 있음과 더불어, 사용자-아이템 상호작용 밀집도가 희소한 데이터 셋일수록 비교 모형 대비 추가적인 성능 개선 효과가 있음을 확인하였다.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼