RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Dimensionality Reduction Considered Harmful (Some of the Time) = 차원축소는 위험하다 (생각보다)

    한글로보기

    https://www.riss.kr/link?id=T17450109

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    차원 축소(Dimensionality Reduction, DR)는 시각 분석에서 가장 널리 사용되지만 동시에 가장 쉽게 오해되는 도구 중 하나이다. 따라서 DR을 활용한 시각 분석은 신뢰성을 잃기 쉽다. 다시 말해, 분석에서 얻은 인사이트가 실제 데이터를 정확히 반영하지 못해 잘못된 지식 형성과 의사결정으로 이어질 수 있다. 이러한 문제는 DR 투영이 원래 데이터의 모든 특성을 본질적으로 포착할 수 없으며, 이러한 한계가 분석 과정에서 충분히 고려되지 않는 데에서 비롯된다.

    본 학위논문에서는 DR 기반 시각 분석의 신뢰성을 향상시키는 방법을 제안한다. 먼저, 시각 분석에서 DR을 사용할 때 실무자들이 직면하는 신뢰성 문제를 이해한다. 우리는 인터뷰 연구와 문헌 조사를 결합하여 실무자들이 실제로 DR을 어떻게 활용하는지 상세히 분석하였다. 이후 세 가지 주요 문제를 해결하기 위한 기술을 설계하였다. 첫째, 대표적인 DR 기법인 t-SNE와 UMAP의 만연한 오용을 완화한다. 이 기법들은 군집 및 클래스 분리도를 과도하게 부각시키기 때문에 심미적으로 보기 좋다는 이유로 부적절한 분석 작업에 잘못 사용된다. 기존 DR 평가 지표들은 클래스 레이블을 기반으로 하기 때문에, 클래스를 곧바로 군집의 정답으로 간주하며 이러한 편향을 더욱 증폭시킨다. 이에 우리는 실무자들이 군집 분석을 더 신뢰성 있게 지원하는 투영을 식별할 수 있도록 하는 새로운 평가 지표를 제안한다. 둘째, DR 하이퍼파라미터의 체리 피킹 문제를 해결한다. 적절한 평가 지표가 있더라도, DR 기법 선택과 하이퍼파라미터 최적화는 많은 시행착오를 요구하며, 이는 실무자들이 기본 설정에 의존하거나 임의로 하이퍼파라미터를 골라 쓰게 만든다. 우리는 데이터셋 특성을 자동으로 반영하여 최적의 투영을 효율적으로 탐색하는 적응형 최적화 워크플로우를 제안하여 분석가들이 체계적으로 최적화를 시행하도록 유도한다. 셋째, DR 투영에서의 상호작용 오류를 줄인다. 고차원 공간은 저차원 공간보다 훨씬 높은 자유도를 가지므로, 적절히 최적화된 DR 투영에서도 왜곡을 완전히 피할 수 없다. 이러한 왜곡은 DR 투영 상에서 사용자가 수행하는 브러싱과 같은 상호작용을 부정확하게 만든다. 우리는 사용자가 군집을 탐색할 때 왜곡을 보정해주는 Distortion-aware Brushing 기법을 제안하여, 사용자가 목표한 고차원 군집을 정확하게 포착할 수 있도록 지원한다.

    마지막으로, DR 기반 시각 분석의 신뢰성을 근본적으로 강화할 수 있는 미래 연구 방향을 제시한다. 이는 관련 담론의 활성화부터 최적의 DR 투영을 자동으로 선택하는 방향까지를 포괄한다. 본 논문은 신뢰가능한 시각 분석의 더 폭넓은 적용을 위한 기반을 마련하며 마무리된다.
    번역하기

    차원 축소(Dimensionality Reduction, DR)는 시각 분석에서 가장 널리 사용되지만 동시에 가장 쉽게 오해되는 도구 중 하나이다. 따라서 DR을 활용한 시각 분석은 신뢰성을 잃기 쉽다. 다시 말해, 분석...

    차원 축소(Dimensionality Reduction, DR)는 시각 분석에서 가장 널리 사용되지만 동시에 가장 쉽게 오해되는 도구 중 하나이다. 따라서 DR을 활용한 시각 분석은 신뢰성을 잃기 쉽다. 다시 말해, 분석에서 얻은 인사이트가 실제 데이터를 정확히 반영하지 못해 잘못된 지식 형성과 의사결정으로 이어질 수 있다. 이러한 문제는 DR 투영이 원래 데이터의 모든 특성을 본질적으로 포착할 수 없으며, 이러한 한계가 분석 과정에서 충분히 고려되지 않는 데에서 비롯된다.

    본 학위논문에서는 DR 기반 시각 분석의 신뢰성을 향상시키는 방법을 제안한다. 먼저, 시각 분석에서 DR을 사용할 때 실무자들이 직면하는 신뢰성 문제를 이해한다. 우리는 인터뷰 연구와 문헌 조사를 결합하여 실무자들이 실제로 DR을 어떻게 활용하는지 상세히 분석하였다. 이후 세 가지 주요 문제를 해결하기 위한 기술을 설계하였다. 첫째, 대표적인 DR 기법인 t-SNE와 UMAP의 만연한 오용을 완화한다. 이 기법들은 군집 및 클래스 분리도를 과도하게 부각시키기 때문에 심미적으로 보기 좋다는 이유로 부적절한 분석 작업에 잘못 사용된다. 기존 DR 평가 지표들은 클래스 레이블을 기반으로 하기 때문에, 클래스를 곧바로 군집의 정답으로 간주하며 이러한 편향을 더욱 증폭시킨다. 이에 우리는 실무자들이 군집 분석을 더 신뢰성 있게 지원하는 투영을 식별할 수 있도록 하는 새로운 평가 지표를 제안한다. 둘째, DR 하이퍼파라미터의 체리 피킹 문제를 해결한다. 적절한 평가 지표가 있더라도, DR 기법 선택과 하이퍼파라미터 최적화는 많은 시행착오를 요구하며, 이는 실무자들이 기본 설정에 의존하거나 임의로 하이퍼파라미터를 골라 쓰게 만든다. 우리는 데이터셋 특성을 자동으로 반영하여 최적의 투영을 효율적으로 탐색하는 적응형 최적화 워크플로우를 제안하여 분석가들이 체계적으로 최적화를 시행하도록 유도한다. 셋째, DR 투영에서의 상호작용 오류를 줄인다. 고차원 공간은 저차원 공간보다 훨씬 높은 자유도를 가지므로, 적절히 최적화된 DR 투영에서도 왜곡을 완전히 피할 수 없다. 이러한 왜곡은 DR 투영 상에서 사용자가 수행하는 브러싱과 같은 상호작용을 부정확하게 만든다. 우리는 사용자가 군집을 탐색할 때 왜곡을 보정해주는 Distortion-aware Brushing 기법을 제안하여, 사용자가 목표한 고차원 군집을 정확하게 포착할 수 있도록 지원한다.

    마지막으로, DR 기반 시각 분석의 신뢰성을 근본적으로 강화할 수 있는 미래 연구 방향을 제시한다. 이는 관련 담론의 활성화부터 최적의 DR 투영을 자동으로 선택하는 방향까지를 포괄한다. 본 논문은 신뢰가능한 시각 분석의 더 폭넓은 적용을 위한 기반을 마련하며 마무리된다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Dimensionality reduction (DR) is one of the most commonly used yet most easily misinterpreted tools in visual analytics. Visual analytics using DR can thus easily be unreliable: insights derived from analysis may not accurately reflect the underlying data, potentially leading to flawed knowledge and decision-making. This problem occurs because DR projections inherently cannot capture all characteristics of the original data, yet these limitations are often not adequately accounted for during analysis.

    In this dissertation, we enhance the reliability of visual analytics with DR. At the beginning, we understand the reliability challenges practitioners encounter when using DR for visual analytics. Following a human-centric approach, we detail how practitioners leverage DR in practice by combining interview studies and literature review. We then address three challenges by designing technical solutions. At first, we mitigate the common misuse of famous DR techniques: t-SNE and UMAP. These techniques are misused for unsuitable analytical tasks as they exaggerate cluster and class separability, which practitioners perceive as "aesthetically pleasing". We find that existing DR evaluation metrics that leverage class labels amplify this bias as they favor projections that well separate the classes. This is because these existing metrics assume classes as ground truth clusters. We propose new metrics that escape from this assumption, provoking that classes are not clusters, enabling practitioners to identify projections that more reliably support cluster analysis. Second, we address the prevalent cherry-picking of hyperparameters. Even with proper evaluation metrics, selecting appropriate DR techniques and optimizing hyperparameters to maximize metric scores requires extensive trial and error, leading practitioners to rely on default settings or cherry-pick hyperparameters. We introduce a dataset-adaptive optimization workflow that significantly reduces the computational cost of optimizing DR projections, motivating practitioners to avoid cherry-picking hyperparameters and instead systematically optimize projections. Third, we make interactions in DR projections less erroneous. High-dimensional space has a significantly higher degree of freedom compared to a low-dimensional space. Thus, even properly optimized DR projections cannot fully escape from distortions in representing the original structure. Such distortions cause interactions on DR projections, such as brushing, to erroneously reflect users’ intentions. We address this problem by proposing a new brushing technique called Distortion-aware brushing, which corrects distortions as users investigate clusters, thereby helping them precisely capture the high-dimensional clusters they target.

    Building on the insights from these studies, we outline future directions that can fundamentally enhance the reliability of DR-based visual analytics---spanning from facilitating relevant discourse to fully automating the selection of optimal DR projections. We conclude the thesis by discussing how our contributions lay the foundation for achieving more reliable visual analytics practices.
    번역하기

    Dimensionality reduction (DR) is one of the most commonly used yet most easily misinterpreted tools in visual analytics. Visual analytics using DR can thus easily be unreliable: insights derived from analysis may not accurately reflect the underlying ...

    Dimensionality reduction (DR) is one of the most commonly used yet most easily misinterpreted tools in visual analytics. Visual analytics using DR can thus easily be unreliable: insights derived from analysis may not accurately reflect the underlying data, potentially leading to flawed knowledge and decision-making. This problem occurs because DR projections inherently cannot capture all characteristics of the original data, yet these limitations are often not adequately accounted for during analysis.

    In this dissertation, we enhance the reliability of visual analytics with DR. At the beginning, we understand the reliability challenges practitioners encounter when using DR for visual analytics. Following a human-centric approach, we detail how practitioners leverage DR in practice by combining interview studies and literature review. We then address three challenges by designing technical solutions. At first, we mitigate the common misuse of famous DR techniques: t-SNE and UMAP. These techniques are misused for unsuitable analytical tasks as they exaggerate cluster and class separability, which practitioners perceive as "aesthetically pleasing". We find that existing DR evaluation metrics that leverage class labels amplify this bias as they favor projections that well separate the classes. This is because these existing metrics assume classes as ground truth clusters. We propose new metrics that escape from this assumption, provoking that classes are not clusters, enabling practitioners to identify projections that more reliably support cluster analysis. Second, we address the prevalent cherry-picking of hyperparameters. Even with proper evaluation metrics, selecting appropriate DR techniques and optimizing hyperparameters to maximize metric scores requires extensive trial and error, leading practitioners to rely on default settings or cherry-pick hyperparameters. We introduce a dataset-adaptive optimization workflow that significantly reduces the computational cost of optimizing DR projections, motivating practitioners to avoid cherry-picking hyperparameters and instead systematically optimize projections. Third, we make interactions in DR projections less erroneous. High-dimensional space has a significantly higher degree of freedom compared to a low-dimensional space. Thus, even properly optimized DR projections cannot fully escape from distortions in representing the original structure. Such distortions cause interactions on DR projections, such as brushing, to erroneously reflect users’ intentions. We address this problem by proposing a new brushing technique called Distortion-aware brushing, which corrects distortions as users investigate clusters, thereby helping them precisely capture the high-dimensional clusters they target.

    Building on the insights from these studies, we outline future directions that can fundamentally enhance the reliability of DR-based visual analytics---spanning from facilitating relevant discourse to fully automating the selection of optimal DR projections. We conclude the thesis by discussing how our contributions lay the foundation for achieving more reliable visual analytics practices.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • 1 Introduction 1
    • 1.1 Backgrounds and Motivation 1
    • 1.2 Thesis Contribution 3
    • 1.3 Prior Publications and Authorship 6
    • Abstract i
    • 1 Introduction 1
    • 1.1 Backgrounds and Motivation 1
    • 1.2 Thesis Contribution 3
    • 1.3 Prior Publications and Authorship 6
    • 2 Background: Dimensionality Reduction 8
    • 2.1 Overview 8
    • 2.2 Distortions in DR Projections 9
    • 2.3 Hyperparameters of DR Techniques 11
    • 3 Visual Analytics with Dimensionality Reduction 13
    • 3.1 Workflow Model 13
    • 3.1.1 Analysts and Machines 14
    • 3.1.2 Workflow Stages 16
    • 3.2 Impact of Our Contributions 18
    • 4 Identifying Reliability Challenges 20
    • 4.1 Introduction 20
    • 4.2 Related Work 21
    • 4.2.1 Investigation from Visualization Researchers 21
    • 4.2.2 Investigation from Domain Researchers 23
    • 4.3 Literature Review on the Visual Analytics Field 23
    • 4.3.1 Protocol 23
    • 4.3.2 Paper Search and Categorization Results 26
    • 4.3.3 Suitability of DR Techniques to Analytic Tasks 30
    • 4.3.4 Findings 33
    • 4.4 Literature Review Beyond the Visualization Field 35
    • 4.4.1 Paper Search Protocol 36
    • 4.4.2 Findings 38
    • 4.5 Interview Study 39
    • 4.5.1 Study Design 39
    • 4.5.2 Analysis Procedure 40
    • 4.5.3 Findings 40
    • 4.6 Summary of Identified Challenges 41
    • 4.7 Conclusion 42
    • 5 Classes are Not Clusters: Improving Label-based Evaluation of Dimensionality Reduction 43
    • 5.1 Introduction 43
    • 5.2 Related Works 45
    • 5.2.1 Evaluating DR Projections Using Class Labels 45
    • 5.2.2 Clustering Quality Metrics 47
    • 5.2.3 Evaluation of Cluster-Label Matching 47
    • 5.3 Proposed Label-Based Evaluation Process 49
    • 5.4 Adjusted Clustering Quality Metrics 49
    • 5.4.1 Axioms 50
    • 5.4.2 Protocols 53
    • 5.4.3 Adjusting Internal Clustering Quality Metrics 56
    • 5.4.4 Adjusting the Calinski-Harabasz Index 56
    • 5.4.5 Adjusting the Remaining IVMs 59
    • 5.4.6 Quantitative Evaluation 59
    • 5.5 Label-Trustworthiness and Label-Continuity 65
    • 5.5.1 Metric Design 65
    • 5.5.2 Selecting Internal Clustering Quality Metrics for Label-T&C 66
    • 5.5.3 Guidelines to Interpret Label-T&C 67
    • 5.5.4 Sensitivity Analysis 68
    • 5.5.5 Runtime Analysis 76
    • 5.6 Case Studies 77
    • 5.6.1 Examining the Effect of t-SNE Perplexity 77
    • 5.6.2 Analyzing DR Techniques’ Performance in Detail 80
    • 5.7 Conclusion 83
    • 6 Dataset-Adaptive Workflow for Accelerating Dimensionality Reduction Optimization 84
    • 6.1 Introduction 84
    • 6.2 Related Work 86
    • 6.2.1 Dataset-Adaptive Machine Learning 86
    • 6.2.2 Intrinsic Dimensionality Metrics 87
    • 6.3 Conventional Workflow for Finding Optimal DR Projections 88
    • 6.3.1 Workflow 89
    • 6.3.2 Problems 89
    • 6.4 Dataset-Adaptive Workflow 89
    • 6.4.1 Structural Complexity and Complexity Metrics 91
    • 6.4.2 Workflow 93
    • 6.5 Structural Complexity Metrics for Dataset-Adaptive Workflow 94
    • 6.5.1 Pairwise Distance Shift (Pds) 94
    • 6.5.2 Mutual Neighbor Consistency (Mnc) 97
    • 6.5.3 Pds+Mnc 102
    • 6.5.4 Implementation 102
    • 6.6 Experiment 1: Validity of Pds+Mnc 102
    • 6.6.1 Objectives and Study Design 102
    • 6.6.2 Results and Discussions 104
    • 6.7 Experiment 2: Suitability of Pds+Mnc for the Dataset-Adaptive Workflow 107
    • 6.7.1 Evaluation on Pretraining Regression Models 107
    • 6.7.2 Evaluation on Predicting Effective DR Techniques 109
    • 6.7.3 Evaluation on Early Terminating Optimization 111
    • 6.8 Experiment 3: Effectiveness of the Dataset-Adaptive Workflow 114
    • 6.8.1 Objectives and Study Design 114
    • 6.9 Discussions 116
    • 6.9.1 Leveraging the Tradeoffs between Predictive Power and Efficiency 116
    • 6.9.2 Complementing Our Structural Complexity Metrics 117
    • 6.9.3 Enhancing Reproducibility 117
    • 6.9.4 Exploring Additional Use Cases 117
    • 6.10 Conclusion 118
    • 7 Distortion-aware Brushing for Reliable Cluster Analysis in Dimensionality Reduction Projections 119
    • 7.1 Introduction 119
    • 7.2 Background and Related Work 122
    • 7.2.1 Interactive Points Relocation 122
    • 7.2.2 Brushing High-dimensional Data 122
    • 7.3 Design Objectives 124
    • 7.4 Distortion-aware brushing 126
    • 7.4.1 Workflow 126
    • 7.4.2 Features for Enhanced Usability 135
    • 7.4.3 Implementation 136
    • 7.5 Evaluating the Robustness of Distortion-Aware Brushing 136
    • 7.5.1 User Study 1: Robustness Against Distortions 136
    • 7.5.2 User Study 2: Robustness Against the Non-Triviality of High-dimensional Cluster Shape 146
    • 7.5.3 Post-Study Interview 149
    • 7.6 Use Case 1: Geospatial Cluster Analysis 151
    • 7.6.1 Procedure 152
    • 7.6.2 Scenario 153
    • 7.7 Use Case 2: Visual-Interactive Labeling 156
    • 7.7.1 Procedure 156
    • 7.7.2 Scenario 157
    • 7.7.3 Comparison with Existing Systems 159
    • 7.8 Discussions 160
    • 7.8.1 Comparison to Automatic Clustering Techniques 160
    • 7.8.2 Visual and Computational Scalability 161
    • 7.8.3 Reproducibility of Cluster Analysis 161
    • 7.8.4 Extension for Subspace Analysis 162
    • 7.8.5 Additional Usage Scenarios 163
    • 7.8.6 Limitations 163
    • 7.9 Conclusion 165
    • 8 Discussions 166
    • 8.1 What will be the Ideal Future Solution? 166
    • 8.1.1 Automating DR Configuration 166
    • 8.1.2 Facilitating Discourse 168
    • 8.1.3 Leveraging Mixed-Initiative Approaches 169
    • 8.1.4 Which Direction Should We Pursue? 169
    • 8.2 What are Reliability Challenges Beyond Faithfulness? 170
    • 8.2.1 The Lack of Interpretability 170
    • 8.2.2 Visual Ambiguity 171
    • 8.2.3 Instability 172
    • 8.2.4 Perceptual Misalignment 173
    • 8.3 How Can Visual Analytics Better Reflect Reality? 174
    • 8.3.1 Dataset Metadata 174
    • 8.3.2 Situated Analytics 175
    • 9 Conclusion 176
    • 9.1 Review of Thesis Contributions 176
    • 9.2 Summary of Thesis Impact 178
    • 9.3 Final Remarks 178
    • Abstract (Korean) 203
    • Acknowledgments 204
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼