RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Constructing SemiCon-MMLU: A Domain-Specific Benchmark for Evaluating LLMs and Prompt Strategies in Semiconductor Manufacturing = SemiCon-MMLU: 반도체 제조 도메인 특화 벤치마크 구축 및 LLM 프롬프트 전략 평가

    한글로보기

    https://www.riss.kr/link?id=T17449969

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large Language Models (LLMs) have demonstrated strong performance on general benchmarks; however, their effectiveness in specialized industrial domains such as
    semiconductor engineering has not been systematically evaluated. This study addresses this gap by developing SemiCon-MMLU, a domain-specific benchmark
    comprising 600 expert-validated multiple-choice questions spanning three functional areas: fundamental physics and device principles (Core), manufacturing processes
    (Fab), and system-level integration (Application).

    Using this benchmark, we evaluated eleven state-of-the-art LLMs, including GPT-4o, Claude-4-Sonnet, Gemini-2.5-Pro, and Llama-3.3-70B, under three prompting
    strategies: zero-shot, few-shot, and chain-of-thought (CoT). Contrary to findings in general reasoning tasks, our results reveal a “CoT Paradox”: chain-of-thought prompting substantially degraded performance (e.g., GPT-4o by roughly 25 percentage points), which we hypothesize is related to hallucination propagation in
    knowledge-intensive domains. Frontier models converged around 78–80% accuracy, notably below their performance on general benchmarks, highlighting a persistent
    domain-specific knowledge gap.

    Furthermore, we propose the Semiconductor Utility Score (SUS), a framework incorporating accuracy, security (on-premise deployability), and computational cost
    into deployment decisions. The SUS analysis demonstrates that model rankings shift considerably depending on organizational priorities: under security-prioritized scenarios, open-weight models such as Llama-3.3-70B achieve higher utility scores than proprietary models despite lower raw accuracy.

    This study contributes: (1) a publicly describable, expert-validated semiconductor domain benchmark; (2) empirical evidence that advanced prompting strategies
    can harm performance in specialized technical domains; and (3) a practical evaluation framework for industrial LLM deployment. By enabling efficient pre-screening
    of candidate models before costly RAG or fine-tuning investments, these findings offer guidance for semiconductor organizations seeking to adopt LLMs while balancing
    accuracy, security, and cost constraints.
    번역하기

    Large Language Models (LLMs) have demonstrated strong performance on general benchmarks; however, their effectiveness in specialized industrial domains such as semiconductor engineering has not been systematically evaluated. This study addresses this ...

    Large Language Models (LLMs) have demonstrated strong performance on general benchmarks; however, their effectiveness in specialized industrial domains such as
    semiconductor engineering has not been systematically evaluated. This study addresses this gap by developing SemiCon-MMLU, a domain-specific benchmark
    comprising 600 expert-validated multiple-choice questions spanning three functional areas: fundamental physics and device principles (Core), manufacturing processes
    (Fab), and system-level integration (Application).

    Using this benchmark, we evaluated eleven state-of-the-art LLMs, including GPT-4o, Claude-4-Sonnet, Gemini-2.5-Pro, and Llama-3.3-70B, under three prompting
    strategies: zero-shot, few-shot, and chain-of-thought (CoT). Contrary to findings in general reasoning tasks, our results reveal a “CoT Paradox”: chain-of-thought prompting substantially degraded performance (e.g., GPT-4o by roughly 25 percentage points), which we hypothesize is related to hallucination propagation in
    knowledge-intensive domains. Frontier models converged around 78–80% accuracy, notably below their performance on general benchmarks, highlighting a persistent
    domain-specific knowledge gap.

    Furthermore, we propose the Semiconductor Utility Score (SUS), a framework incorporating accuracy, security (on-premise deployability), and computational cost
    into deployment decisions. The SUS analysis demonstrates that model rankings shift considerably depending on organizational priorities: under security-prioritized scenarios, open-weight models such as Llama-3.3-70B achieve higher utility scores than proprietary models despite lower raw accuracy.

    This study contributes: (1) a publicly describable, expert-validated semiconductor domain benchmark; (2) empirical evidence that advanced prompting strategies
    can harm performance in specialized technical domains; and (3) a practical evaluation framework for industrial LLM deployment. By enabling efficient pre-screening
    of candidate models before costly RAG or fine-tuning investments, these findings offer guidance for semiconductor organizations seeking to adopt LLMs while balancing
    accuracy, security, and cost constraints.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델(LLM)은 일반적인 벤치마크에서 우수한 성능을 보여왔으나, 반도체 공학과 같은 전문 산업 도메인에서의 효용성은 체계적으로 평가되지 않았다. 본 연구는
    이러한 한계를 해결하기 위해 반도체 도메인 특화 벤치마크인 SemiCon-MMLU를 개발하였다. SemiCon-MMLU는 전문가 검증을 거친 600개의 객관식 문항으로 구성되
    며, 기초 물리 및 소자 원리(Core), 제조 공정(Fab), 시스템 수준 통합(Application)의 세 가지 기능 영역을 포괄한다.

    본 벤치마크를 활용하여 GPT-4o, Claude-4-Sonnet, Gemini-2.5-Pro, Llama-3.3-70B를 포함한 11개의 최신 LLM을 제로샷(zero-shot), 퓨샷(few-shot), 사고의 연쇄 (Chain-of-Thought, CoT) 등 세 가지 프롬프트 전략 하에서 평가하였다. 일반적인 추론 과제에서의 기존 연구 결과와 달리, 본 연구에서는 “CoT 역설(CoT Paradox)”을 발견하였다: 사고의 연쇄 프롬프팅이 오히려 성능을 크게 저하시켰으며(예: GPT-4o에서 약 25 퍼센트포인트 하락), 이는 지식 집약적 도메인에서의 환각 전파(hallucination propagation)와 관련이 있는 것으로 추정된다. 최상위 모델들은 78–80% 수준의 정확도에서 수렴하였으며, 이는 일반 벤치마크 대비 현저히 낮은 수치로서 도메인 특화 지식의 격차가 지속적으로 존재함을 시사한다.

    나아가 본 연구는 정확도, 보안성(온프레미스 배포 가능 여부), 연산 비용을 통합한 다기준 평가 프레임워크인 반도체 유용성 점수(Semiconductor Utility Score, SUS)를
    제안한다. SUS 분석 결과, 가중치 설정에 따라 모델 순위가 상당히 변동함을 확인하였다: 보안을 우선시하는 시나리오에서는 Llama-3.3-70B와 같은 오픈웨이트 모델이
    원시 정확도가 낮음에도 불구하고 상용 모델보다 높은 유용성 점수를 달성하였다. 이는 조직의 우선순위에 따라 최적 모델 선택이 달라질 수 있음을 시사한다.

    본 연구의 기여는 다음과 같다: (1) 공개 가능한 방법론에 기반한 전문가 검증 반도체 도메인 벤치마크 개발, (2) 전문 기술 도메인에서 고급 프롬프팅 전략이 오히려 성능을
    저하시킬 수 있다는 실증적 증거 제시, (3) 산업 현장의 LLM 도입을 위한 실용적 평가 프레임워크 제안. 고비용의 RAG 또는 파인튜닝 투자에 앞서 후보 모델을 효율적으로
    사전 선별할 수 있도록 함으로써, 본 연구 결과는 정확도, 보안, 비용의 균형을 고려하며 LLM 도입을 모색하는 반도체 기업들에게 실질적인 지침을 제공한다.
    번역하기

    대규모 언어 모델(LLM)은 일반적인 벤치마크에서 우수한 성능을 보여왔으나, 반도체 공학과 같은 전문 산업 도메인에서의 효용성은 체계적으로 평가되지 않았다. 본 연구는 이러한 한계를 해...

    대규모 언어 모델(LLM)은 일반적인 벤치마크에서 우수한 성능을 보여왔으나, 반도체 공학과 같은 전문 산업 도메인에서의 효용성은 체계적으로 평가되지 않았다. 본 연구는
    이러한 한계를 해결하기 위해 반도체 도메인 특화 벤치마크인 SemiCon-MMLU를 개발하였다. SemiCon-MMLU는 전문가 검증을 거친 600개의 객관식 문항으로 구성되
    며, 기초 물리 및 소자 원리(Core), 제조 공정(Fab), 시스템 수준 통합(Application)의 세 가지 기능 영역을 포괄한다.

    본 벤치마크를 활용하여 GPT-4o, Claude-4-Sonnet, Gemini-2.5-Pro, Llama-3.3-70B를 포함한 11개의 최신 LLM을 제로샷(zero-shot), 퓨샷(few-shot), 사고의 연쇄 (Chain-of-Thought, CoT) 등 세 가지 프롬프트 전략 하에서 평가하였다. 일반적인 추론 과제에서의 기존 연구 결과와 달리, 본 연구에서는 “CoT 역설(CoT Paradox)”을 발견하였다: 사고의 연쇄 프롬프팅이 오히려 성능을 크게 저하시켰으며(예: GPT-4o에서 약 25 퍼센트포인트 하락), 이는 지식 집약적 도메인에서의 환각 전파(hallucination propagation)와 관련이 있는 것으로 추정된다. 최상위 모델들은 78–80% 수준의 정확도에서 수렴하였으며, 이는 일반 벤치마크 대비 현저히 낮은 수치로서 도메인 특화 지식의 격차가 지속적으로 존재함을 시사한다.

    나아가 본 연구는 정확도, 보안성(온프레미스 배포 가능 여부), 연산 비용을 통합한 다기준 평가 프레임워크인 반도체 유용성 점수(Semiconductor Utility Score, SUS)를
    제안한다. SUS 분석 결과, 가중치 설정에 따라 모델 순위가 상당히 변동함을 확인하였다: 보안을 우선시하는 시나리오에서는 Llama-3.3-70B와 같은 오픈웨이트 모델이
    원시 정확도가 낮음에도 불구하고 상용 모델보다 높은 유용성 점수를 달성하였다. 이는 조직의 우선순위에 따라 최적 모델 선택이 달라질 수 있음을 시사한다.

    본 연구의 기여는 다음과 같다: (1) 공개 가능한 방법론에 기반한 전문가 검증 반도체 도메인 벤치마크 개발, (2) 전문 기술 도메인에서 고급 프롬프팅 전략이 오히려 성능을
    저하시킬 수 있다는 실증적 증거 제시, (3) 산업 현장의 LLM 도입을 위한 실용적 평가 프레임워크 제안. 고비용의 RAG 또는 파인튜닝 투자에 앞서 후보 모델을 효율적으로
    사전 선별할 수 있도록 함으로써, 본 연구 결과는 정확도, 보안, 비용의 균형을 고려하며 LLM 도입을 모색하는 반도체 기업들에게 실질적인 지침을 제공한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ix
    • List of Tables xi
    • List of Figures xv
    • Abstract i
    • Contents ix
    • List of Tables xi
    • List of Figures xv
    • Chapter 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Research Problem and Motivation 4
    • 1.2.1 Rationale for Pre-trained Model Evaluation 6
    • 1.3 Research Objectives and Scope 7
    • 1.4 Research Contributions 9
    • 1.5 Structure of the Thesis 11
    • Chapter 2 Literature Review 12
    • 2.1 Overview 12
    • 2.2 Semiconductor Knowledge Structure 12
    • 2.2.1 Academic and Theoretical Frameworks 13
    • 2.2.2 Industrial and Practical Taxonomies 13
    • 2.2.3 Implications for Benchmarking 14
    • 2.3 General-Purpose LLM Evaluation Benchmarks 15
    • 2.3.1 Limitations of Broad Knowledge Benchmarks 15
    • 2.3.2 Reasoning-Focused Benchmarks 16
    • 2.4 Domain-Specific Benchmark Development 17
    • 2.4.1 Methodological Lessons from Other Fields 17
    • 2.4.2 Existing Efforts in Hardware Engineering 18
    • 2.5 Prompt Engineering Strategies 19
    • 2.5.1 In-Context Learning and Chain-of-Thought 19
    • 2.5.2 Cost-Benefit Trade-offs 21
    • 2.6 LLM Applications in Semiconductor Contexts 22
    • 2.7 Research Gap and Motivation 24
    • 2.8 Chapter Summary 26
    • Chapter 3 Methodology 28
    • 3.1 Overview 28
    • 3.2 Benchmark Design 29
    • 3.2.1 Design Principles and Framework Selection 29
    • 3.2.2 Data Sources and Collection Strategy 32
    • 3.2.3 Domain Taxonomy and Classification 37
    • 3.2.4 Item Development and Validation Process 41
    • 3.2.5 Difficulty Stratification and Cognitive Complexity 47
    • 3.2.6 Benchmark Quality Analysis 49
    • 3.3 Chapter Summary 52
    • Chapter 4 Experiment 55
    • 4.1 Target Models 55
    • 4.1.1 Selection Criteria 55
    • 4.1.2 Selected Models 56
    • 4.2 Prompt Engineering Strategies 57
    • 4.3 Experimental Implementation 58
    • 4.3.1 Computing Environment 59
    • 4.3.2 Hyperparameter Control 59
    • 4.3.3 Response Processing 59
    • 4.4 Evaluation Metrics 60
    • 4.4.1 Quantitative Accuracy 60
    • 4.4.2 Prompt Efficiency Index (PEI) 60
    • 4.4.3 Reasoning Depth Metric 61
    • 4.4.4 Semiconductor Utility Score (SUS) 61
    • 4.5 Limitations 62
    • Chapter 5 Results 64
    • 5.1 Overview 64
    • 5.2 Overall Model Performance 65
    • 5.3 Fine-grained Performance Analysis 66
    • 5.3.1 Performance by Functional Domain 66
    • 5.3.2 Performance by Data Source 68
    • 5.3.3 Performance by Difficulty Level 69
    • 5.4 Prompt Strategy Analysis 70
    • 5.4.1 Strategy Comparison 70
    • 5.4.2 Prompt Efficiency Analysis 71
    • 5.4.3 Domain-Strategy Interaction 71
    • 5.5 Model Tier Analysis 72
    • 5.6 Semiconductor Utility Score (SUS) Analysis 72
    • 5.6.1 SUS Ranking 72
    • 5.6.2 Scenario-Dependent Ranking 72
    • 5.6.3 Sensitivity Analysis 73
    • 5.7 Summary of Results 74
    • Chapter 6 Discussion 76
    • 6.1 Interpretation of Performance Patterns 76
    • 6.1.1 The ∼80% Performance Ceiling 76
    • 6.1.2 Domain Hierarchy: Why Application Is Most Challenging 77
    • 6.1.3 Differential Generalization: The Llama Inverse Pattern 78
    • 6.2 The CoT Paradox: When Reasoning Hurts Performance 79
    • 6.2.1 Hallucination Propagation Hypothesis 80
    • 6.2.2 Domain-Specific Evidence 82
    • 6.2.3 Implications for Prompt Engineering 82
    • 6.3 Prompt Strategy Implications 83
    • 6.3.1 Few-shot Divergence 83
    • 6.3.2 Cost-Benefit Analysis 84
    • 6.4 Industrial Utility: A Multi-Criteria Perspective 84
    • 6.4.1 Trade-off Structure in Model Selection 84
    • 6.4.2 The Security Imperative 85
    • 6.4.3 Sensitivity to Weight Configuration 85
    • 6.4.4 Strategic Implications 87
    • 6.4.5 Pre-RAG Screening Value 88
    • 6.5 Benchmark Validation 89
    • 6.5.1 Difficulty Calibration 89
    • 6.5.2 Source Differentiation 89
    • 6.6 Limitations 90
    • 6.6.1 Unimodal Evaluation 90
    • 6.6.2 Absence of Operational Data 90
    • 6.6.3 Static Knowledge Snapshot 91
    • 6.6.4 Single-Run Evaluation 91
    • 6.6.5 SUS Weight Subjectivity 91
    • 6.6.6 Temporal Scope of Model Evaluation 92
    • 6.7 Future Work 92
    • 6.7.1 Multimodal Benchmark Extension 92
    • 6.7.2 Temporal Knowledge Assessment 92
    • 6.7.3 Domain-Specific Fine-Tuning 92
    • 6.7.4 Longitudinal Model Tracking 93
    • 6.7.5 Operational Integration Studies 93
    • 6.8 Chapter Summary 93
    • Chapter 7 Conclusion 95
    • 7.1 Summary of Research 95
    • 7.1.1 Key Findings 96
    • 7.2 Research Contributions 97
    • 7.2.1 Methodological Contribution 97
    • 7.2.2 Empirical Contribution 98
    • 7.2.3 Practical Contribution 98
    • 7.3 Limitations 99
    • 7.4 Future Work 99
    • 7.5 Closing Remarks 100
    • Chapter A Example Questions from SemiCon-MMLU 101
    • A.1 Public Examination Questions 101
    • A.2 Expert-Authored Questions (Proprietary) 103
    • A.3 LLM-Augmented Questions 104
    • Chapter B Statistical Validation 107
    • B.1 Benchmark Dataset Statistics 107
    • B.1.1 Answer Position Distribution 107
    • B.1.2 Domain and Difficulty Distribution 108
    • B.1.3 Data Source Distribution 108
    • B.2 Model Performance Statistics 109
    • B.2.1 Overall Accuracy with Confidence Intervals 109
    • B.3 Pairwise Model Comparisons (McNemar Test) 110
    • B.3.1 Frontier Model Comparisons 110
    • B.4 Difficulty Level Analysis (χ2 Tests) 111
    • B.4.1 Contingency Table Construction 111
    • B.4.2 Results by Model 111
    • B.5 Prompt Strategy Analysis 112
    • B.5.1 Accuracy and Token Usage by Strategy 112
    • B.5.2 Token Cost Multipliers 112
    • Chapter C SUS Framework Details 114
    • C.1 Score Definition and Calculations 114
    • C.1.1 Score Definition 114
    • C.1.2 Normalization Parameters 115
    • C.1.3 Calculated SUS Values 115
    • C.1.4 Sensitivity Analysis 115
    • C.2 Complete Model Profiles 116
    • Chapter D Experimental Environment 118
    • D.1 Software Environment 118
    • D.2 Model Configurations 118
    • D.3 Inference Parameters 119
    • Bibliography 120
    • 국문초록 128
    • 감사의 글 130
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼