RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Scaffolding Deliberation: A Human-in-the-Loop Framework for Guiding LLM Reasoning = 스캐폴딩 기반 인간 참여형 LLM 추론 지도 프레임워크

    한글로보기

    https://www.riss.kr/link?id=T17450907

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large Language Models (LLMs) perform well on complex reasoning tasks but
    remain vulnerable to error propagation, hallucination, and brittle multi-step inference. While inference-time techniques such as Chain-of-Thought, Self-Consistency,
    and Tree-of-Thought improve accuracy, they lack external supervision and often fail
    once incorrect assumptions arise. We propose an inference-time Human-in-the-Loop
    (HITL) reasoning framework that integrates reward-based monitoring, reflection, and
    targeted feedback to guide reasoning without retraining or increasing model size.
    Our framework uses either a Process Reward Model (PRM) or an LLM-based
    reward evaluator to assess intermediate reasoning steps and trigger intervention only
    when necessary. To address the limitations of step-level rewards, we introduce a structured feedback schema that separates error detection from correction, enabling both
    local and global reasoning repair. We systematically analyze the effects of feedback
    quality, reward models, thresholds, and self-consistency on accuracy, computational
    cost, and human workload.
    Experiments on challenging GSM8K subsets show that HITL consistently outperforms baseline inference. Stronger reasoning models benefit more from HITL, higherquality feedback reduces intervention frequency, and self-consistency further lowers
    human workload while controlling token growth. Moreover, domain-general LLMs
    can serve as effective reward and feedback models even without pretrained PRMs,
    highlighting the practicality of HITL in new domains.
    번역하기

    Large Language Models (LLMs) perform well on complex reasoning tasks but remain vulnerable to error propagation, hallucination, and brittle multi-step inference. While inference-time techniques such as Chain-of-Thought, Self-Consistency, and Tree-of-T...

    Large Language Models (LLMs) perform well on complex reasoning tasks but
    remain vulnerable to error propagation, hallucination, and brittle multi-step inference. While inference-time techniques such as Chain-of-Thought, Self-Consistency,
    and Tree-of-Thought improve accuracy, they lack external supervision and often fail
    once incorrect assumptions arise. We propose an inference-time Human-in-the-Loop
    (HITL) reasoning framework that integrates reward-based monitoring, reflection, and
    targeted feedback to guide reasoning without retraining or increasing model size.
    Our framework uses either a Process Reward Model (PRM) or an LLM-based
    reward evaluator to assess intermediate reasoning steps and trigger intervention only
    when necessary. To address the limitations of step-level rewards, we introduce a structured feedback schema that separates error detection from correction, enabling both
    local and global reasoning repair. We systematically analyze the effects of feedback
    quality, reward models, thresholds, and self-consistency on accuracy, computational
    cost, and human workload.
    Experiments on challenging GSM8K subsets show that HITL consistently outperforms baseline inference. Stronger reasoning models benefit more from HITL, higherquality feedback reduces intervention frequency, and self-consistency further lowers
    human workload while controlling token growth. Moreover, domain-general LLMs
    can serve as effective reward and feedback models even without pretrained PRMs,
    highlighting the practicality of HITL in new domains.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델은 복잡한 추론 과제에서 우수한 성능을 보이지만, 여전히 오류
    전파, 환각, 그리고 취약한 다단계 추론 문제에 취약하다. Chain-of-Thought, SelfConsistency, Tree-of-Thought와 같은 추론 시점 기법들은 정확도를 향상시키지만,
    외부 감독 없이 작동하기 때문에 잘못된 가정이 도입될 경우 쉽게 실패한다. 본 연
    구에서는 모델 재학습이나 파라미터 확장 없이, 보상 기반 모니터링, 반성, 그리고
    선택적 피드백을 결합한 추론 시점 Human-in-the-Loop 추론 프레임워크를 제안한다.
    제안하는 프레임워크는 Process Reward Model(PRM) 또는 LLM 기반 보상 평
    가기를 사용하여 중간 추론 단계를 평가하고, 필요할 때에만 개입을 유도한다. 또한
    오류 탐지와 교정을 분리한 구조화된 피드백 스키마를 도입하여 국소적 및 전역
    적 추론 복구를 가능하게 하며, 피드백 품질, 보상 모델, 임계값, self-consistency가
    정확도, 계산 비용, 그리고 인간 개입 부담에 미치는 영향을 분석한다.
    GSM8K의 도전적인 하위 데이터셋에 대한 실험 결과, HITL 추론은 기본 추론
    방식 대비 일관되게 우수한 성능을 보였다. 특히 더 강력한 추론 모델일수록 HITL의
    효과가 커지며, 고품질 피드백은 개입 빈도를 감소시키고, self-consistency는 토큰
    사용량의 증가를 억제하면서 인간의 개입 부담을 더욱 줄인다. 또한 사전 학습된
    PRM이 없는 환경에서도 범용 LLM이 보상 및 피드백 모델로 효과적으로 활용될 수
    있음을 보여주며, 새로운 도메인에서의 HITL 적용 가능성을 제시한다.
    번역하기

    대규모 언어 모델은 복잡한 추론 과제에서 우수한 성능을 보이지만, 여전히 오류 전파, 환각, 그리고 취약한 다단계 추론 문제에 취약하다. Chain-of-Thought, SelfConsistency, Tree-of-Thought와 같은 추...

    대규모 언어 모델은 복잡한 추론 과제에서 우수한 성능을 보이지만, 여전히 오류
    전파, 환각, 그리고 취약한 다단계 추론 문제에 취약하다. Chain-of-Thought, SelfConsistency, Tree-of-Thought와 같은 추론 시점 기법들은 정확도를 향상시키지만,
    외부 감독 없이 작동하기 때문에 잘못된 가정이 도입될 경우 쉽게 실패한다. 본 연
    구에서는 모델 재학습이나 파라미터 확장 없이, 보상 기반 모니터링, 반성, 그리고
    선택적 피드백을 결합한 추론 시점 Human-in-the-Loop 추론 프레임워크를 제안한다.
    제안하는 프레임워크는 Process Reward Model(PRM) 또는 LLM 기반 보상 평
    가기를 사용하여 중간 추론 단계를 평가하고, 필요할 때에만 개입을 유도한다. 또한
    오류 탐지와 교정을 분리한 구조화된 피드백 스키마를 도입하여 국소적 및 전역
    적 추론 복구를 가능하게 하며, 피드백 품질, 보상 모델, 임계값, self-consistency가
    정확도, 계산 비용, 그리고 인간 개입 부담에 미치는 영향을 분석한다.
    GSM8K의 도전적인 하위 데이터셋에 대한 실험 결과, HITL 추론은 기본 추론
    방식 대비 일관되게 우수한 성능을 보였다. 특히 더 강력한 추론 모델일수록 HITL의
    효과가 커지며, 고품질 피드백은 개입 빈도를 감소시키고, self-consistency는 토큰
    사용량의 증가를 억제하면서 인간의 개입 부담을 더욱 줄인다. 또한 사전 학습된
    PRM이 없는 환경에서도 범용 LLM이 보상 및 피드백 모델로 효과적으로 활용될 수
    있음을 보여주며, 새로운 도메인에서의 HITL 적용 가능성을 제시한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • Chapter 1 Introduction 1
    • Chapter 2 Background 3
    • Section 2.1 LLM Hallucination and Flaw Repetition 3
    • Abstract i
    • Contents ii
    • Chapter 1 Introduction 1
    • Chapter 2 Background 3
    • Section 2.1 LLM Hallucination and Flaw Repetition 3
    • Section 2.2 Scaffolding in reasoning 4
    • Section 2.3 Key Design Questions for HITL Reasoning 5
    • Subsection 2.3.1 When should humans intervene? 5
    • Subsection 2.3.2 What should humans provide? 5
    • Chapter 3 Main Architecture 6
    • Section 3.1 Self-Consistency with Human-in-the-Loop 6
    • Section 3.2 Reward Model 8
    • Subsection 3.2.1 Process Reward Model 8
    • Subsection 3.2.2 LLM-based Reward Model 11
    • Section 3.3 Reflection Activity 13
    • Section 3.4 Feedback 15
    • Subsection 3.4.1 Feedback Schema 15
    • Subsection 3.4.2 Feedback Types 16
    • Chapter 4 Experiments 18
    • Section 4.1 Dataset & Model 18
    • Subsection 4.1.1 Data Selection 18
    • Subsection 4.1.2 Erroneous Samples in the gsm8k Dataset 20
    • Section 4.2 Experimental Results 20
    • Subsection 4.2.1 Reasoning Model 20
    • Subsection 4.2.2 Feedback 25
    • Subsection 4.2.3 Reward 31
    • Subsection 4.2.4 Self-Consistency 35
    • Chapter 5 Conclusion Future Work 38
    • Section 5.1 Conclusion 38
    • Section 5.2 Future Work 40
    • Bibliography 43
    • Abstract (In Korean) 45
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼