RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Automatic Routing Between Symbol-Prediction and Cloze Scoring = 다지선다형 LLM 평가에서 Symbol과 Cloze 채점 방식 간 자동 라우팅

    한글로보기

    https://www.riss.kr/link?id=T17449861

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are systematically attributable to task characteristics: natural language continuation benefits from likelihood scoring, whereas explicit comparison is better suited to symbol-based selection. These trends are consistent across various decoder-based LLMs, indicating model-agnostic effects. To address these inconsistencies, a dynamic format-alignment strategy is introduced that employs a lightweight classifier trained on latent model-preference signals. In contrast to human-designed heuristics, which often degrade performance, this approach uses model-generated signals to determine the optimal format for each problem instance. The proposed method achieves substantial and consistent improvements in zero-shot accuracy across reasoning and knowledge benchmarks, better revealing the models' latent capabilities.
    번역하기

    Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are systematically attributable to task characteristics: natural language continu...

    Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are systematically attributable to task characteristics: natural language continuation benefits from likelihood scoring, whereas explicit comparison is better suited to symbol-based selection. These trends are consistent across various decoder-based LLMs, indicating model-agnostic effects. To address these inconsistencies, a dynamic format-alignment strategy is introduced that employs a lightweight classifier trained on latent model-preference signals. In contrast to human-designed heuristics, which often degrade performance, this approach uses model-generated signals to determine the optimal format for each problem instance. The proposed method achieves substantial and consistent improvements in zero-shot accuracy across reasoning and knowledge benchmarks, better revealing the models' latent capabilities.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    다지선다 과제에서 대규모 언어 모델(LLM)의 성능은 Symbol 평가 형식과 Cloze 평가 형식 사이에서 현저하게 달라진다. 관찰되는 차이는 벤치마크 특성에 의해 체계적으로 설명된다. 즉, 자연어를 이어 쓰는(continuation) 유형의 벤치마크는 Cloze 방식에서 이점을 얻는 반면, 명시적 비교(explicit comparison)는 Symbol 방식에 더 적합하다. 이러한 경향은 decoder 기반 다양한 LLM 전반에서 일관되게 나타난다. 이에 잠재적 모델 선호 신호에 기반해 학습된 경량 분류기(classifier)를 활용하는 동적 format-alignment 전략을 제안한다. 성능을 저하시키는 경우가 잦은 인간 설계 휴리스틱과 달리, 이 접근법은 각 문제 인스턴스마다 최적의 형식을 결정하기 위해 모델이 생성한 신호를 사용한다. 제안 방법은 추론 및 지식 벤치마크에서 zero-shot 정확도를 크게 그리고 일관되게 향상시키며, 모델의 잠재 역량을 더 잘 드러낸다.
    번역하기

    다지선다 과제에서 대규모 언어 모델(LLM)의 성능은 Symbol 평가 형식과 Cloze 평가 형식 사이에서 현저하게 달라진다. 관찰되는 차이는 벤치마크 특성에 의해 체계적으로 설명된다. 즉, 자연어...

    다지선다 과제에서 대규모 언어 모델(LLM)의 성능은 Symbol 평가 형식과 Cloze 평가 형식 사이에서 현저하게 달라진다. 관찰되는 차이는 벤치마크 특성에 의해 체계적으로 설명된다. 즉, 자연어를 이어 쓰는(continuation) 유형의 벤치마크는 Cloze 방식에서 이점을 얻는 반면, 명시적 비교(explicit comparison)는 Symbol 방식에 더 적합하다. 이러한 경향은 decoder 기반 다양한 LLM 전반에서 일관되게 나타난다. 이에 잠재적 모델 선호 신호에 기반해 학습된 경량 분류기(classifier)를 활용하는 동적 format-alignment 전략을 제안한다. 성능을 저하시키는 경우가 잦은 인간 설계 휴리스틱과 달리, 이 접근법은 각 문제 인스턴스마다 최적의 형식을 결정하기 위해 모델이 생성한 신호를 사용한다. 제안 방법은 추론 및 지식 벤치마크에서 zero-shot 정확도를 크게 그리고 일관되게 향상시키며, 모델의 잠재 역량을 더 잘 드러낸다.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • Chapter 2. Related Works 4
    • Chapter 3. Symbol vs. Cloze: A Baseline 6
    • Chapter 4. Evaluation via Model-Preferred Formats 12
    • Chapter 5. Evaluation Results 19
    • Chapter 1. Introduction 1
    • Chapter 2. Related Works 4
    • Chapter 3. Symbol vs. Cloze: A Baseline 6
    • Chapter 4. Evaluation via Model-Preferred Formats 12
    • Chapter 5. Evaluation Results 19
    • Chapter 6. Conclusion 23
    • Acknowledgements 24
    • Bibliography 25
    • Appendix 31
    • Abstract in Korean 54
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼