연구자들은 대규모 언어 모델(Large Language Model, LLM)의 응답에 반영된 성격 특성과 가치를 측정하기 위해 기존의 심리측정 설문지(예: BFI, PVQ)를 적용해왔다. 그러나 이러한 인간을 위해 설계된...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
연구자들은 대규모 언어 모델(Large Language Model, LLM)의 응답에 반영된 성격 특성과 가치를 측정하기 위해 기존의 심리측정 설문지(예: BFI, PVQ)를 적용해왔다. 그러나 이러한 인간을 위해 설계된...
연구자들은 대규모 언어 모델(Large Language Model, LLM)의 응답에 반영된 성격 특성과 가치를 측정하기 위해 기존의 심리측정 설문지(예: BFI, PVQ)를 적용해왔다. 그러나 이러한 인간을 위해 설계된 설문지를 LLM에 적용하는 것에 대한 우려가 제기되어 왔다. 그 중 하나는 생태학적 타당도의 부족인데, 이는 LLM이 사용자 질의에 응답하는 실제 상황을 설문 문항이 얼마나 적절히 반영하는지에 관한 것이다. 하지만 기존 설문지와 생태학적으로 타당한 설문지가 결과에서 어떻게 다른지, 그리고 이러한 차이가 어떤 통찰을 제공하는지는 명확하지 않은 상태이다.
본 논문에서 우리는 두 유형의 설문지에 대한 포괄적인 비교 분석을 수행한다. 우리의 분석은 기존 설문지가 (1) 생태학적으로 타당한 설문지와 상당히 다른 LLM 프로파일을 산출하며, 사용자 질의 맥락에서 표현되는 심리적 특성에서 벗어나고, (2) 안정적인 측정을 위한 문항 수가 불충분하며, (3) LLM이 안정적인 구성개념을 지니고 있다는 오해를 불러일으키고, (4) 페르소나가 부여된 LLM에 대해 과장된 프로파일을 산출함을 밝힌다. 전반적으로 우리의 연구는 LLM에 대한 기존 심리학적 설문지 사용에 경계할 것을 권고한다.
다국어 초록 (Multilingual Abstract)
Researchers have applied established psychometric questionnaires (e.g., BFI, PVQ) to measure the personality traits and values reflected in the responses of Large Language Models (LLMs). However, concerns have been raised about applying these human-de...
Researchers have applied established psychometric questionnaires (e.g., BFI, PVQ) to measure the personality traits and values reflected in the responses of Large Language Models (LLMs). However, concerns have been raised about applying these human-designed questionnaires to LLMs. One such concern is their lack of ecological validity—the extent to which survey questions adequately reflect and resemble real-world contexts in which LLMs generate texts in response to user queries. However, it remains unclear how established questionnaires and ecologically valid questionnaires differ in their outcomes, and what insights these differences may provide.
In this paper, we conduct a comprehensive comparative analysis of the two types of questionnaires. Our analysis reveals that established questionnaires (1) yield substantially different profiles of LLMs from ecologically valid ones, deviating from the psychological characteristics expressed in the context of user queries, (2) suffer from insufficient items for stable measurement, (3) create misleading impressions that LLMs possess stable constructs, and (4) yield exaggerated profiles for persona-prompted LLMs. Overall, our work cautions against the use of established psychological questionnaires for LLMs.
목차 (Table of Contents)