최근 대규모 언어 모델(LLM)은 방대한 데이터를 바탕으로 뛰어난 추론 능력을 보여주며, 검색 엔진이나 코드 인터프리터와 같은 외부 도구를 활용하는 연구로 확장되고 있다. 그러나 기존 도...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17451387
서울 : 서울대학교 대학원, 2026
학위논문(석사) -- 서울대학교 대학원 , 협동과정인공지능전공 , 2026. 2
2026
한국어
LLM ; 도구 학습 ; 컨텍스트 엔지니어링 ; 로봇 액션
006.3
서울
iii, 40 ; 26 cm
지도교수: 장병탁
I804:11032-000000193828
0
상세조회0
다운로드최근 대규모 언어 모델(LLM)은 방대한 데이터를 바탕으로 뛰어난 추론 능력을 보여주며, 검색 엔진이나 코드 인터프리터와 같은 외부 도구를 활용하는 연구로 확장되고 있다. 그러나 기존 도...
최근 대규모 언어 모델(LLM)은 방대한 데이터를 바탕으로 뛰어난 추론 능력을 보여주며, 검색 엔진이나 코드 인터프리터와 같은 외부 도구를 활용하는 연구로 확장되고 있다. 그러나 기존 도구 학습은 주로 가상 환경의 소프트웨어적 도구에 국한되어, 로봇이 물리적 환경에서 수행하는 행동을 LLM의 도구 관점에서 해석하려는 시도는 상대적으로 미비하였다. 이에 본 논문은 3D 가상환경 내에서 로봇의 행동을 LLM이 호출할 수 있는 함수 형태의 외부 도구로 정의하고, 이를 통해 에이전트가 인간 사용자와의 다중 턴 대화를 통해 복잡한 가사 과업을 수행하는 체화된 추론 프레임워크를 제안한다. 본 연구는 시각적 인식의 한계를 보완하기 위해 시뮬레이터가 제공하는 메타데이터를 활용하여 LLM에 최적화된 텍스트 기반의 상태 추상화 및 공간정보 증강 방법론을 적용하였다. 특히 제한된 컨텍스트 윈도우 내에서 추론 효율성을 극대화하기 위해, 프롬프트를 환경 상태, 상호작용 이력, 도구 목록의세 가지 축으로 구조화하는 컨텍스트 엔지니어링 체계를 구축하다. TEACh 데이터셋을 기반으로 한 실험 결과, 다중 턴 상호작용 과업에서 객체 선별 시 의미적 연관성을 고려한 필터링 전략이 더 높은 정확도가 보임을 확인하였다. 또한, LLM 전용 기능인 Function Calling을 활용했을 때 텍스트 생성 방식보다 정확도가 향상됨을 보았으며, 전체 이력 요약과 최근 상세 로그를 결합한 하이브리드 전략의 가능성을 보았다. 이러한 결과는 LLM 외부 도구로써의 로봇 액션의 가능성을 보여준다.
다국어 초록 (Multilingual Abstract)
Recent Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities based on vast amounts of data, expanding into research that utilizes external tools such as search engines and code interpreters. However, existing tool learning ...
Recent Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities based on vast amounts of data, expanding into research that utilizes external tools such as search engines and code interpreters. However, existing tool learning has primarily been confined to software-based tools within virtual environments; consequently, attempts to interpret robot actions performed in physical environments from the perspective of LLM tools have been relatively scarce. To address this, we propose an embodied reasoning framework in which robot actions within a 3D virtual environment are defined as external tools in the form of functions callable by the LLM. This enables an agent to perform complex household tasks through multi-turn dialogue with human users. To compensate for the limitations of visual perception, we incorporate simulator-provided metadata and introduce text-based state abstraction and spatial-information augmentation methods optimized for LLMs. In particular, to maximize reasoning efficiency within a limited context window, we established a context engineering system that structures prompts along three axes: environmental state, interaction history, and tool list. Experimental results based on the TEACh dataset demonstrated that a filtering strategy considering semantic relevance during object selection yielded higher accuracy in multi-turn interaction tasks. Furthermore, we observed that utilizing the LLM-specific Function Calling feature improved accuracy compared to standard text generation methods. We also identified the potential of a hybrid strategy that combines a summary of the full history with recent detailed logs. These findings demonstrate the potential of integrating robot actions as external tools for LLMs.
목차 (Table of Contents)