RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    도구 증강 LLM 기반 가상 체화 에이전트의 멀티턴 과제 수행 = Tool-Augmented LLMs for Multi-turn Task Execution and Reasoning in Virtual Embodied Agents

    한글로보기

    https://www.riss.kr/link?id=T17451387

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 대규모 언어 모델(LLM)은 방대한 데이터를 바탕으로 뛰어난 추론 능력을 보여주며, 검색 엔진이나 코드 인터프리터와 같은 외부 도구를 활용하는 연구로 확장되고 있다. 그러나 기존 도구 학습은 주로 가상 환경의 소프트웨어적 도구에 국한되어, 로봇이 물리적 환경에서 수행하는 행동을 LLM의 도구 관점에서 해석하려는 시도는 상대적으로 미비하였다. 이에 본 논문은 3D 가상환경 내에서 로봇의 행동을 LLM이 호출할 수 있는 함수 형태의 외부 도구로 정의하고, 이를 통해 에이전트가 인간 사용자와의 다중 턴 대화를 통해 복잡한 가사 과업을 수행하는 체화된 추론 프레임워크를 제안한다. 본 연구는 시각적 인식의 한계를 보완하기 위해 시뮬레이터가 제공하는 메타데이터를 활용하여 LLM에 최적화된 텍스트 기반의 상태 추상화 및 공간정보 증강 방법론을 적용하였다. 특히 제한된 컨텍스트 윈도우 내에서 추론 효율성을 극대화하기 위해, 프롬프트를 환경 상태, 상호작용 이력, 도구 목록의세 가지 축으로 구조화하는 컨텍스트 엔지니어링 체계를 구축하다. TEACh 데이터셋을 기반으로 한 실험 결과, 다중 턴 상호작용 과업에서 객체 선별 시 의미적 연관성을 고려한 필터링 전략이 더 높은 정확도가 보임을 확인하였다. 또한, LLM 전용 기능인 Function Calling을 활용했을 때 텍스트 생성 방식보다 정확도가 향상됨을 보았으며, 전체 이력 요약과 최근 상세 로그를 결합한 하이브리드 전략의 가능성을 보았다. 이러한 결과는 LLM 외부 도구로써의 로봇 액션의 가능성을 보여준다.
    번역하기

    최근 대규모 언어 모델(LLM)은 방대한 데이터를 바탕으로 뛰어난 추론 능력을 보여주며, 검색 엔진이나 코드 인터프리터와 같은 외부 도구를 활용하는 연구로 확장되고 있다. 그러나 기존 도...

    최근 대규모 언어 모델(LLM)은 방대한 데이터를 바탕으로 뛰어난 추론 능력을 보여주며, 검색 엔진이나 코드 인터프리터와 같은 외부 도구를 활용하는 연구로 확장되고 있다. 그러나 기존 도구 학습은 주로 가상 환경의 소프트웨어적 도구에 국한되어, 로봇이 물리적 환경에서 수행하는 행동을 LLM의 도구 관점에서 해석하려는 시도는 상대적으로 미비하였다. 이에 본 논문은 3D 가상환경 내에서 로봇의 행동을 LLM이 호출할 수 있는 함수 형태의 외부 도구로 정의하고, 이를 통해 에이전트가 인간 사용자와의 다중 턴 대화를 통해 복잡한 가사 과업을 수행하는 체화된 추론 프레임워크를 제안한다. 본 연구는 시각적 인식의 한계를 보완하기 위해 시뮬레이터가 제공하는 메타데이터를 활용하여 LLM에 최적화된 텍스트 기반의 상태 추상화 및 공간정보 증강 방법론을 적용하였다. 특히 제한된 컨텍스트 윈도우 내에서 추론 효율성을 극대화하기 위해, 프롬프트를 환경 상태, 상호작용 이력, 도구 목록의세 가지 축으로 구조화하는 컨텍스트 엔지니어링 체계를 구축하다. TEACh 데이터셋을 기반으로 한 실험 결과, 다중 턴 상호작용 과업에서 객체 선별 시 의미적 연관성을 고려한 필터링 전략이 더 높은 정확도가 보임을 확인하였다. 또한, LLM 전용 기능인 Function Calling을 활용했을 때 텍스트 생성 방식보다 정확도가 향상됨을 보았으며, 전체 이력 요약과 최근 상세 로그를 결합한 하이브리드 전략의 가능성을 보았다. 이러한 결과는 LLM 외부 도구로써의 로봇 액션의 가능성을 보여준다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities based on vast amounts of data, expanding into research that utilizes external tools such as search engines and code interpreters. However, existing tool learning has primarily been confined to software-based tools within virtual environments; consequently, attempts to interpret robot actions performed in physical environments from the perspective of LLM tools have been relatively scarce. To address this, we propose an embodied reasoning framework in which robot actions within a 3D virtual environment are defined as external tools in the form of functions callable by the LLM. This enables an agent to perform complex household tasks through multi-turn dialogue with human users. To compensate for the limitations of visual perception, we incorporate simulator-provided metadata and introduce text-based state abstraction and spatial-information augmentation methods optimized for LLMs. In particular, to maximize reasoning efficiency within a limited context window, we established a context engineering system that structures prompts along three axes: environmental state, interaction history, and tool list. Experimental results based on the TEACh dataset demonstrated that a filtering strategy considering semantic relevance during object selection yielded higher accuracy in multi-turn interaction tasks. Furthermore, we observed that utilizing the LLM-specific Function Calling feature improved accuracy compared to standard text generation methods. We also identified the potential of a hybrid strategy that combines a summary of the full history with recent detailed logs. These findings demonstrate the potential of integrating robot actions as external tools for LLMs.
    번역하기

    Recent Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities based on vast amounts of data, expanding into research that utilizes external tools such as search engines and code interpreters. However, existing tool learning ...

    Recent Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities based on vast amounts of data, expanding into research that utilizes external tools such as search engines and code interpreters. However, existing tool learning has primarily been confined to software-based tools within virtual environments; consequently, attempts to interpret robot actions performed in physical environments from the perspective of LLM tools have been relatively scarce. To address this, we propose an embodied reasoning framework in which robot actions within a 3D virtual environment are defined as external tools in the form of functions callable by the LLM. This enables an agent to perform complex household tasks through multi-turn dialogue with human users. To compensate for the limitations of visual perception, we incorporate simulator-provided metadata and introduce text-based state abstraction and spatial-information augmentation methods optimized for LLMs. In particular, to maximize reasoning efficiency within a limited context window, we established a context engineering system that structures prompts along three axes: environmental state, interaction history, and tool list. Experimental results based on the TEACh dataset demonstrated that a filtering strategy considering semantic relevance during object selection yielded higher accuracy in multi-turn interaction tasks. Furthermore, we observed that utilizing the LLM-specific Function Calling feature improved accuracy compared to standard text generation methods. We also identified the potential of a hybrid strategy that combines a summary of the full history with recent detailed logs. These findings demonstrate the potential of integrating robot actions as external tools for LLMs.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 제 1 절 연구의 배경 1
    • 제 2 절 연구의 목표 2
    • 제 2 장 관련 연구 4
    • 제 1 절 대규모 언어모델의 도구 학습 4
    • 제 1 장 서론 1
    • 제 1 절 연구의 배경 1
    • 제 2 절 연구의 목표 2
    • 제 2 장 관련 연구 4
    • 제 1 절 대규모 언어모델의 도구 학습 4
    • 제 2 절 체화된 에이전트와 대화 7
    • 제 3 장 방법론 8
    • 제 1 절 TEACh 데이터셋 8
    • 제 2 절 데이터셋 변환 및 프롬프트 구성 12
    • 제 4 장 실험 방법 및 결과 20
    • 제 1 절 실험 방법 20
    • 제 2 절 실험 결과 26
    • 제 5 장 결론 33
    • 제 1 절 연구 요약 및 시사점 33
    • 제 2 절 한계점 및 향후 연구 방향 34
    • 참고문헌 36
    • Abstract 39
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼