RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents = 시간 및 환경 변화를 반영한 스마트 홈 대규모 언어 모델 에이전트 벤치마크

    한글로보기

    https://www.riss.kr/link?id=T17450940

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델(LLM) 에이전트는 다단계, 도구 증강 작업에서 탁월한 성능을 보인다. 그러나 스마트 홈 환경은 잠재적 사용자 의도, 시간적 종속성, 장치 제약, 스케줄링 등을 처리해야 하는 독특한 과제를 제시한다. 이러한 과제를 해결하기 위한 스마트 홈 에이전트 개발의 주요 병목은 두 가지이다. 첫째, 에이전트가 장치와 상호작용하고 결과를 관찰할 수 있는 현실적인 시뮬레이션 환경의 부재이며, 둘째, 이를 평가할 수 있는 도전적인 벤치마크의 부재이다. 이를 해결하기 위해 본 연구는 SimuHome을 제안한다. 이는 스마트 장치를 시뮬레이션하고 API 호출을 지원하며 환경 변수의 변화를 반영하는 시간 가속 홈 환경이다. 시뮬레이터를 스마트 홈 통신의 글로벌 산업 표준인 Matter 프로토콜 기반으로 구축함으로써, SimuHome은 고충실도 환경을 제공하며, 여기서 검증된 에이전트는 최소한의 적응만으로 실제 Matter 호환 장치에 배포 가능하다. 본 연구는 앞서 언급한 역량들을 요구하는 12가지 사용자 질의 유형에 걸쳐 600개의 에피소드로 구성된 도전적인 벤치마크를 제공한다. ReAct 프레임워크 하에서 16개 에이전트를 평가한 결과, 모델 간 뚜렷한 능력과 한계를 확인하였다. 70억 파라미터 미만의 모델들은 모든 질의 유형에서 무시할 만큼 미미한 성능을 보였다. 최고 성능을 기록한 표준 모델인 GPT-4.1조차도 암묵적 의도 추론, 상태 검증, 특히 시간적 스케줄링에서 어려움을 겪었다. 반면 GPT-5.1과 같은 추론 모델은 모든 질의 유형에서 표준 모델을 일관되게 능가했으나, 다만 평균 추론 시간이 3배 이상 소요되어 실시간 스마트 홈 애플리케이션에는 현실적으로 적용하기 어렵다. 이는 작업 성능과 실제 실용성 간의 중요한 상충관계를 강조한다.
    번역하기

    대규모 언어 모델(LLM) 에이전트는 다단계, 도구 증강 작업에서 탁월한 성능을 보인다. 그러나 스마트 홈 환경은 잠재적 사용자 의도, 시간적 종속성, 장치 제약, 스케줄링 등을 처리해야 하는...

    대규모 언어 모델(LLM) 에이전트는 다단계, 도구 증강 작업에서 탁월한 성능을 보인다. 그러나 스마트 홈 환경은 잠재적 사용자 의도, 시간적 종속성, 장치 제약, 스케줄링 등을 처리해야 하는 독특한 과제를 제시한다. 이러한 과제를 해결하기 위한 스마트 홈 에이전트 개발의 주요 병목은 두 가지이다. 첫째, 에이전트가 장치와 상호작용하고 결과를 관찰할 수 있는 현실적인 시뮬레이션 환경의 부재이며, 둘째, 이를 평가할 수 있는 도전적인 벤치마크의 부재이다. 이를 해결하기 위해 본 연구는 SimuHome을 제안한다. 이는 스마트 장치를 시뮬레이션하고 API 호출을 지원하며 환경 변수의 변화를 반영하는 시간 가속 홈 환경이다. 시뮬레이터를 스마트 홈 통신의 글로벌 산업 표준인 Matter 프로토콜 기반으로 구축함으로써, SimuHome은 고충실도 환경을 제공하며, 여기서 검증된 에이전트는 최소한의 적응만으로 실제 Matter 호환 장치에 배포 가능하다. 본 연구는 앞서 언급한 역량들을 요구하는 12가지 사용자 질의 유형에 걸쳐 600개의 에피소드로 구성된 도전적인 벤치마크를 제공한다. ReAct 프레임워크 하에서 16개 에이전트를 평가한 결과, 모델 간 뚜렷한 능력과 한계를 확인하였다. 70억 파라미터 미만의 모델들은 모든 질의 유형에서 무시할 만큼 미미한 성능을 보였다. 최고 성능을 기록한 표준 모델인 GPT-4.1조차도 암묵적 의도 추론, 상태 검증, 특히 시간적 스케줄링에서 어려움을 겪었다. 반면 GPT-5.1과 같은 추론 모델은 모든 질의 유형에서 표준 모델을 일관되게 능가했으나, 다만 평균 추론 시간이 3배 이상 소요되어 실시간 스마트 홈 애플리케이션에는 현실적으로 적용하기 어렵다. 이는 작업 성능과 실제 실용성 간의 중요한 상충관계를 강조한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large Language Model (LLM) agents excel at multi-step, tool-augmented tasks. However, smart homes introduce distinct challenges, requiring agents to handle latent user intents, temporal dependencies, device constraints, scheduling, and more. The main bottlenecks for developing smart home agents with such capabilities include the lack of a realistic simulation environment where agents can interact with devices and observe the results, as well as a challenging benchmark to evaluate them. To address this, we introduce SimuHome, a time-accelerated home environment that simulates smart devices, supports API calls, and reflects changes in environmental variables. By building the simulator on the Matter protocol, the global industry standard for smart home communication, SimuHome provides a high-fidelity environment, and agents validated in SimuHome can be deployed on real Matter-compliant devices with minimal adaptation. We provide a challenging benchmark of 600 episodes across twelve user query types that require the aforementioned capabilities. Our evaluation of 16 agents under a unified ReAct framework reveals distinct capabilities and limitations across models. Models under 7B parameters exhibited negligible performance across all query types. Even GPT-4.1, the best-performing standard model, struggled with implicit intent inference, state verification, and particularly temporal scheduling. While reasoning models such as GPT-5.1 consistently outperformed standard models on every query type, they required over three times the average inference time, which can be prohibitive for real-time smart home applications. This highlights a critical trade-off between task performance and real-world practicality.
    번역하기

    Large Language Model (LLM) agents excel at multi-step, tool-augmented tasks. However, smart homes introduce distinct challenges, requiring agents to handle latent user intents, temporal dependencies, device constraints, scheduling, and more. The main ...

    Large Language Model (LLM) agents excel at multi-step, tool-augmented tasks. However, smart homes introduce distinct challenges, requiring agents to handle latent user intents, temporal dependencies, device constraints, scheduling, and more. The main bottlenecks for developing smart home agents with such capabilities include the lack of a realistic simulation environment where agents can interact with devices and observe the results, as well as a challenging benchmark to evaluate them. To address this, we introduce SimuHome, a time-accelerated home environment that simulates smart devices, supports API calls, and reflects changes in environmental variables. By building the simulator on the Matter protocol, the global industry standard for smart home communication, SimuHome provides a high-fidelity environment, and agents validated in SimuHome can be deployed on real Matter-compliant devices with minimal adaptation. We provide a challenging benchmark of 600 episodes across twelve user query types that require the aforementioned capabilities. Our evaluation of 16 agents under a unified ReAct framework reveals distinct capabilities and limitations across models. Models under 7B parameters exhibited negligible performance across all query types. Even GPT-4.1, the best-performing standard model, struggled with implicit intent inference, state verification, and particularly temporal scheduling. While reasoning models such as GPT-5.1 consistently outperformed standard models on every query type, they required over three times the average inference time, which can be prohibitive for real-time smart home applications. This highlights a critical trade-off between task performance and real-world practicality.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Chapter 1. Introduction 1
    • Chapter 2. Related Work 4
    • Chapter 3. SimuHome: A Smart Home Simulator 6
    • 3.1 Motivation
    • Abstract i
    • Chapter 1. Introduction 1
    • Chapter 2. Related Work 4
    • Chapter 3. SimuHome: A Smart Home Simulator 6
    • 3.1 Motivation
    • 3.2 Simulator Architecture and Operation
    • 3.3 Task Definition
    • Chaper 4. Benchmark Design 9
    • 4.1 Query Types
    • 4.2 Episode Generation
    • 4.3 Evaluation Methods
    • Chaper 5. Experiments 15
    • 5.1 Main Results
    • 5.2 Experiments
    • 5.2.1 Error Analysis
    • 5.2.2 Role of Tool Feedback
    • 5.2.3 Performance-Latency Trade-off
    • 5.2.4 Disentangling Framework Limitation from Model Capabilities
    • Chaper 6. Conclusion 22
    • Appendix A. Infeasible Query Types 23
    • Appendix B. List of Tools 25
    • Appendix C. List of Matter Clusters 27
    • Appendix D. List of Device Types 30
    • Appendix E. Error Analysis 32
    • E.1 Error Taxonomy Details
    • E.2 Error Type Distributions
    • E.3 Distribution of API Response Errors
    • Appendix F. Multi-turn Interactive Dialogue Experiments 34
    • Appendix G. Dynamic Re-evaluation with Post-Execution Failure Notice 36
    • Appendix H. Fine-tuning Experiment 38
    • Appendix I. Analysis of GPT-4.1 Performance on QT2-F 39
    • Appendix J. Addressing Deferred Feedback through Simulation-based Pre-validation 40
    • Appendix K. Discussion on Complex Envrionmental Interactions 41
    • Appendix L. Goal Examples 42
    • Appendix M. LLM Judge Validation 45
    • Appendix N. Experimental Setup 46
    • Appendix O. Prompts 47
    • O.1 ReAct Prompt
    • O.1.1 Baseline Prompt
    • O.1.2 Concise Prompt
    • O.2 LLM Judge Prompt
    • O.2.1 QT1 Feasible Judge Prompt
    • O.2.2 QT1 Infeasible Judge Prompt
    • O.2.3 QT2 Infeasible Judge Prompt
    • O.2.4 QT2 Infeasible-Nonexistence Judge Prompt
    • O.2.5 QT3 Infeasible Judge Prompt
    • O.2.6 QT4-1 Infeasible Judge Prompt
    • O.2.7 QT4-2 Infeasible Judge Prompt
    • O.2.8 QT4-3 Infeasible Judge Prompt
    • Appendix P. Prompt Variation Analysis 62
    • Appendix Q. Equations of Aggregators 64
    • 초록 71
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼