대규모 언어 모델(LLM) 에이전트는 다단계, 도구 증강 작업에서 탁월한 성능을 보인다. 그러나 스마트 홈 환경은 잠재적 사용자 의도, 시간적 종속성, 장치 제약, 스케줄링 등을 처리해야 하는...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
대규모 언어 모델(LLM) 에이전트는 다단계, 도구 증강 작업에서 탁월한 성능을 보인다. 그러나 스마트 홈 환경은 잠재적 사용자 의도, 시간적 종속성, 장치 제약, 스케줄링 등을 처리해야 하는...
대규모 언어 모델(LLM) 에이전트는 다단계, 도구 증강 작업에서 탁월한 성능을 보인다. 그러나 스마트 홈 환경은 잠재적 사용자 의도, 시간적 종속성, 장치 제약, 스케줄링 등을 처리해야 하는 독특한 과제를 제시한다. 이러한 과제를 해결하기 위한 스마트 홈 에이전트 개발의 주요 병목은 두 가지이다. 첫째, 에이전트가 장치와 상호작용하고 결과를 관찰할 수 있는 현실적인 시뮬레이션 환경의 부재이며, 둘째, 이를 평가할 수 있는 도전적인 벤치마크의 부재이다. 이를 해결하기 위해 본 연구는 SimuHome을 제안한다. 이는 스마트 장치를 시뮬레이션하고 API 호출을 지원하며 환경 변수의 변화를 반영하는 시간 가속 홈 환경이다. 시뮬레이터를 스마트 홈 통신의 글로벌 산업 표준인 Matter 프로토콜 기반으로 구축함으로써, SimuHome은 고충실도 환경을 제공하며, 여기서 검증된 에이전트는 최소한의 적응만으로 실제 Matter 호환 장치에 배포 가능하다. 본 연구는 앞서 언급한 역량들을 요구하는 12가지 사용자 질의 유형에 걸쳐 600개의 에피소드로 구성된 도전적인 벤치마크를 제공한다. ReAct 프레임워크 하에서 16개 에이전트를 평가한 결과, 모델 간 뚜렷한 능력과 한계를 확인하였다. 70억 파라미터 미만의 모델들은 모든 질의 유형에서 무시할 만큼 미미한 성능을 보였다. 최고 성능을 기록한 표준 모델인 GPT-4.1조차도 암묵적 의도 추론, 상태 검증, 특히 시간적 스케줄링에서 어려움을 겪었다. 반면 GPT-5.1과 같은 추론 모델은 모든 질의 유형에서 표준 모델을 일관되게 능가했으나, 다만 평균 추론 시간이 3배 이상 소요되어 실시간 스마트 홈 애플리케이션에는 현실적으로 적용하기 어렵다. 이는 작업 성능과 실제 실용성 간의 중요한 상충관계를 강조한다.
다국어 초록 (Multilingual Abstract)
Large Language Model (LLM) agents excel at multi-step, tool-augmented tasks. However, smart homes introduce distinct challenges, requiring agents to handle latent user intents, temporal dependencies, device constraints, scheduling, and more. The main ...
Large Language Model (LLM) agents excel at multi-step, tool-augmented tasks. However, smart homes introduce distinct challenges, requiring agents to handle latent user intents, temporal dependencies, device constraints, scheduling, and more. The main bottlenecks for developing smart home agents with such capabilities include the lack of a realistic simulation environment where agents can interact with devices and observe the results, as well as a challenging benchmark to evaluate them. To address this, we introduce SimuHome, a time-accelerated home environment that simulates smart devices, supports API calls, and reflects changes in environmental variables. By building the simulator on the Matter protocol, the global industry standard for smart home communication, SimuHome provides a high-fidelity environment, and agents validated in SimuHome can be deployed on real Matter-compliant devices with minimal adaptation. We provide a challenging benchmark of 600 episodes across twelve user query types that require the aforementioned capabilities. Our evaluation of 16 agents under a unified ReAct framework reveals distinct capabilities and limitations across models. Models under 7B parameters exhibited negligible performance across all query types. Even GPT-4.1, the best-performing standard model, struggled with implicit intent inference, state verification, and particularly temporal scheduling. While reasoning models such as GPT-5.1 consistently outperformed standard models on every query type, they required over three times the average inference time, which can be prohibitive for real-time smart home applications. This highlights a critical trade-off between task performance and real-world practicality.
목차 (Table of Contents)