자연어 질의에 내재된 모호성은 Text-to-SQL 분야의 핵심적인 난제로 남아 있으며, 이는 빈번히 불완전하거나 부정확한 쿼리 생성으로 이어진다. 본 연구에서는 계층적 후보 생성, 로그 확률(log...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
자연어 질의에 내재된 모호성은 Text-to-SQL 분야의 핵심적인 난제로 남아 있으며, 이는 빈번히 불완전하거나 부정확한 쿼리 생성으로 이어진다. 본 연구에서는 계층적 후보 생성, 로그 확률(log...
자연어 질의에 내재된 모호성은 Text-to-SQL 분야의 핵심적인 난제로 남아 있으며, 이는 빈번히 불완전하거나 부정확한 쿼리 생성으로 이어진다. 본 연구에서는 계층적 후보 생성, 로그 확률(log-probability) 기반 정렬, 그리고 해석을 보존하는 가지치기 기법을 활용하여, 네 가지 주요 모호성 유형(Syntactic, Linking, Filtering, SELECT)을 체계적으로 해결하는 구조적 프레임워크인 BEACON을 제안한다. 이 프레임워크는 대형 언어 모델에 대한 변경을 최소화하면서 이러한 모호성 해소 단계들을 추론 파이프라인에 통합하며, 별도의 태스크 특화 파인튜닝(fine-tuning)을 필요로 하지 않는다. BEACON은 모호성 중심의 벤치마크(예: AMBROSIA, AmbiQT)에서 최고 수준의 성능을 달성하였으며, 실제 환경을 반영한 Text-to-SQL 데이터셋인 BIRD에서도 커버리지(coverage)와 재현율(recall)을 향상시켰다. 소거 연구 결과, 확률적 정렬과 해석 보존 기법이 커버리지 및 순위 품질 향상의 주요 요인임이 확인되었으며, 이는 다양한 도메인에 적용 가능한 일반적인 모호성 해결 방안임을 시사한다. 결론적으로, 이러한 결과들은 BEACON이 Text-to-SQL의 포괄적인 모호성 해결을 위한 실용적이고 범용적인 접근 방식임을 입증한다.
다국어 초록 (Multilingual Abstract)
Ambiguity in natural language questions remains a core obstacle for Text-to-SQL, often yielding incomplete or incorrect queries. We introduce BEACON, a structured framework that systematically addresses four major ambiguity types, namely Syntactic, Li...
Ambiguity in natural language questions remains a core obstacle for Text-to-SQL, often yielding incomplete or incorrect queries. We introduce BEACON, a structured framework that systematically addresses four major ambiguity types, namely Syntactic, Linking, Filtering, and SELECT, using hierarchical candidate generation, log-probability-guided ordering, and pruning that preserves promising yet deferred interpretations. The framework integrates these disambiguation steps into the inference pipeline with minimal changes to the underlying LLM and requires no task-specific fine-tuning. BEACON attains state-of-the-art performance on ambiguity-focused benchmarks (e.g., AMBROSIA and AmbiQT) and improves coverage and recall on a realistic Text-to-SQL dataset (BIRD). Ablations indicate that probabilistic ordering and interpretation retention are the primary contributors to coverage and ranking quality, suggesting a generalizable recipe for ambiguity resolution across domains. Together, these results position BEACON as a practical, domain-agnostic approach for comprehensive ambiguity resolution in Text-to-SQL.
목차 (Table of Contents)