최근 비정형 데이터에 대한 검색 수요가 증가함에 따라, 벡터 유사도 검색(vector similarity search)에 메타데이터 필터를 결합한 필터 기반 근사 최근접 이웃 탐색(filtered-ANNS)의 중요성이 커지고 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17451143
서울 : 서울대학교 대학원, 2026
2026
한국어
벡터 검색 ; 필터 기반 ANNS ; 그래프 인덱스 ; 대규모 벡터–라벨 데이터셋 ; StackOverflow10M ; UNG+
621.39
서울
vi, 56 ; 26 cm
지도교수: 김진수
I804:11032-000000194526
0
상세조회0
다운로드최근 비정형 데이터에 대한 검색 수요가 증가함에 따라, 벡터 유사도 검색(vector similarity search)에 메타데이터 필터를 결합한 필터 기반 근사 최근접 이웃 탐색(filtered-ANNS)의 중요성이 커지고 ...
최근 비정형 데이터에 대한 검색 수요가 증가함에 따라, 벡터 유사도 검색(vector similarity search)에 메타데이터 필터를 결합한 필터 기반 근사 최근접 이웃 탐색(filtered-ANNS)의 중요성이 커지고 있다. 그러나 filtered-ANNS를 현실적으로 평가하기 위해 필요한, 수천만 개 벡터와 수만 개 고카디널리티(high-cardinality) 라벨이 결합된 대규모 벡터--라벨 데이터셋은 부족한 상황이다. 그 결과 기존 filtered-ANNS 연구는 소규모 데이터셋 또는 SIFT·DEEP 벡터에 합성 라벨을 부여한 데이터셋에 기반해 성능을 검증하는 경우가 많았다.
본 논문에서는 먼저 공개 Stack Overflow 데이터를 기반으로, 약 1천만(10M) 개의 질문 임베딩과 약 6만 개 수준의 태그를 결합한 대규모 벡터--라벨 데이터셋(StackOverflow10M)을 구축한다. 각 질문의 제목과 본문을 벡터로 임베딩하고, 사람이 부여한 태그를 다중 라벨 필터로 사용함으로써 합성 라벨에 의존하지 않는 현실적인 filtered-ANNS 평가 환경을 제공한다. 이어서 기존 Unified Navigating Graph(UNG)를 확장하여, LNG 트리 하향 탐색 과정에서 발견되는 distant superset 라벨 그룹에 extra cross-group edge를 추가하는 \textbf{UNG+}를 제안하고, 브루트포스(brute-force) 구축 비용을 줄이기 위해 기존 벡터 그래프를 재활용하는 변형을 함께 제시한다.
실험 결과, SIFT10M과 DEEP10M에서는 동일 재현율(recall) 기준으로 UNG+가 평균 약 9.9\% 더 높은 처리량(throughput)을 보였고, StackOverflow10M에서는 동일 처리량 기준으로 재현율이 평균 약 4.5\% 향상되었다. 이는 제안 기법이 합성 라벨 및 실제 라벨 환경 모두에서, 고카디널리티 라벨을 갖는 대규모 벡터 데이터셋에 대해 보다 우수한 recall--QPS 트레이드오프를 제공함을 보여준다.
다국어 초록 (Multilingual Abstract)
Recent growth in the demand for searching unstructured data has increased the importance of filtered approximate nearest neighbor search (filtered-ANNS), which combines vector similarity search with metadata filtering. However, large-scale vector–la...
Recent growth in the demand for searching unstructured data has increased the importance of filtered approximate nearest neighbor search (filtered-ANNS), which combines vector similarity search with metadata filtering. However, large-scale vector–label datasets that realistically reflect production conditions—tens of millions of vectors paired with tens of thousands of high-cardinality labels—remain scarce. As a result, prior filtered-ANNS studies have often validated performance using small-scale datasets or datasets created by assigning synthetic labels to SIFT/DEEP vectors.
This thesis constructs a large-scale vector--label dataset, StackOverflow10M, by combining approximately 10 million question embeddings with around 60 thousand tags from publicly available Stack Overflow data. The title and body of each question are embedded into vectors, and human-assigned tags serve as multi-label filters, providing a realistic filtered-ANNS evaluation setting without relying on synthetic labels. To improve Unified Navigating Graph (UNG), UNG+ adds extra cross-group edges to distant superset label groups discovered during the top-down traversal of the LNG tree. To reduce brute-force construction cost, a variant that reuses an existing vector graph is also introduced.
Experimental results show that, on SIFT10M and DEEP10M, UNG+ achieves on average about 9.9\% higher throughput at the same recall, and on StackOverflow10M, it improves recall by about 4.5\% on average at the same throughput. These results demonstrate that the proposed method provides a better recall--QPS trade-off for large-scale vector datasets with high-cardinality labels in both synthetic-label and real-label settings.
목차 (Table of Contents)