신약 개발 초기 단계의 고질적인 데이터 희소성(Cold Start) 문제는 딥러닝 모델의 현장 적용을 가로막는 주요 장벽이다. 이에 별도의 학습 없이 소수의 예제(Shot)만으로 새로운 과업에 적응하...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17372371
서울 : 국민대학교 일반대학원, 2025
학위논문(석사) -- 국민대학교 일반대학원 , 컴퓨터공학과 컴퓨터공학전공 , 2026. 2
2025
한국어
서울
v, 34 ; 26 cm
지도교수: 박하명
I804:11014-200000962201
0
상세조회0
다운로드신약 개발 초기 단계의 고질적인 데이터 희소성(Cold Start) 문제는 딥러닝 모델의 현장 적용을 가로막는 주요 장벽이다. 이에 별도의 학습 없이 소수의 예제(Shot)만으로 새로운 과업에 적응하...
신약 개발 초기 단계의 고질적인 데이터 희소성(Cold Start) 문제는 딥러닝 모델의 현장 적용을 가로막는 주요 장벽이다. 이에 별도의 학습 없이 소수의 예제(Shot)만으로 새로운 과업에 적응하는 거대 언어 모델(Large Language Model; LLM) 기반의 In-Context Learning (ICL)이 유력한 대안으로 부상하고 있다. 그러나 기존의 단순 유사도 기반(Instance-level) 샷 선택 전략은 구조적으로 유사한 분자가 활성이 반대인 활성 절벽(Activity Cliff) 현상을 간과하며, 파편화된 문맥으로 인해 모델의 추론 성능을 저해한다는 한계가 있다. 본 연구에서는 이를 해결하기 위해 Pool-Clustering이라는 새로운 Set-level 샷 선택 프레임워크를 제안한다. 제안 방법은 전체 탐색 공간을 분자 골격(Scaffold) 기준으로 사전 군집화하고, 희소한 스캐폴드를 병합하여 구조적 문맥을 보존한 뒤, 질의(Query)와 연관된 클러스터 전체를 샷 집합으로 제공한다. 이는 LLM에게 동일 골격 내 구조적 변형에 따른 활성 변화의 인과 관계를 맥락적으로 학습시킨다. MoleculeNet 벤치마크의 분류 데이터셋(BACE, BBBP)에 대한 실험 결과, 제안 방법은 기존 학습 가능한 모델들과 대등한 경쟁력을 보였다. 특히, 기존 전략인 Top-k가 취약했던 활성 절벽 구간(Hard Case)에서 성능 하락을 효과적으로 억제하며 우수한 강건성과 추론 효율성을 입증하였다.
다국어 초록 (Multilingual Abstract)
Data sparsity, or the Cold Start problem, significantly hinders the application of deep learning in early-stage drug discovery. While LLM (Large Language Model)-based In-Context Learning (ICL) offers a training-free alternative, existing instance-leve...
Data sparsity, or the Cold Start problem, significantly hinders the application of deep learning in early-stage drug discovery. While LLM (Large Language Model)-based In-Context Learning (ICL) offers a training-free alternative, existing instance-level selection strategies fail to address "Activity Cliffs" and suffer from context fragmentation. To overcome this, we propose Pool-Clustering, a novel Set-leve} shot selection framework. By pre-clustering the search space (Pool) based on molecular scaffolds and retrieving entire clusters as context, our method enables LLMs to learn the causal relationship between structural variations and biological activity. Experiments on MoleculeNet (BACE, BBBP) demonstrate that our approach achieves performance comparable to supervised models and exhibits superior robustness in hard cases involving activity cliffs.
목차 (Table of Contents)