RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Activity Cliff Aware Active Learning with Multi-View Molecular Representation for Accelerating Hit Discovery = 활성 절벽 기반 능동학습과 다중 관정 분자 표현을 통한 히트 탐색 가속화

    한글로보기

    https://www.riss.kr/link?id=T17449770

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    능동학습은 제한된 라벨링 예산 하에서 분자 탐색을 가속화할 수 있는 효과적인 전략으로 주목받고 있다. 그러나 기존의 방법들은 분자 구조의 미세한 변화로 인해 생체 활성이 급격히 변하는 활성 절벽 현상을 충분히 포착하지 못해, 비효율적인 샘플 선택으로 이어지는 한계가 있다. 본 연구에서는 이러한 문제를 해결하기 위해 활성 절벽를 고려한 능동학습 프레임워크를 제안한다. 제안한 방법은 1D SMILES 기반 Molformer, 2D fingerprint, 3D UniMol 표현을 통합한 다중 관점 분자 표현을 활용하여 분자 구조를 상호 보완적인 관점에서 효과적으로 인코딩한다. 또한 분자 쌍 간 관계를 학습하는 활성 절벽 예측 모듈을 추가하여 국소적 구조-활성 민감도를 보다 정교하게 반영하고, 이를 기반으로 구조적으로 유의미한 분자를 우선적으로 선택하는 활성 절벽을 고려한 획득 함수를 설계하였다.

    제안한 프레임워크는 LIT-PCBA 벤치마크의 세 가지 단백질 표적을 대상으로 총 1,000개의 라벨링 예산 하에서 평가되었다. 실험 결과, 제안한 방법은 모든 표적에서 기존의 획득 전략 대비 더 많은 활성 화합물을 발견하며, 높은 enrichment factor를 달성하고 초기 소량 데이터 환경에서도 뛰어난 안정성을 보였다. 실험을 통해 다중 관점 인코더와 활성 절벽 예측 모듈이 모두 성능 향상에 기여함을 확인하였다. 전체적으로, 본 연구는 활성 절벽 인지와 보완적인 분자 표현을 능동학습에 통합하는 것이 데이터가 제한된 신약개발 환경에서 히트 탐색을 가속화하는 데 중요한 역할을 함을 강조한다.
    번역하기

    능동학습은 제한된 라벨링 예산 하에서 분자 탐색을 가속화할 수 있는 효과적인 전략으로 주목받고 있다. 그러나 기존의 방법들은 분자 구조의 미세한 변화로 인해 생체 활성이 급격히 변하...

    능동학습은 제한된 라벨링 예산 하에서 분자 탐색을 가속화할 수 있는 효과적인 전략으로 주목받고 있다. 그러나 기존의 방법들은 분자 구조의 미세한 변화로 인해 생체 활성이 급격히 변하는 활성 절벽 현상을 충분히 포착하지 못해, 비효율적인 샘플 선택으로 이어지는 한계가 있다. 본 연구에서는 이러한 문제를 해결하기 위해 활성 절벽를 고려한 능동학습 프레임워크를 제안한다. 제안한 방법은 1D SMILES 기반 Molformer, 2D fingerprint, 3D UniMol 표현을 통합한 다중 관점 분자 표현을 활용하여 분자 구조를 상호 보완적인 관점에서 효과적으로 인코딩한다. 또한 분자 쌍 간 관계를 학습하는 활성 절벽 예측 모듈을 추가하여 국소적 구조-활성 민감도를 보다 정교하게 반영하고, 이를 기반으로 구조적으로 유의미한 분자를 우선적으로 선택하는 활성 절벽을 고려한 획득 함수를 설계하였다.

    제안한 프레임워크는 LIT-PCBA 벤치마크의 세 가지 단백질 표적을 대상으로 총 1,000개의 라벨링 예산 하에서 평가되었다. 실험 결과, 제안한 방법은 모든 표적에서 기존의 획득 전략 대비 더 많은 활성 화합물을 발견하며, 높은 enrichment factor를 달성하고 초기 소량 데이터 환경에서도 뛰어난 안정성을 보였다. 실험을 통해 다중 관점 인코더와 활성 절벽 예측 모듈이 모두 성능 향상에 기여함을 확인하였다. 전체적으로, 본 연구는 활성 절벽 인지와 보완적인 분자 표현을 능동학습에 통합하는 것이 데이터가 제한된 신약개발 환경에서 히트 탐색을 가속화하는 데 중요한 역할을 함을 강조한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Active learning has emerged as an effective strategy for accelerating molecular discovery under limited labeling budgets. However, conventional methods often fail to capture activity cliffs—sharp changes in bioactivity induced by small structural perturbations—which can lead to suboptimal sample selection. In this work, we introduce an activity cliff–aware active learning framework that integrates multi-view molecular representation to enhance hit discovery efficiency. Our method combines 1D SMILES-based Molformer, 2D fingerprint, and 3D UniMol representations into a unified multi-view encoder, caturing molecular structure from multiple complementary views. To better capture local structure–activity sensitivity, we further incorporate an activity cliff prediction module trained on pairwise molecular relationships, and use its output to design a cliff-aware acquisition function that prioritizes structurally informative molecules for labeling.
    We evaluate our framework on three protein targets from the LIT-PCBA benchmark using a fixed 1,000-sample labeling budget. Across all targets, our method consistently discovers more active compounds than baseline acquisition strategies, achieving higher enrichment factors and demonstrating superior robustness in early-stage data-scarce regimes. Ablation studies confirm that both the multi-view encoder and the cliff prediction module contribute to performance gains. Overall, our results highlight the importance of incorporating activity-cliff awareness and complementary molecular representations within active learning frameworks, underscoring their combined contribution to accelerating hit discovery in data-limited drug discovery pipelines.
    번역하기

    Active learning has emerged as an effective strategy for accelerating molecular discovery under limited labeling budgets. However, conventional methods often fail to capture activity cliffs—sharp changes in bioactivity induced by small structural pe...

    Active learning has emerged as an effective strategy for accelerating molecular discovery under limited labeling budgets. However, conventional methods often fail to capture activity cliffs—sharp changes in bioactivity induced by small structural perturbations—which can lead to suboptimal sample selection. In this work, we introduce an activity cliff–aware active learning framework that integrates multi-view molecular representation to enhance hit discovery efficiency. Our method combines 1D SMILES-based Molformer, 2D fingerprint, and 3D UniMol representations into a unified multi-view encoder, caturing molecular structure from multiple complementary views. To better capture local structure–activity sensitivity, we further incorporate an activity cliff prediction module trained on pairwise molecular relationships, and use its output to design a cliff-aware acquisition function that prioritizes structurally informative molecules for labeling.
    We evaluate our framework on three protein targets from the LIT-PCBA benchmark using a fixed 1,000-sample labeling budget. Across all targets, our method consistently discovers more active compounds than baseline acquisition strategies, achieving higher enrichment factors and demonstrating superior robustness in early-stage data-scarce regimes. Ablation studies confirm that both the multi-view encoder and the cliff prediction module contribute to performance gains. Overall, our results highlight the importance of incorporating activity-cliff awareness and complementary molecular representations within active learning frameworks, underscoring their combined contribution to accelerating hit discovery in data-limited drug discovery pipelines.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Chapter 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Challenges 2
    • 1.3 Problem Definition 2
    • Abstract i
    • Chapter 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Challenges 2
    • 1.3 Problem Definition 2
    • 1.4 Our Approach 3
    • Chapter 2 Related works 5
    • 2.1 Active Learning for Molecular Discovery 5
    • 2.2 Molecular Representation Learning 6
    • 2.3 Activity Cliffs in Drug Discovery 6
    • Chapter 3 Method 8
    • 3.1 Multi-view Molecular Representation 8
    • 3.1.1 Molecular Representation 8
    • 3.1.2 Attention-based Fusion Module 10
    • 3.2 Activity Cliff Prediction Module 12
    • 3.2.1 Activity Cliff Pair Construction 12
    • 3.2.2 Module Architecture 13
    • 3.3 Sample Selection 14
    • Chapter 4 Experiments 16
    • 4.1 Dataset 16
    • 4.2 Experimental Setup 17
    • Chapter 5 Results 19
    • 5.1 Performance Comparison 19
    • 5.2 Cliff Pair Ordering Evaluation 20
    • 5.3 Ablation Study 24
    • 5.3.1 Impact of Cliff Prediction Module 24
    • 5.3.2 Impact of Multi-view Representation Module 25
    • Chapter 6 Conclusion and Future Works 29
    • Chapter 7 Appendix 36
    • 7.1 Datasets 36
    • 7.2 Baseline detail 36
    • 국문초록 38
    • 감사의 글 39
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼