RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Accurate Discrimination of True Antibody-Antigen Structures in AlphaFold 3 via Multi-Metric Machine Learning = AlphaFold 3 기반 항체-항원 결합 구조의 신뢰성 평가 및 정밀 판별을 위한 통합 기계학습 프레임워크

    한글로보기

    https://www.riss.kr/link?id=T17449765

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 딥러닝 기술의 발전은 거대분자 구조 예측 분야에 상당한 진전을 가져왔으나, 항체-항원 상호작용의 대규모 가상 스크리닝(HTVS)을 위한 모델의 실질적인 적용은 여전히 계산 비효율성 및 신뢰도 보정의 한계로 인해 제한적이다. 본 연구에서는 AlphaFold 3(AF3)의 성능을 다중 메트릭 학습, 무질서도 인지 평가, 그리고 통계적으로 최적화된 샘플링 기법과 통합하여, 정확하고 비용 효율적인 대규모 스크리닝을 가능하게 하는 통합 계산 프레임워크를 제시한다.

    본 연구에서는 선별된 항원-항체 복합체 데이터셋을 사용하여 AF3 파생 구조 및 신뢰도 메트릭을 체계적으로 평가하였다. 그 결과, 인터페이스 예측 TM-점수(ipTM)가 결합 정확도를 예측하는 가장 강력한 단일 예측 변수(ROC-AUC = 0.770)로 나타났으나, 고정된 ipTM 임계값에 의존하는 방식은 높은 정밀도에도 불구하고 치명적으로 낮은 재현율을 보였다. 이러한 한계를 극복하기 위해, 본 연구에서는 XGBoost, 랜덤 포레스트, 서포트 벡터 머신, 로지스틱 회귀 등 기계 학습 분류기를 개발하였다. 이 모델들은 전역 pTM, 인터페이스 pLDDT, fraction disordered, 사슬 간 PAE, 접촉 확률 통계와 같은 상호 보완적인 특징들을 통합한다. L2-정규화된 로지스틱 회귀 모델이 정밀도와 재현율 간의 가장 우수한 균형(F1 = 0.560)을 달성하였으며, 이는 AF3 기본 모델 대비 26% 향상된 수치이다.

    SHapley Additive exPlanations(SHAP)을 사용한 특징 기여도 분석 결과, 인터페이스 주변의 국소적 무질서도가 결합 확률을 결정하는 가장 지배적인 요인임이 밝혀졌다. 무질서도와 예측된 결합 확률 간에 관찰된 임계값과 유사한 의존성은, 인터페이스에서의 구조적 엔트로피 증가가 안정적인 복합체 형성을 저해한다는 '접힘-결합 연동 모델(coupled folding-and-binding models)'의 생물리학적 특성과 일치한다. 이 발견은 계산 스크리닝 프레임워크에 구조적 무질서도를 통합하는 것의 중요성을 강조한다.

    또한, AF3 추론 과정을 합리화하기 위한 통계적 샘플링 프로토콜을 개발하였다. 정량적 수렴 분석을 통해 15-20개의 무작위 시드가 구조적 다양성의 95% 이상을 포착함을 입증하였으며, 이를 통해 정확도와 계산 비용 간의 원칙적인 균형점을 확립하였다.

    종합적으로, 본 연구는 AF3를 확률론적 구조 생성기에서 치료용 항체 발굴을 위한 실용적이고 의사결정이 가능한 스크리닝 도구로 변환시켰다. 본 논문에서 제안한 프레임워크는 예측 정확도와 해석 가능성을 향상시킬 뿐만 아니라, 다른 생체분자 도킹 및 설계 작업에도 적용 가능한 일반화된 방법론적 원칙을 수립한다. 따라서 본 연구는 신뢰도 보정과 계산 효율을 동시에 달성하는 확장형 스크리닝 프레임워크를 제안하여, 대규모 항체-항원 가상 스크리닝의 방법론적 기반을 강화한다.
    번역하기

    최근 딥러닝 기술의 발전은 거대분자 구조 예측 분야에 상당한 진전을 가져왔으나, 항체-항원 상호작용의 대규모 가상 스크리닝(HTVS)을 위한 모델의 실질적인 적용은 여전히 계산 비효율성 ...

    최근 딥러닝 기술의 발전은 거대분자 구조 예측 분야에 상당한 진전을 가져왔으나, 항체-항원 상호작용의 대규모 가상 스크리닝(HTVS)을 위한 모델의 실질적인 적용은 여전히 계산 비효율성 및 신뢰도 보정의 한계로 인해 제한적이다. 본 연구에서는 AlphaFold 3(AF3)의 성능을 다중 메트릭 학습, 무질서도 인지 평가, 그리고 통계적으로 최적화된 샘플링 기법과 통합하여, 정확하고 비용 효율적인 대규모 스크리닝을 가능하게 하는 통합 계산 프레임워크를 제시한다.

    본 연구에서는 선별된 항원-항체 복합체 데이터셋을 사용하여 AF3 파생 구조 및 신뢰도 메트릭을 체계적으로 평가하였다. 그 결과, 인터페이스 예측 TM-점수(ipTM)가 결합 정확도를 예측하는 가장 강력한 단일 예측 변수(ROC-AUC = 0.770)로 나타났으나, 고정된 ipTM 임계값에 의존하는 방식은 높은 정밀도에도 불구하고 치명적으로 낮은 재현율을 보였다. 이러한 한계를 극복하기 위해, 본 연구에서는 XGBoost, 랜덤 포레스트, 서포트 벡터 머신, 로지스틱 회귀 등 기계 학습 분류기를 개발하였다. 이 모델들은 전역 pTM, 인터페이스 pLDDT, fraction disordered, 사슬 간 PAE, 접촉 확률 통계와 같은 상호 보완적인 특징들을 통합한다. L2-정규화된 로지스틱 회귀 모델이 정밀도와 재현율 간의 가장 우수한 균형(F1 = 0.560)을 달성하였으며, 이는 AF3 기본 모델 대비 26% 향상된 수치이다.

    SHapley Additive exPlanations(SHAP)을 사용한 특징 기여도 분석 결과, 인터페이스 주변의 국소적 무질서도가 결합 확률을 결정하는 가장 지배적인 요인임이 밝혀졌다. 무질서도와 예측된 결합 확률 간에 관찰된 임계값과 유사한 의존성은, 인터페이스에서의 구조적 엔트로피 증가가 안정적인 복합체 형성을 저해한다는 '접힘-결합 연동 모델(coupled folding-and-binding models)'의 생물리학적 특성과 일치한다. 이 발견은 계산 스크리닝 프레임워크에 구조적 무질서도를 통합하는 것의 중요성을 강조한다.

    또한, AF3 추론 과정을 합리화하기 위한 통계적 샘플링 프로토콜을 개발하였다. 정량적 수렴 분석을 통해 15-20개의 무작위 시드가 구조적 다양성의 95% 이상을 포착함을 입증하였으며, 이를 통해 정확도와 계산 비용 간의 원칙적인 균형점을 확립하였다.

    종합적으로, 본 연구는 AF3를 확률론적 구조 생성기에서 치료용 항체 발굴을 위한 실용적이고 의사결정이 가능한 스크리닝 도구로 변환시켰다. 본 논문에서 제안한 프레임워크는 예측 정확도와 해석 가능성을 향상시킬 뿐만 아니라, 다른 생체분자 도킹 및 설계 작업에도 적용 가능한 일반화된 방법론적 원칙을 수립한다. 따라서 본 연구는 신뢰도 보정과 계산 효율을 동시에 달성하는 확장형 스크리닝 프레임워크를 제안하여, 대규모 항체-항원 가상 스크리닝의 방법론적 기반을 강화한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advances in deep learning have significantly improved macromolecular structure prediction, yet the practical deployment of these models for high-throughput virtual screening (HTVS) of antibody–antigen interactions remains constrained by computational inefficiency and limited confidence calibration. In this study, we present an integrated computational framework that augments AlphaFold 3 (AF3) with multi-metric learning, disorder-aware evaluation, and statistically optimized sampling to achieve accurate and cost-effective screening at scale.

    Using a curated dataset of antigen–antibody complexes, we systematically evaluated over twenty AF3-derived structural and confidence metrics. While the interface predicted TM-score (ipTM) emerged as the strongest single predictor of binding accuracy (ROC-AUC = 0.770), reliance on a fixed ipTM threshold resulted in high precision but critically low recall. To address this limitation, we developed machine learning classifiers—including XGBoost, Random Forest, Support Vector Machine, and Logistic Regression—that integrate complementary features such as global pTM, interface pLDDT, fraction disordered, inter-chain PAE, and contact probability statistics. The L2-regularized Logistic Regression model achieved the best balance between precision and recall (F1 = 0.560), representing a 26% improvement over the AF3 baseline.

    Feature attribution analyses using SHapley Additive exPlanations (SHAP) revealed that interface-local disorder is the dominant determinant of binding probability. The observed threshold-like dependence between disorder and predicted binding probability is biophysically consistent with coupled folding-and-binding models, in which increased conformational entropy at the interface disfavors stable complex formation. This finding underscores the importance of incorporating structural disorder into computational screening frameworks.

    A statistical sampling protocol was also developed to rationalize AF3 inference. Quantitative convergence analyses demonstrated that 15–20 random seeds capture over 95% of conformational diversity, establishing a principled balance between accuracy and computational cost.

    Overall, the proposed framework supports the use of AF3 as a practical screening tool for therapeutic antibody discovery. The proposed framework not only enhances predictive accuracy and interpretability but also establishes generalizable methodological principles applicable to other biomolecular docking and design tasks. This work contributes toward scalable, confidence-aware computational biophysics.
    번역하기

    Recent advances in deep learning have significantly improved macromolecular structure prediction, yet the practical deployment of these models for high-throughput virtual screening (HTVS) of antibody–antigen interactions remains constrained by compu...

    Recent advances in deep learning have significantly improved macromolecular structure prediction, yet the practical deployment of these models for high-throughput virtual screening (HTVS) of antibody–antigen interactions remains constrained by computational inefficiency and limited confidence calibration. In this study, we present an integrated computational framework that augments AlphaFold 3 (AF3) with multi-metric learning, disorder-aware evaluation, and statistically optimized sampling to achieve accurate and cost-effective screening at scale.

    Using a curated dataset of antigen–antibody complexes, we systematically evaluated over twenty AF3-derived structural and confidence metrics. While the interface predicted TM-score (ipTM) emerged as the strongest single predictor of binding accuracy (ROC-AUC = 0.770), reliance on a fixed ipTM threshold resulted in high precision but critically low recall. To address this limitation, we developed machine learning classifiers—including XGBoost, Random Forest, Support Vector Machine, and Logistic Regression—that integrate complementary features such as global pTM, interface pLDDT, fraction disordered, inter-chain PAE, and contact probability statistics. The L2-regularized Logistic Regression model achieved the best balance between precision and recall (F1 = 0.560), representing a 26% improvement over the AF3 baseline.

    Feature attribution analyses using SHapley Additive exPlanations (SHAP) revealed that interface-local disorder is the dominant determinant of binding probability. The observed threshold-like dependence between disorder and predicted binding probability is biophysically consistent with coupled folding-and-binding models, in which increased conformational entropy at the interface disfavors stable complex formation. This finding underscores the importance of incorporating structural disorder into computational screening frameworks.

    A statistical sampling protocol was also developed to rationalize AF3 inference. Quantitative convergence analyses demonstrated that 15–20 random seeds capture over 95% of conformational diversity, establishing a principled balance between accuracy and computational cost.

    Overall, the proposed framework supports the use of AF3 as a practical screening tool for therapeutic antibody discovery. The proposed framework not only enhances predictive accuracy and interpretability but also establishes generalizable methodological principles applicable to other biomolecular docking and design tasks. This work contributes toward scalable, confidence-aware computational biophysics.

    더보기

    목차 (Table of Contents)

    • Introduction 1
    • Body 7
    • Conclusion 55
    • Bibliography 85
    • Abstract in Korean 89
    • Introduction 1
    • Body 7
    • Conclusion 55
    • Bibliography 85
    • Abstract in Korean 89
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼