RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    잠재 공간 기반 리간드 결합 자리 유사 구조 검색의 가속화 = Accelerated Template Based Search via Latent Mapping of Binding Site

    한글로보기

    https://www.riss.kr/link?id=T17451760

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    리간드 결합 자리 유사 구조 탐색은 단백질의 리간드 결합 자리를 규명하여 후속 과정의 검색 공간을 절약하는 신약 개발 과정에서의 중요한 단계 중 하나로, 실험적으로 규명된 리간드 결합 자리와의 비교를 통해 목표 단백질에서 결합 자리를 찾아낸다. 그러나 기존의 방식은 리간드 결합 자리의 구조 정보를 직접 비교하거나, 결합 자리 정보를 축약하여 비교하는 방식을 취하고 있어 확장성과 정밀성 간의 상충 문제를 겪고 있다.
    본 연구에서는 해당 문제를 극복하기 위해 그래프 신경망 기반 변분 오토인코더를 활용한 새로운 리간드 결합 자리 검색 방식, LEN-Seek을 제안한다. LEN-Seek은 단백질 결합 자리의 3차원 구조 및 물리화학적 정보를 확률적 잠재 공간으로 압축하여, 낮은 차원의 벡터 공간에서의 유사도 검색을 가능케 하였다.
    리간드 결합 자리는 불연속적으로 분포하는 아미노산 잔기의 집합으로 간주되어 각 아미노산 잔기를 하나의 정점으로 하는 그래프 형태로 표현되었다. 정점 특성은 단백질 언어 모델 Ankh의 임베딩을 활용하였고, 간선 특성으로는 아미노산 잔기 간 거리, 방향, 상대적 회전 정보를 포함하였다. 모든 기하학적 정보의 정의 및 그 계산은 회전-병진 불변성을 유지하도록 설계되어, 데이터 증강이나 등변 신경망을 사용하지 않음으로써 학습 효율을 향상시켰다. 모델은 잠재 벡터로부터 원본 정점 특성, 좌표 및 국소 좌표계를 복원하도록 학습되었으며, 그 결과 구조 및 문맥적 정보를 압축 및 그 압축된 벡터로부터 성공적으로 리간드 결합 자리 정보를 복원해 낼 수 있음을 확인하였다.
    기존의 유사 구조 검색 방법인 ProBiS와의 비교 평가의 일환으로서, LEN-Seek이 잠재 공간에서 검색해 낸 상위 결과의 기존의 방식과 높은 상관관계를 가짐을 확인하였으며, 또한 유사한 리간드를 가진 결합 자리를 찾아낼 수 있음을 확인하였다. 본 연구는 이러한 결과를 통하여 대규모 단백질 구조 데이터베이스에 대한 리간드 결합 자리의 유사 구조 검색의 가능성을 제시한다.
    번역하기

    리간드 결합 자리 유사 구조 탐색은 단백질의 리간드 결합 자리를 규명하여 후속 과정의 검색 공간을 절약하는 신약 개발 과정에서의 중요한 단계 중 하나로, 실험적으로 규명된 리간드 결...

    리간드 결합 자리 유사 구조 탐색은 단백질의 리간드 결합 자리를 규명하여 후속 과정의 검색 공간을 절약하는 신약 개발 과정에서의 중요한 단계 중 하나로, 실험적으로 규명된 리간드 결합 자리와의 비교를 통해 목표 단백질에서 결합 자리를 찾아낸다. 그러나 기존의 방식은 리간드 결합 자리의 구조 정보를 직접 비교하거나, 결합 자리 정보를 축약하여 비교하는 방식을 취하고 있어 확장성과 정밀성 간의 상충 문제를 겪고 있다.
    본 연구에서는 해당 문제를 극복하기 위해 그래프 신경망 기반 변분 오토인코더를 활용한 새로운 리간드 결합 자리 검색 방식, LEN-Seek을 제안한다. LEN-Seek은 단백질 결합 자리의 3차원 구조 및 물리화학적 정보를 확률적 잠재 공간으로 압축하여, 낮은 차원의 벡터 공간에서의 유사도 검색을 가능케 하였다.
    리간드 결합 자리는 불연속적으로 분포하는 아미노산 잔기의 집합으로 간주되어 각 아미노산 잔기를 하나의 정점으로 하는 그래프 형태로 표현되었다. 정점 특성은 단백질 언어 모델 Ankh의 임베딩을 활용하였고, 간선 특성으로는 아미노산 잔기 간 거리, 방향, 상대적 회전 정보를 포함하였다. 모든 기하학적 정보의 정의 및 그 계산은 회전-병진 불변성을 유지하도록 설계되어, 데이터 증강이나 등변 신경망을 사용하지 않음으로써 학습 효율을 향상시켰다. 모델은 잠재 벡터로부터 원본 정점 특성, 좌표 및 국소 좌표계를 복원하도록 학습되었으며, 그 결과 구조 및 문맥적 정보를 압축 및 그 압축된 벡터로부터 성공적으로 리간드 결합 자리 정보를 복원해 낼 수 있음을 확인하였다.
    기존의 유사 구조 검색 방법인 ProBiS와의 비교 평가의 일환으로서, LEN-Seek이 잠재 공간에서 검색해 낸 상위 결과의 기존의 방식과 높은 상관관계를 가짐을 확인하였으며, 또한 유사한 리간드를 가진 결합 자리를 찾아낼 수 있음을 확인하였다. 본 연구는 이러한 결과를 통하여 대규모 단백질 구조 데이터베이스에 대한 리간드 결합 자리의 유사 구조 검색의 가능성을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Ligand binding site identification is a step of computational drug discovery, which diminishes search space of following task in computational drug discovery and thus increases efficiency of drug discovery process. Template based binding site search method is one of the binding site search methods which identifies potential ligand binding site in a target protein by comparing the site to known ligand binding sites. However, conventional methods are based on one-by-one comparison on 3-dimensional structure or comparison on simplified vector, resulting in trade-off problem of scalability and accuracy. To address this limitation, we propose LEN-Seek, a novel template based binding site search framework based on a graph neural network (GNN)-driven variational autoencoder (VAE). LEN-Seek encodes structural and biochemical information of ligand binding site into probabilistic latent space, enabling similarity search within a low-dimensional vector space. Ligand binding site is represented as a set of discontinuous amino acid residues in 3D space, modeled as a graph in which each residue serves as a single node. Node features were derived from the embeddings of the protein language model Ankh, while edges were represented by inter-residue distance, orientation, and rotation. All geometric information was represented into roto-translational invariant features, which improved computational efficiency by avoiding data augmentation or SE(3)-equivariant models. LEN-Seek was trained to reconstruct the original node feature, coordinate, and local frame from the latent vector, confirming its ability to encode both structural and physicochemical context. Evaluation against the existent method ProBiS demonstrated that LEN-Seek successfully retrieved a substantial portion of similar ligand binding sites identified by ProBiS, while confirming the ability to find binding sites binding to similar ligands. These findings highlight the potential of the proposed approach for high-speed template based binding site search in large scale protein structure database.
    번역하기

    Ligand binding site identification is a step of computational drug discovery, which diminishes search space of following task in computational drug discovery and thus increases efficiency of drug discovery process. Template based binding site search m...

    Ligand binding site identification is a step of computational drug discovery, which diminishes search space of following task in computational drug discovery and thus increases efficiency of drug discovery process. Template based binding site search method is one of the binding site search methods which identifies potential ligand binding site in a target protein by comparing the site to known ligand binding sites. However, conventional methods are based on one-by-one comparison on 3-dimensional structure or comparison on simplified vector, resulting in trade-off problem of scalability and accuracy. To address this limitation, we propose LEN-Seek, a novel template based binding site search framework based on a graph neural network (GNN)-driven variational autoencoder (VAE). LEN-Seek encodes structural and biochemical information of ligand binding site into probabilistic latent space, enabling similarity search within a low-dimensional vector space. Ligand binding site is represented as a set of discontinuous amino acid residues in 3D space, modeled as a graph in which each residue serves as a single node. Node features were derived from the embeddings of the protein language model Ankh, while edges were represented by inter-residue distance, orientation, and rotation. All geometric information was represented into roto-translational invariant features, which improved computational efficiency by avoiding data augmentation or SE(3)-equivariant models. LEN-Seek was trained to reconstruct the original node feature, coordinate, and local frame from the latent vector, confirming its ability to encode both structural and physicochemical context. Evaluation against the existent method ProBiS demonstrated that LEN-Seek successfully retrieved a substantial portion of similar ligand binding sites identified by ProBiS, while confirming the ability to find binding sites binding to similar ligands. These findings highlight the potential of the proposed approach for high-speed template based binding site search in large scale protein structure database.

    더보기

    목차 (Table of Contents)

    • 제1장 서 론 1
    • 제1절 연구의 배경 1
    • 제2절 연구의 목적 12
    • 제2장 본문 15
    • 제1절 방법 15
    • 제1장 서 론 1
    • 제1절 연구의 배경 1
    • 제2절 연구의 목적 12
    • 제2장 본문 15
    • 제1절 방법 15
    • 제2절 결과 33
    • 제3장 고찰 43
    • 제4장 결론 57
    • 참고문헌 59
    • Abstract 68
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼