RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    코드 사전학습 언어 모델을 활용한 결함 위치 추정 및 패치 생성 기법 = Fault Localization and Patch Generation Techniques Using Code Pre-trained Language Models

    한글로보기

    https://www.riss.kr/link?id=T17379888

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Automated Program Repair (APR) is a technique that can effectively reduce the cost of the software development process. APR is generally composed of two stages: fault localization and patch generation. Recently, research on APR has been actively conducted by leveraging Code Pretrained Language Model (CodePLM). Within this context, this paper conducts two studies. First, existing studies that fine-tune CodePLMs for APR mainly focus on patch generation and do not address fault localization and patch generation in an integrated manner. However, in realistic scenarios, a model should be capable of handling both tasks jointly. To this end, this paper proposes a fine-tuning methodology that performs fault localization and patch generation simultaneously, and applies it to encoder-based, decoder-based, and encoder-decoder-based models to analyze their performance. Exeperimental results show that the proposed methodology achieves the best performance with the encoder-decoder-based model, and further confirmed the critical role of fault localization. Second, motivated by the observation that different layers of language models capture information at different levels and with different characteristics, a fine-tuning methodology that utilizes early layers has been proposed for code classification tasks. While this approach has shown performance improvements in bug detection tasks, its effectiveness for fault localization—an area closely related to bug detection—has not yet been explored. Accordingly, this paper analyzes whether the early layers of CodePLMs are effective for fault localization. Experimental results indicate that performance improvements are not consistently observed and that the magnitude of improvement is limited, suggesting that this approach is not effective. In summary, although strong fault localization capability is essential for jointly performing fault localization and patch generation, the results suggest that current CodePLMs have inherent limitations in sufficiently capturing the information required for accurate fault localization.
    번역하기

    Automated Program Repair (APR) is a technique that can effectively reduce the cost of the software development process. APR is generally composed of two stages: fault localization and patch generation. Recently, research on APR has been actively condu...

    Automated Program Repair (APR) is a technique that can effectively reduce the cost of the software development process. APR is generally composed of two stages: fault localization and patch generation. Recently, research on APR has been actively conducted by leveraging Code Pretrained Language Model (CodePLM). Within this context, this paper conducts two studies. First, existing studies that fine-tune CodePLMs for APR mainly focus on patch generation and do not address fault localization and patch generation in an integrated manner. However, in realistic scenarios, a model should be capable of handling both tasks jointly. To this end, this paper proposes a fine-tuning methodology that performs fault localization and patch generation simultaneously, and applies it to encoder-based, decoder-based, and encoder-decoder-based models to analyze their performance. Exeperimental results show that the proposed methodology achieves the best performance with the encoder-decoder-based model, and further confirmed the critical role of fault localization. Second, motivated by the observation that different layers of language models capture information at different levels and with different characteristics, a fine-tuning methodology that utilizes early layers has been proposed for code classification tasks. While this approach has shown performance improvements in bug detection tasks, its effectiveness for fault localization—an area closely related to bug detection—has not yet been explored. Accordingly, this paper analyzes whether the early layers of CodePLMs are effective for fault localization. Experimental results indicate that performance improvements are not consistently observed and that the magnitude of improvement is limited, suggesting that this approach is not effective. In summary, although strong fault localization capability is essential for jointly performing fault localization and patch generation, the results suggest that current CodePLMs have inherent limitations in sufficiently capturing the information required for accurate fault localization.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    프로그램 자동 수정은 소프트웨어 개발 과정의 비용을 효과적으로 절감할 수 있는 기술로, 크게 결함 위치 추정과 패치 생성의 두 단계로 구성된다. 최근에는 코드 사전학습 언어 모델을 활용하여 프로그램 자동 수정 연구가 활발히 진행되고 있다. 본 논문은 이러한 흐름 속에서 두 가지 연구를 수행하였다. 먼저, 기존의 코드 사전학습 언어 모델을 미세조정하여 프로그램 자동 수정에 활용하는 연구들은 주로 패치 생성에 집중되어 있으며 결함 위치 추정과 패치 생성을 통합적으로 다루지 않았다. 그러나 실제 상황을 고려한다면 모델은 두 과제를 통합적으로 다룰 수 있어야 한다. 이에 본 논문에서는 결함 위치 추정과 패치 생성을 한 번에 수행하기 위한 미세조정 방법론을 제안하고, 인코더 기반 모델, 디코더 기반 모델, 인코더-디코더 기반 모델에 적용하여 성능을 분석하였다. 실험 결과, 제안하는 방법론은 인코더-디코더 기반 모델에서 가장 우수한 성능을 보였으며, 결함 위치 추정의 중요성을 확인하였다. 다음으로, 언어 모델의 계층별 표현이 서로 다룬 수준과 특성의 정보를 포착한다는 점에 착안하여, 코드 분류 과제에서 초기 계층을 활용하는 미세조정 방법론이 제안되었다. 해당 연구는 버그 탐지 과제에서 성능 향상을 보였으나, 버그 탐지와 밀접한 관련이 있는 결함 위치 추정에 대해서는 아직 탐구되지 않았다. 이에 본 논문에서는 코드 사전학습 언어 모델의 초기 계층이 결함 위치 추정에 효과적인지를 분석하였다. 실험 결과, 성능 향상이 일관되게 나타나지 않았으며 성능 향상 폭이 크지 않아 효과적이지 않음을 확인하였다. 종합하면, 결함 위치 추정과 패치 생성을 통합적으로 수행하기 위해서는 결함 위치 추정 능력이 핵심적이나, 현재의 코드 사전학습 언어 모델은 결함 위치를 추정하는 데 필요한 정보를 충분히 확보하는 데 한계가 있음을 시사한다.
    번역하기

    프로그램 자동 수정은 소프트웨어 개발 과정의 비용을 효과적으로 절감할 수 있는 기술로, 크게 결함 위치 추정과 패치 생성의 두 단계로 구성된다. 최근에는 코드 사전학습 언어 모델을 활...

    프로그램 자동 수정은 소프트웨어 개발 과정의 비용을 효과적으로 절감할 수 있는 기술로, 크게 결함 위치 추정과 패치 생성의 두 단계로 구성된다. 최근에는 코드 사전학습 언어 모델을 활용하여 프로그램 자동 수정 연구가 활발히 진행되고 있다. 본 논문은 이러한 흐름 속에서 두 가지 연구를 수행하였다. 먼저, 기존의 코드 사전학습 언어 모델을 미세조정하여 프로그램 자동 수정에 활용하는 연구들은 주로 패치 생성에 집중되어 있으며 결함 위치 추정과 패치 생성을 통합적으로 다루지 않았다. 그러나 실제 상황을 고려한다면 모델은 두 과제를 통합적으로 다룰 수 있어야 한다. 이에 본 논문에서는 결함 위치 추정과 패치 생성을 한 번에 수행하기 위한 미세조정 방법론을 제안하고, 인코더 기반 모델, 디코더 기반 모델, 인코더-디코더 기반 모델에 적용하여 성능을 분석하였다. 실험 결과, 제안하는 방법론은 인코더-디코더 기반 모델에서 가장 우수한 성능을 보였으며, 결함 위치 추정의 중요성을 확인하였다. 다음으로, 언어 모델의 계층별 표현이 서로 다룬 수준과 특성의 정보를 포착한다는 점에 착안하여, 코드 분류 과제에서 초기 계층을 활용하는 미세조정 방법론이 제안되었다. 해당 연구는 버그 탐지 과제에서 성능 향상을 보였으나, 버그 탐지와 밀접한 관련이 있는 결함 위치 추정에 대해서는 아직 탐구되지 않았다. 이에 본 논문에서는 코드 사전학습 언어 모델의 초기 계층이 결함 위치 추정에 효과적인지를 분석하였다. 실험 결과, 성능 향상이 일관되게 나타나지 않았으며 성능 향상 폭이 크지 않아 효과적이지 않음을 확인하였다. 종합하면, 결함 위치 추정과 패치 생성을 통합적으로 수행하기 위해서는 결함 위치 추정 능력이 핵심적이나, 현재의 코드 사전학습 언어 모델은 결함 위치를 추정하는 데 필요한 정보를 충분히 확보하는 데 한계가 있음을 시사한다.

    더보기

    목차 (Table of Contents)

    • 목 차 iii
    • I. 서 론 1
    • II. 배경 지식 3
    • 1. 자동 프로그램 수정(Automated Program Repair, APR) 3
    • 1) 결함 위치 추정(Fault Localization) 3
    • 목 차 iii
    • I. 서 론 1
    • II. 배경 지식 3
    • 1. 자동 프로그램 수정(Automated Program Repair, APR) 3
    • 1) 결함 위치 추정(Fault Localization) 3
    • 2) 패치 생성(Patch Generation) 4
    • 2. 코드 사전학습 언어 모델(Code Pre-trained Language Model, CodePLM) 5
    • 1) 인코더 기반 모델 6
    • 2) 디코더 기반 모델 7
    • 3) 인코더-디코더 기반 모델 7
    • III. 통합적 결함 위치 추정 및 패치 생성 8
    • 1. 관련 연구 9
    • 2. 데이터 구축 11
    • 1) 원본 데이터 11
    • 2) 구축 데이터 12
    • 3. CodePLM 선정 12
    • 4. 방법론 14
    • 5. 실험 15
    • 1) 실험 환경 15
    • 2) 실험 평가 15
    • 3) 실험 방법 16
    • 4) 실험 결과 17
    • IV. 초기 계층을 활용한 결함 위치 추정 19
    • 1. 관련 연구 20
    • 2. 실험 21
    • 1) 데이터 21
    • 2) 분류 모델 구성 방법 21
    • 3) 분류 기반의 결함 위치 추정 방법 22
    • 4) 실험 환경 23
    • 5) 실험 평가 23
    • 6) 실험 결과 23
    • V. 결 론 25
    • VI. 참고 문헌 27
    • Abstract 31
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼