RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    고장률 데이터북의 품목 분류를 위한 문자열 유사도 알고리즘 성능 비교에 관한 연구 = Performance Comparison of String Similarity Algorithms for Component Classification in a Failure Databook

    한글로보기

    https://www.riss.kr/link?id=A109444623

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Purpose: A failure databook provides failure rate information to assess system reliability. Such databooks have been developed and adopted in industries across several countries. Recently, a Korean-style failure rate databook, based on the field failure data of weapon systems, has been developed and distributed to defense industries. A critical aspect of a databook's development is the categorization of items, which is a time-intensive process. This study aims to support the classification of new items by evaluating the accuracy of five string similarity algorithms.
    Methods: The similarity algorithms considered in this study include the Jaro‒Winkler similarity, longest common subsequence (LCS) similarity, N-gram similarity, Levenshtein similarity, and Hamming similarity. To assess the accuracy of these algorithms, verification items were extracted through stratified sampling from the Korean failure rate databook, with the remaining items designated as reference data. Simulations were conducted by treating the verification items as new entries and evaluating how accurately each algorithm recommended the optimal major classification for these items.
    Results: The simulation results revealed that the Jaro‒Winkler similarity algorithm achieves the highest overall recommendation accuracy, followed by the LCS and N-gram similarity algorithms. Furthermore, as the number of recommended categories increases, the difference in accuracy between the Jaro‒Winkler and LCS algorithms becomes negligible.
    Conclusion: The findings of this study can be applied to automate item classification during the revision of failure rate databooks. Additionally, collecting supplementary data beyond item names could enable the development of advanced item classification recommendation techniques using various big data analysis methods.
    번역하기

    Purpose: A failure databook provides failure rate information to assess system reliability. Such databooks have been developed and adopted in industries across several countries. Recently, a Korean-style failure rate databook, based on the field failu...

    Purpose: A failure databook provides failure rate information to assess system reliability. Such databooks have been developed and adopted in industries across several countries. Recently, a Korean-style failure rate databook, based on the field failure data of weapon systems, has been developed and distributed to defense industries. A critical aspect of a databook's development is the categorization of items, which is a time-intensive process. This study aims to support the classification of new items by evaluating the accuracy of five string similarity algorithms.
    Methods: The similarity algorithms considered in this study include the Jaro‒Winkler similarity, longest common subsequence (LCS) similarity, N-gram similarity, Levenshtein similarity, and Hamming similarity. To assess the accuracy of these algorithms, verification items were extracted through stratified sampling from the Korean failure rate databook, with the remaining items designated as reference data. Simulations were conducted by treating the verification items as new entries and evaluating how accurately each algorithm recommended the optimal major classification for these items.
    Results: The simulation results revealed that the Jaro‒Winkler similarity algorithm achieves the highest overall recommendation accuracy, followed by the LCS and N-gram similarity algorithms. Furthermore, as the number of recommended categories increases, the difference in accuracy between the Jaro‒Winkler and LCS algorithms becomes negligible.
    Conclusion: The findings of this study can be applied to automate item classification during the revision of failure rate databooks. Additionally, collecting supplementary data beyond item names could enable the development of advanced item classification recommendation techniques using various big data analysis methods.

    더보기

    참고문헌 (Reference)

    1 R. B. Abernethy, "Weibull Analysis Handbook" Dr. Robert B. Abernethy 2006

    2 W. E. Winkler, "String comparator metrics and enhanced decision rules in the Fellegi-Sunter model of record linkage" 354-359, 1990

    3 OREDA Participants, "Offshore Reliability Data (OREDA)"

    4 C. D. Manning, "Introduction to Information Retrieval" Cambridge University Press 2008

    5 International Atomic Energy Agency, "Generic Component Reliability Data for Research Reactor PSA" IAEA 1988

    6 R. W. Hamming, "Error detecting and error correcting codes" 29 (29): 147-160, 1950

    7 Quanterion Solutions, "Electronic Parts Reliability Data (EPRD) / Nonelectronic Parts Reliability Data (NPRD)"

    8 정상원 ; 정기창, "Comparing string similarity algorithms for recognizing task names found in construction documents" 21 (21): 125-134, 2020

    9 V. I. Levenshtein, "Binary codes capable of correcting deletions, insertions, and reversals" 10 (10): 707-710, 1966

    10 D. S. Hirschberg, "Algorithms for the longest common subsequence problem" 24 (24): 664-675, 1977

    1 R. B. Abernethy, "Weibull Analysis Handbook" Dr. Robert B. Abernethy 2006

    2 W. E. Winkler, "String comparator metrics and enhanced decision rules in the Fellegi-Sunter model of record linkage" 354-359, 1990

    3 OREDA Participants, "Offshore Reliability Data (OREDA)"

    4 C. D. Manning, "Introduction to Information Retrieval" Cambridge University Press 2008

    5 International Atomic Energy Agency, "Generic Component Reliability Data for Research Reactor PSA" IAEA 1988

    6 R. W. Hamming, "Error detecting and error correcting codes" 29 (29): 147-160, 1950

    7 Quanterion Solutions, "Electronic Parts Reliability Data (EPRD) / Nonelectronic Parts Reliability Data (NPRD)"

    8 정상원 ; 정기창, "Comparing string similarity algorithms for recognizing task names found in construction documents" 21 (21): 125-134, 2020

    9 V. I. Levenshtein, "Binary codes capable of correcting deletions, insertions, and reversals" 10 (10): 707-710, 1966

    10 D. S. Hirschberg, "Algorithms for the longest common subsequence problem" 24 (24): 664-675, 1977

    11 M. Pecht, "A critique of MIL-HDBK-217E reliability prediction methods" 37 (37): 453-457, 1988

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼