RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재후보

    WWW 환경에서 중복문서의 검출 기법에 대한 고찰 = A Survey on Detecting Duplicate Documents in World Wide Web Environment

    한글로보기

    https://www.riss.kr/link?id=A103722621

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recently, as the number of documents in the WWW(World Wide Web) increases, it becomes crucial
    to treat duplicate documents. In this article, we survey previous research results related to handling
    duplicate documents in WWW environment. First, we introduce a variety of methods for determining
    whether given two documents are duplicated. Second, we address methods for detecting duplicate
    documents efficiently from a large document database. Finally, we suggest further research directions.
    번역하기

    Recently, as the number of documents in the WWW(World Wide Web) increases, it becomes crucial to treat duplicate documents. In this article, we survey previous research results related to handling duplicate documents in WWW environment. First, we intr...

    Recently, as the number of documents in the WWW(World Wide Web) increases, it becomes crucial
    to treat duplicate documents. In this article, we survey previous research results related to handling
    duplicate documents in WWW environment. First, we introduce a variety of methods for determining
    whether given two documents are duplicated. Second, we address methods for detecting duplicate
    documents efficiently from a large document database. Finally, we suggest further research directions.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 들어 웹 문서가 증가함에 따라 중복문서 검출의 중요성이 점차 커지고 있다. 본 논문에서는 WWW 환경에서
    중복문서를 검출하는 기법에 관련된 기존의 연구 현황에 대하여 소개한다. 먼저, 두 개의 문서가 주어졌을 때 중복인
    지의 여부를 판정하는 기법들을 소개한다. 두 번째로는 대용량의 문서 데이터베이스에서 중복문서들을 효율적으로
    검출하는 기법들에 대해 논한다. 마지막으로 향후 연구 방향에 대하여 제시한다.
    번역하기

    최근 들어 웹 문서가 증가함에 따라 중복문서 검출의 중요성이 점차 커지고 있다. 본 논문에서는 WWW 환경에서 중복문서를 검출하는 기법에 관련된 기존의 연구 현황에 대하여 소개한다. 먼...

    최근 들어 웹 문서가 증가함에 따라 중복문서 검출의 중요성이 점차 커지고 있다. 본 논문에서는 WWW 환경에서
    중복문서를 검출하는 기법에 관련된 기존의 연구 현황에 대하여 소개한다. 먼저, 두 개의 문서가 주어졌을 때 중복인
    지의 여부를 판정하는 기법들을 소개한다. 두 번째로는 대용량의 문서 데이터베이스에서 중복문서들을 효율적으로
    검출하는 기법들에 대해 논한다. 마지막으로 향후 연구 방향에 대하여 제시한다.

    더보기

    참고문헌 (Reference)

    1 S. Schleimer, "Winnowing: Local Algorithms for Document Fingerprinting" SIGMOD 76-85, 2003

    2 N. Jain, "Using Bloom Filters to Refine Web Search Results" WebDB 25-30, 2005

    3 J. Zobel, "The Case of the Duplicate Documents Measurement, Search, and Science" APWeb 2006

    4 S. Brin, "The Anatomy of a Largescale Hypertextual Web Search Engine" 30 : 107-117, 1998

    5 A. Pereira Jr., "Syntactic Similarity of Web Documents" LAWEB 194-121, 2003

    6 A. Broder, "Syntactic Clustering of the Web" WWW 391-404, 1997

    7 S. Jonathan, "SpotSigs: Near Duplicate Detection in Web Page Collections" SIGIR 2007

    8 M. Charikar, "Similarity Estimation Techniques from Rounding Algorithms" 380-388, 2002

    9 S. Lawrence, "Searching the World Wide Web" 280 (280): 98-100, 1998

    10 A. Spink, "Searching the Web: the Public and Their Queries" 52 (52): 226-234, 2001

    1 S. Schleimer, "Winnowing: Local Algorithms for Document Fingerprinting" SIGMOD 76-85, 2003

    2 N. Jain, "Using Bloom Filters to Refine Web Search Results" WebDB 25-30, 2005

    3 J. Zobel, "The Case of the Duplicate Documents Measurement, Search, and Science" APWeb 2006

    4 S. Brin, "The Anatomy of a Largescale Hypertextual Web Search Engine" 30 : 107-117, 1998

    5 A. Pereira Jr., "Syntactic Similarity of Web Documents" LAWEB 194-121, 2003

    6 A. Broder, "Syntactic Clustering of the Web" WWW 391-404, 1997

    7 S. Jonathan, "SpotSigs: Near Duplicate Detection in Web Page Collections" SIGIR 2007

    8 M. Charikar, "Similarity Estimation Techniques from Rounding Algorithms" 380-388, 2002

    9 S. Lawrence, "Searching the World Wide Web" 280 (280): 98-100, 1998

    10 A. Spink, "Searching the Web: the Public and Their Queries" 52 (52): 226-234, 2001

    11 T. Haveliwala, "Scalable Techniques for Clustering the Web" WebDB 129-134, 2000

    12 N. Heintze, "Scalable Document Fingerprinting" 191-200, 1996

    13 N. Shivakumar, "SCAM: A Copy Detection Mechanism for Digital Documents" DL 155-163, 1995

    14 J. Conrad, "Online Duplicate Document Detection: Signature Reliability in a Dynamic Retrieval Environment" CIKM 443-452, 2003

    15 A. Broder, "On the Resemblance and Containment of Documents" 21-29, 1998

    16 D. Fetterly, "On the Evolution of Clusters of Near-Duplicate Web Pages" LA-WEB 37-45, 2003

    17 H. Yang, "Near-Duplicate Detection for eRulemaking" DGO 15-18, 2005

    18 H. Yang, "Near-Duplicate Detection by Instance-level Constrained Clustering" SIGIR 421-428, 2006

    19 K. Bharat, "Mirror, Mirror on the Web: A Study of Host Pairs with Replicated Content" WWW 1579-1590, 1999

    20 A. Broder, "Min-Wise Independent Permutations" 60 (60): 630-659, 2000

    21 T. Hoad, "Methods for Identifying Versioned and Plagiarized Documents" 54 (54): 203-215, 2003

    22 A. Kolcz, "Improved Robustness of Signature-based Near-replica Detection via Lexicon Randomization" SIGKDD 605-610, 2004

    23 A. Broder, "Identifying and Filtering Near- Duplicate Documents" CPM 1-10, 2000

    24 M. Rabin, "Fingerprinting by Random Polynomials" Harvard University 1981

    25 U. Manber, "Finding Similar Files in a Large File System" 1-10, 1994

    26 J. Dean, "Finding Related Pages in the World Wide Web" 314 : 1467-1479, 1999

    27 N. Shivakumar, "Finding Near-Replicas of Documents on the Web" WebDB 204-212, 1998

    28 M. Henzinger, "Finding Near-Duplicate Web Pages: A Large-Scale Evaluation of Algorithms" SIGIR 284-291, 2006

    29 T. Haveliwala, "Evaluating Strategies for Similarity Search on the Web" WWW 432-442, 2002

    30 J. Cooper, "Detecting Similar Documents Using Salient Terms" CIKM 245-251, 2002

    31 G. Manku, "Detecting Near-Duplicates for Web Crawling" WWW 141-149, 2007

    32 S. Brin, "Copy Detection Mechanisms for Digital Documents" SIGMOD 398-409, 1995

    33 J. Conrad, "Constructing a Text Corpus for Inexact Duplicate Detection" SIGIR 582-583, 2004

    34 A. Chowdhury, "Collection Statistics forFast Duplicate Document Detection" 20 (20): 171-191, 2002

    35 S. Park, "Analysis of Lexical Signatures for Finding Lost or Related Documents,”" SIGIR 11-18, 2002

    36 S. Ye, "A Systematic Study of Parameter Correlations in Large Scale Duplicate Document Detection" PAKDD 275-284, 2006

    37 S. Ye, "A Query-Dependent Duplicate Detection Approach for Large Scale Search Engines" APWeb 48-58, 2004

    더보기

    동일학술지(권/호) 다른 논문

    동일학술지 더보기

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2026 평가 재인증평가 신청대상 (재인증)
    2020-01-01 등재 등재학술지 유지 (재인증) KCI등재
    2017-01-01 등재 등재학술지 유지 (계속평가) KCI등재
    2013-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2010-01-01 등재 등재학술지 선정 (등재후보2차) KCI등재
    2009-01-01 등재 등재후보 1차 PASS (등재후보1차) KCI등재후보
    2007-01-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 0.02 0.02 0.01
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    0.02 0.02 0.183 0.03
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼