1 S. Schleimer, "Winnowing: Local Algorithms for Document Fingerprinting" SIGMOD 76-85, 2003
2 N. Jain, "Using Bloom Filters to Refine Web Search Results" WebDB 25-30, 2005
3 J. Zobel, "The Case of the Duplicate Documents Measurement, Search, and Science" APWeb 2006
4 S. Brin, "The Anatomy of a Largescale Hypertextual Web Search Engine" 30 : 107-117, 1998
5 A. Pereira Jr., "Syntactic Similarity of Web Documents" LAWEB 194-121, 2003
6 A. Broder, "Syntactic Clustering of the Web" WWW 391-404, 1997
7 S. Jonathan, "SpotSigs: Near Duplicate Detection in Web Page Collections" SIGIR 2007
8 M. Charikar, "Similarity Estimation Techniques from Rounding Algorithms" 380-388, 2002
9 S. Lawrence, "Searching the World Wide Web" 280 (280): 98-100, 1998
10 A. Spink, "Searching the Web: the Public and Their Queries" 52 (52): 226-234, 2001
1 S. Schleimer, "Winnowing: Local Algorithms for Document Fingerprinting" SIGMOD 76-85, 2003
2 N. Jain, "Using Bloom Filters to Refine Web Search Results" WebDB 25-30, 2005
3 J. Zobel, "The Case of the Duplicate Documents Measurement, Search, and Science" APWeb 2006
4 S. Brin, "The Anatomy of a Largescale Hypertextual Web Search Engine" 30 : 107-117, 1998
5 A. Pereira Jr., "Syntactic Similarity of Web Documents" LAWEB 194-121, 2003
6 A. Broder, "Syntactic Clustering of the Web" WWW 391-404, 1997
7 S. Jonathan, "SpotSigs: Near Duplicate Detection in Web Page Collections" SIGIR 2007
8 M. Charikar, "Similarity Estimation Techniques from Rounding Algorithms" 380-388, 2002
9 S. Lawrence, "Searching the World Wide Web" 280 (280): 98-100, 1998
10 A. Spink, "Searching the Web: the Public and Their Queries" 52 (52): 226-234, 2001
11 T. Haveliwala, "Scalable Techniques for Clustering the Web" WebDB 129-134, 2000
12 N. Heintze, "Scalable Document Fingerprinting" 191-200, 1996
13 N. Shivakumar, "SCAM: A Copy Detection Mechanism for Digital Documents" DL 155-163, 1995
14 J. Conrad, "Online Duplicate Document Detection: Signature Reliability in a Dynamic Retrieval Environment" CIKM 443-452, 2003
15 A. Broder, "On the Resemblance and Containment of Documents" 21-29, 1998
16 D. Fetterly, "On the Evolution of Clusters of Near-Duplicate Web Pages" LA-WEB 37-45, 2003
17 H. Yang, "Near-Duplicate Detection for eRulemaking" DGO 15-18, 2005
18 H. Yang, "Near-Duplicate Detection by Instance-level Constrained Clustering" SIGIR 421-428, 2006
19 K. Bharat, "Mirror, Mirror on the Web: A Study of Host Pairs with Replicated Content" WWW 1579-1590, 1999
20 A. Broder, "Min-Wise Independent Permutations" 60 (60): 630-659, 2000
21 T. Hoad, "Methods for Identifying Versioned and Plagiarized Documents" 54 (54): 203-215, 2003
22 A. Kolcz, "Improved Robustness of Signature-based Near-replica Detection via Lexicon Randomization" SIGKDD 605-610, 2004
23 A. Broder, "Identifying and Filtering Near- Duplicate Documents" CPM 1-10, 2000
24 M. Rabin, "Fingerprinting by Random Polynomials" Harvard University 1981
25 U. Manber, "Finding Similar Files in a Large File System" 1-10, 1994
26 J. Dean, "Finding Related Pages in the World Wide Web" 314 : 1467-1479, 1999
27 N. Shivakumar, "Finding Near-Replicas of Documents on the Web" WebDB 204-212, 1998
28 M. Henzinger, "Finding Near-Duplicate Web Pages: A Large-Scale Evaluation of Algorithms" SIGIR 284-291, 2006
29 T. Haveliwala, "Evaluating Strategies for Similarity Search on the Web" WWW 432-442, 2002
30 J. Cooper, "Detecting Similar Documents Using Salient Terms" CIKM 245-251, 2002
31 G. Manku, "Detecting Near-Duplicates for Web Crawling" WWW 141-149, 2007
32 S. Brin, "Copy Detection Mechanisms for Digital Documents" SIGMOD 398-409, 1995
33 J. Conrad, "Constructing a Text Corpus for Inexact Duplicate Detection" SIGIR 582-583, 2004
34 A. Chowdhury, "Collection Statistics forFast Duplicate Document Detection" 20 (20): 171-191, 2002
35 S. Park, "Analysis of Lexical Signatures for Finding Lost or Related Documents,”" SIGIR 11-18, 2002
36 S. Ye, "A Systematic Study of Parameter Correlations in Large Scale Duplicate Document Detection" PAKDD 275-284, 2006
37 S. Ye, "A Query-Dependent Duplicate Detection Approach for Large Scale Search Engines" APWeb 48-58, 2004