Efficient Identification of Duplicate Bibliographical References

In this work we present an approach to extract and to structure bibliographical references from BibTex files, allowing the identification of the duplicate ones, which can appear slightly different in different files. To deal with this problem, existing systems use classifiers, clustering or others algorithms, allied with an Edit Distance metric, to distinguish between duplicate and nonduplicate records. The main challenge is to identify the duplicate records in database where the volume of the references can reach millions, in an efficient computational time. The technique proposed constructs a key (string) with information from each reference and stores them in a metric data structure called Slim-Tree. The Slim-Tree structure allows the minimization of the comparisons between references (being close to O(n log (n))), considering only the most similar keys to a given one.

author = {de Melo, Vin\’{i}cius Veloso and de Andrade Lopes, Alneu},
title = {Efficient Identification of Duplicate Bibliographical References},
booktitle = {Proceedings of the 2005 conference on Advances in Logic Based Intelligent Systems: Selected Papers of LAPTEC 2005},
year = {2005},
isbn = {1-58603-568-1},
pages = {169–176},
numpages = {8},
url = {http://dl.acm.org/citation.cfm?id=1565899.1565923},
acmid = {1565923},
publisher = {IOS Press},
address = {Amsterdam, The Netherlands, The Netherlands},
keywords = {Bibliographical References, Duplicate Record Detection, Metric Trees},