Real-Time Fuzzy Record-Matching Similarity Metric and Optimal Q-Gram Filter
ČlánekOmezený přístuppeer-reviewedpostprintNačítá se...
Datum
Vedoucí práce
Oponent
Název časopisu
Název svazku
Nakladatel
Abstrakt
In this paper, we introduce an advanced Fuzzy Record Similarity Metric (FRMS) that improves approximate record matching and models human perception of record similarity. The FRMS utilizes a newly developed similarity space with favorable properties combined with a metric space, employing a bag-of-words model with general applications in text mining and cluster analysis. To optimize the FRMS, we propose a two-stage method for approximate string matching and search that outperforms baseline methods in terms of average time complexity and F measure on various datasets. In the first stage, we construct an optimal Q-gram count filter as an optimal lower bound for fuzzy token similarities such as FRMS. The approximated Q-gram count filter achieves a high accuracy rate, filtering over 99% of dissimilar records, with a constant time complexity of aproximate to 0(1). In the second stage, FRMS runs for a polynomial time of approximately approximate to 0(n4) and models human perception of record similarity by maximum weight matching in a bipartite graph. The FRMS architecture has widespread applications in structured document storage such as databases and has already been commercialized by one of the largest IT companies. As a side result, we explain the behavior of the singularity of the Q-gram filter and the advantages of a padding extension. Overall, our method provides a more accurate and efficient approach to approximate string matching and search with real-time runtime.
Rozsah stran
p. 1-32
ISSN
Permanentní identifikátor
Projekt
Časopis nebo seriál
Algorithms, volume 18, issue: 3
Vydavatelská verze
https://www.mdpi.com/1999-4893/18/3/150
Přístup k e-verzi
Pouze v rámci univerzity
Název akce
ISBN
Studijní obor
Studijní program
Signatura tištěné verze
Umístění tištěné verze
Přístup k tištěné verzi
Klíčová slova
fuzzy matching, Q-gram filter, approximate string matching, record linkage, entity resolution, similarity space, fuzzy shoda, Q-gramový filtr, přibližná shoda řetězců, propojení záznamů, rozlišení entit, prostor podobnosti