Towards Universal Dense Blocking for Entity Resolution

Wang, Tianshu; Lin, Hongyu; Han, Xianpei; Chen, Xiaoyang; Cao, Boxi; Sun, Le

Computer Science > Databases

arXiv:2404.14831 (cs)

[Submitted on 23 Apr 2024 (v1), last revised 25 Apr 2024 (this version, v2)]

Title:Towards Universal Dense Blocking for Entity Resolution

Authors:Tianshu Wang, Hongyu Lin, Xianpei Han, Xiaoyang Chen, Boxi Cao, Le Sun

View PDF HTML (experimental)

Abstract:Blocking is a critical step in entity resolution, and the emergence of neural network-based representation models has led to the development of dense blocking as a promising approach for exploring deep semantics in blocking. However, previous advanced self-supervised dense blocking approaches require domain-specific training on the target domain, which limits the benefits and rapid adaptation of these methods. To address this issue, we propose UniBlocker, a dense blocker that is pre-trained on a domain-independent, easily-obtainable tabular corpus using self-supervised contrastive learning. By conducting domain-independent pre-training, UniBlocker can be adapted to various downstream blocking scenarios without requiring domain-specific fine-tuning. To evaluate the universality of our entity blocker, we also construct a new benchmark covering a wide range of blocking tasks from multiple domains and scenarios. Our experiments show that the proposed UniBlocker, without any domain-specific learning, significantly outperforms previous self- and unsupervised dense blocking methods and is comparable and complementary to the state-of-the-art sparse blocking methods.

Comments:	Code and data are available at this this https URL
Subjects:	Databases (cs.DB); Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:	arXiv:2404.14831 [cs.DB]
	(or arXiv:2404.14831v2 [cs.DB] for this version)
	https://doi.org/10.48550/arXiv.2404.14831

Submission history

From: Tianshu Wang [view email]
[v1] Tue, 23 Apr 2024 08:39:29 UTC (407 KB)
[v2] Thu, 25 Apr 2024 06:37:51 UTC (407 KB)

Computer Science > Databases

Title:Towards Universal Dense Blocking for Entity Resolution

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Databases

Title:Towards Universal Dense Blocking for Entity Resolution

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators