URL-BERT: Training Webpage Representations via Social Media Engagements

Qamar, Ayesha; Verma, Chetan; El-Kishky, Ahmed; Binnani, Sumit; Mehta, Sneha; Berg-Kirkpatrick, Taylor

Computer Science > Computation and Language

arXiv:2310.16303 (cs)

[Submitted on 25 Oct 2023]

Title:URL-BERT: Training Webpage Representations via Social Media Engagements

Authors:Ayesha Qamar, Chetan Verma, Ahmed El-Kishky, Sumit Binnani, Sneha Mehta, Taylor Berg-Kirkpatrick

View PDF

Abstract:Understanding and representing webpages is crucial to online social networks where users may share and engage with URLs. Common language model (LM) encoders such as BERT can be used to understand and represent the textual content of webpages. However, these representations may not model thematic information of web domains and URLs or accurately capture their appeal to social media users. In this work, we introduce a new pre-training objective that can be used to adapt LMs to understand URLs and webpages. Our proposed framework consists of two steps: (1) scalable graph embeddings to learn shallow representations of URLs based on user engagement on social media and (2) a contrastive objective that aligns LM representations with the aforementioned graph-based representation. We apply our framework to the multilingual version of BERT to obtain the model URL-BERT. We experimentally demonstrate that our continued pre-training approach improves webpage understanding on a variety of tasks and Twitter internal and external benchmarks.

Subjects:	Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:	arXiv:2310.16303 [cs.CL]
	(or arXiv:2310.16303v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2310.16303

Submission history

From: Ayesha Qamar [view email]
[v1] Wed, 25 Oct 2023 02:22:50 UTC (108 KB)

Computer Science > Computation and Language

Title:URL-BERT: Training Webpage Representations via Social Media Engagements

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:URL-BERT: Training Webpage Representations via Social Media Engagements

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators