EmTaggeR: A Word Embedding Based Novel Method for Hashtag Recommendation on Twitter

Dey, Kuntal; Shrivastava, Ritvik; Kaushik, Saroj; Subramaniam, L. Venkata

Computer Science > Computation and Language

arXiv:1712.01562 (cs)

[Submitted on 5 Dec 2017]

Title:EmTaggeR: A Word Embedding Based Novel Method for Hashtag Recommendation on Twitter

Authors:Kuntal Dey, Ritvik Shrivastava, Saroj Kaushik, L. Venkata Subramaniam

View PDF

Abstract:The hashtag recommendation problem addresses recommending (suggesting) one or more hashtags to explicitly tag a post made on a given social network platform, based upon the content and context of the post. In this work, we propose a novel methodology for hashtag recommendation for microblog posts, specifically Twitter. The methodology, EmTaggeR, is built upon a training-testing framework that builds on the top of the concept of word embedding. The training phase comprises of learning word vectors associated with each hashtag, and deriving a word embedding for each hashtag. We provide two training procedures, one in which each hashtag is trained with a separate word embedding model applicable in the context of that hashtag, and another in which each hashtag obtains its embedding from a global context. The testing phase constitutes computing the average word embedding of the test post, and finding the similarity of this embedding with the known embeddings of the hashtags. The tweets that contain the most-similar hashtag are extracted, and all the hashtags that appear in these tweets are ranked in terms of embedding similarity scores. The top-K hashtags that appear in this ranked list, are recommended for the given test post. Our system produces F1 score of 50.83%, improving over the LDA baseline by around 6.53 times, outperforming the best-performing system known in the literature that provides a lift of 6.42 times. EmTaggeR is a fast, scalable and lightweight system, which makes it practical to deploy in real-life applications.

Comments:	Accepted at the IEEE International Conference on Data Mining (ICDM) 2017 ACUMEN Workshop
Subjects:	Computation and Language (cs.CL); Information Retrieval (cs.IR); Social and Information Networks (cs.SI)
Cite as:	arXiv:1712.01562 [cs.CL]
	(or arXiv:1712.01562v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1712.01562

Submission history

From: Ritvik Shrivastava [view email]
[v1] Tue, 5 Dec 2017 10:29:14 UTC (267 KB)

Computer Science > Computation and Language

Title:EmTaggeR: A Word Embedding Based Novel Method for Hashtag Recommendation on Twitter

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:EmTaggeR: A Word Embedding Based Novel Method for Hashtag Recommendation on Twitter

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators