Can we use Google Scholar to identify highly-cited documents?

Martín-Martín, Alberto; Orduna-Malea, Enrique; Harzing, Anne-Wil; López-Cózar, Emilio Delgado

doi:10.1016/j.joi.2016.11.008

Computer Science > Digital Libraries

arXiv:1804.10439 (cs)

[Submitted on 27 Apr 2018]

Title:Can we use Google Scholar to identify highly-cited documents?

Authors:Alberto Martín-Martín, Enrique Orduna-Malea, Anne-Wil Harzing, Emilio Delgado López-Cózar

View PDF

Abstract:The main objective of this paper is to empirically test whether the identification of highly-cited documents through Google Scholar is feasible and reliable. To this end, we carried out a longitudinal analysis (1950 to 2013), running a generic query (filtered only by year of publication) to minimise the effects of academic search engine optimisation. This gave us a final sample of 64,000 documents (1,000 per year). The strong correlation between a document's citations and its position in the search results (r= -0.67) led us to conclude that Google Scholar is able to identify highly-cited papers effectively. This, combined with Google Scholar's unique coverage (no restrictions on document type and source), makes the academic search engine an invaluable tool for bibliometric research relating to the identification of the most influential scientific documents. We find evidence, however, that Google Scholar ranks those documents whose language (or geographical web domain) matches with the user's interface language higher than could be expected based on citations. Nonetheless, this language effect and other factors related to the Google Scholar's operation, i.e. the proper identification of versions and the date of publication, only have an incidental impact. They do not compromise the ability of Google Scholar to identify the highly-cited papers.

Subjects:	Digital Libraries (cs.DL)
Cite as:	arXiv:1804.10439 [cs.DL]
	(or arXiv:1804.10439v1 [cs.DL] for this version)
	https://doi.org/10.48550/arXiv.1804.10439
Journal reference:	Martín-Martín, A., Orduna-Malea, E., Harzing, A.-W., & Delgado López-Cózar, E. (2017). Can we use Google Scholar to identify highly-cited documents? Journal of Informetrics, 11(1), 152-163
Related DOI:	https://doi.org/10.1016/j.joi.2016.11.008

Submission history

From: Alberto Martín-Martín [view email]
[v1] Fri, 27 Apr 2018 11:02:21 UTC (1,150 KB)

Computer Science > Digital Libraries

Title:Can we use Google Scholar to identify highly-cited documents?

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Digital Libraries

Title:Can we use Google Scholar to identify highly-cited documents?

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators