Leveraging Predictions from Multiple Repositories to Improve Bot Detection

Chidambaram, Natarajan; Decan, Alexandre; Golzadeh, Mehdi

doi:10.1145/3528228.3528403

Computer Science > Software Engineering

arXiv:2203.16987 (cs)

[Submitted on 31 Mar 2022]

Title:Leveraging Predictions from Multiple Repositories to Improve Bot Detection

Authors:Natarajan Chidambaram, Alexandre Decan, Mehdi Golzadeh

View PDF

Abstract:Contemporary social coding platforms such as GitHub facilitate collaborative distributed software development. Developers engaged in these platforms often use machine accounts (bots) for automating effort-intensive or repetitive activities. Determining whether a contributor corresponds to a bot or a human account is important in socio-technical studies, for example, to assess the positive and negative impact of using bots, analyse the evolution of bots and their usage, identify top human contributors, and so on. BoDeGHa is one of the bot detection tools that have been proposed in the literature. It relies on comment activity within a single repository to predict whether an account is driven by a bot or by a human. This paper presents preliminary results on how the effectiveness of BoDeGHa can be improved by combining the predictions obtained from many repositories at once. We found that doing this not only increases the number of cases for which a prediction can be made but that many diverging predictions can be fixed this way. These promising, albeit preliminary, results suggest that the "wisdom of the crowd" principle can improve the effectiveness of bot detection tools.

Comments:	4 pages, 2 figures, 3 tables
Subjects:	Software Engineering (cs.SE)
Cite as:	arXiv:2203.16987 [cs.SE]
	(or arXiv:2203.16987v1 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2203.16987
Related DOI:	https://doi.org/10.1145/3528228.3528403

Submission history

From: Natarajan Chidambaram [view email]
[v1] Thu, 31 Mar 2022 12:19:22 UTC (187 KB)

Computer Science > Software Engineering

Title:Leveraging Predictions from Multiple Repositories to Improve Bot Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Software Engineering

Title:Leveraging Predictions from Multiple Repositories to Improve Bot Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators