DebiasedDTA: A Framework for Improving the Generalizability of Drug-Target Affinity Prediction Models

Özçelik, Rıza; Bağ, Alperen; Atıl, Berk; Barsbey, Melih; Özgür, Arzucan; Özkırımlı, Elif

Quantitative Biology > Quantitative Methods

arXiv:2107.05556 (q-bio)

[Submitted on 4 Jul 2021 (v1), last revised 8 Jan 2023 (this version, v5)]

Title:DebiasedDTA: A Framework for Improving the Generalizability of Drug-Target Affinity Prediction Models

Authors:Rıza Özçelik, Alperen Bağ, Berk Atıl, Melih Barsbey, Arzucan Özgür, Elif Özkırımlı

View PDF

Abstract:Computational models that accurately predict the binding affinity of an input protein-chemical pair can accelerate drug discovery studies. These models are trained on available protein-chemical interaction datasets, which may contain dataset biases that may lead the model to learn dataset-specific patterns, instead of generalizable relationships. As a result, the prediction performance of models drops for previously unseen biomolecules, $\textit{i.e.}$ the prediction models cannot generalize to biomolecules outside of the dataset. The latest approaches that aim to improve model generalizability either have limited applicability or introduce the risk of degrading prediction performance. Here, we present DebiasedDTA, a novel drug-target affinity (DTA) prediction model training framework that addresses dataset biases to improve the generalizability of affinity prediction models. DebiasedDTA reweights the training samples to mitigate the effect of dataset biases and is applicable to most DTA prediction models. The results suggest that models trained in the DebiasedDTA framework can achieve improved generalizability in predicting the interactions of the previously unseen biomolecules, as well as performance improvements on those previously seen. Extensive experiments with different biomolecule representations, model architectures, and datasets demonstrate that DebiasedDTA can upgrade DTA prediction models irrespective of the biomolecule representation, model architecture, and training dataset. Last but not least, we release DebiasedDTA as an open-source python library to enable other researchers to debias their own predictors and/or develop their own debiasing methods. We believe that this python library will corroborate and foster research to develop more generalizable DTA prediction models.

Subjects:	Quantitative Methods (q-bio.QM); Machine Learning (cs.LG)
Cite as:	arXiv:2107.05556 [q-bio.QM]
	(or arXiv:2107.05556v5 [q-bio.QM] for this version)
	https://doi.org/10.48550/arXiv.2107.05556

Submission history

From: Rıza Özçelik [view email]
[v1] Sun, 4 Jul 2021 19:21:37 UTC (782 KB)
[v2] Sat, 17 Jul 2021 21:17:02 UTC (376 KB)
[v3] Thu, 13 Jan 2022 18:39:34 UTC (1,681 KB)
[v4] Tue, 17 May 2022 11:03:47 UTC (607 KB)
[v5] Sun, 8 Jan 2023 19:16:46 UTC (443 KB)

Quantitative Biology > Quantitative Methods

Title:DebiasedDTA: A Framework for Improving the Generalizability of Drug-Target Affinity Prediction Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Quantitative Biology > Quantitative Methods

Title:DebiasedDTA: A Framework for Improving the Generalizability of Drug-Target Affinity Prediction Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators