Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text

Rehana, Hasin; Çam, Nur Bengisu; Basmaci, Mert; Zheng, Jie; Jemiyo, Christianah; He, Yongqun; Özgür, Arzucan; Hur, Junguk

Computer Science > Computation and Language

arXiv:2303.17728 (cs)

[Submitted on 30 Mar 2023 (v1), last revised 13 Dec 2023 (this version, v2)]

Title:Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text

Authors:Hasin Rehana, Nur Bengisu Çam, Mert Basmaci, Jie Zheng, Christianah Jemiyo, Yongqun He, Arzucan Özgür, Junguk Hur

View PDF

Abstract:Detecting protein-protein interactions (PPIs) is crucial for understanding genetic mechanisms, disease pathogenesis, and drug design. However, with the fast-paced growth of biomedical literature, there is a growing need for automated and accurate extraction of PPIs to facilitate scientific knowledge discovery. Pre-trained language models, such as generative pre-trained transformers (GPT) and bidirectional encoder representations from transformers (BERT), have shown promising results in natural language processing (NLP) tasks. We evaluated the performance of PPI identification of multiple GPT and BERT models using three manually curated gold-standard corpora: Learning Language in Logic (LLL) with 164 PPIs in 77 sentences, Human Protein Reference Database with 163 PPIs in 145 sentences, and Interaction Extraction Performance Assessment with 335 PPIs in 486 sentences. BERT-based models achieved the best overall performance, with BioBERT achieving the highest recall (91.95%) and F1-score (86.84%) and PubMedBERT achieving the highest precision (85.25%). Interestingly, despite not being explicitly trained for biomedical texts, GPT-4 achieved commendable performance, comparable to the top-performing BERT models. It achieved a precision of 88.37%, a recall of 85.14%, and an F1-score of 86.49% on the LLL dataset. These results suggest that GPT models can effectively detect PPIs from text data, offering promising avenues for application in biomedical literature mining. Further research could explore how these models might be fine-tuned for even more specialized tasks within the biomedical domain.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2303.17728 [cs.CL]
	(or arXiv:2303.17728v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2303.17728

Submission history

From: Junguk Hur [view email]
[v1] Thu, 30 Mar 2023 22:06:10 UTC (624 KB)
[v2] Wed, 13 Dec 2023 00:18:46 UTC (2,168 KB)

Computer Science > Computation and Language

Title:Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators