TY - GEN
T1 - Biomedical Knowledge Graph Completion with Efficient Contrastive Learning
AU - Qian, Jing
AU - Yue, Yong
AU - Atkinson, Katie
AU - Li, Gangmin
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - Knowledge graphs (KGs) are typically incomplete, but the task of knowledge graph completion (KGC) can address this issue by using existing facts to deduce the missing links. Biomedical KGC aims to automatically predict the head or tail entity in KG triples from biomedical text, which enables data-driven tasks such as drug discovery and disease treatment. With the success of Transformer architecture, textual encoding methods have emerged that utilize pre-trained language models (PLMs) to learn entity and relation representations. Nevertheless, the performance of textual encoding methods still substantially falls behind graph embedding methods, primarily due to the efficient contrastive learning of the latter. This paper is based on SimKGC that employ three distinct types of negatives, respectively, in-batch negatives, pre-batch negatives, and hard negatives to enhance the performance of textual encoding methods in biomedical KGC. Furthermore, we adopt a two-tower model with biomedical PLMs to encode entities and relations, respectively. It is the first attempt to apply efficient contrastive learning in biomedical KGC. Extensive experiments reveal that the combination of InfoNCE loss from contrastive learning and biomedical PLMs can substantially outperform graph embedding methods on two biomedical KGs, UMLS and Hetionet, in terms of automatic evaluation metrics (MR, MRR, and Hits@{1,3,10}).
AB - Knowledge graphs (KGs) are typically incomplete, but the task of knowledge graph completion (KGC) can address this issue by using existing facts to deduce the missing links. Biomedical KGC aims to automatically predict the head or tail entity in KG triples from biomedical text, which enables data-driven tasks such as drug discovery and disease treatment. With the success of Transformer architecture, textual encoding methods have emerged that utilize pre-trained language models (PLMs) to learn entity and relation representations. Nevertheless, the performance of textual encoding methods still substantially falls behind graph embedding methods, primarily due to the efficient contrastive learning of the latter. This paper is based on SimKGC that employ three distinct types of negatives, respectively, in-batch negatives, pre-batch negatives, and hard negatives to enhance the performance of textual encoding methods in biomedical KGC. Furthermore, we adopt a two-tower model with biomedical PLMs to encode entities and relations, respectively. It is the first attempt to apply efficient contrastive learning in biomedical KGC. Extensive experiments reveal that the combination of InfoNCE loss from contrastive learning and biomedical PLMs can substantially outperform graph embedding methods on two biomedical KGs, UMLS and Hetionet, in terms of automatic evaluation metrics (MR, MRR, and Hits@{1,3,10}).
KW - Biomedical pre-trained language models
KW - Contrastive learning
KW - Knowledge graph completion
UR - https://www.scopus.com/pages/publications/105027191942
U2 - 10.1007/978-3-031-93570-1_4
DO - 10.1007/978-3-031-93570-1_4
M3 - Conference Proceeding
AN - SCOPUS:105027191942
SN - 9783031935695
T3 - EAI/Springer Innovations in Communication and Computing
SP - 33
EP - 47
BT - 2nd International Conference on Big Data, IoT, and Cloud Computing - ICBICC 2024
A2 - Jia, Xiaolin
A2 - Zhang, Hui
A2 - Ramayah, Thurasamy
PB - Springer Science and Business Media Deutschland GmbH
T2 - 2nd International Conference on Big Data, IoT, and Cloud Computing, ICBICC 2024
Y2 - 30 December 2024 through 1 January 2025
ER -