PubMed knowledge graph 2.0: Connecting papers, patents, and clinical trials in biomedical science

Jian Xu, Chao Yu, Jiawei Xu, Vetle I. Torvik, Jaewoo Kang, Mujeen Sung, Min Song, Yi Bu, Ying Ding

Published: 17 Jun 2025, Last Modified: 15 Jan 2026Scientific DataEveryoneRevisionsCC BY-SA 4.0
Abstract: Papers, patents, and clinical trials are essential scientific resources in biomedicine, crucial for knowledge sharing and dissemination. However, these documents are often stored in disparate databases with varying management standards and data formats, making it challenging to form systematic and fine-grained connections among them. To address this issue, we construct PKG 2.0, a comprehensive knowledge graph dataset encompassing over 36 million papers, 1.3 million patents, and 0.48 million clinical trials in the biomedical field. PKG 2.0 integrates these dispersed resources through 482 million biomedical entity linkages, 19 million citation linkages, and 7 million project linkages. The construction of PKG 2.0 wove together fine-grained biomedical entity extraction, high-performance author name disambiguation, multi-source citation integration, and high-quality project data from the NIH Exporter. Data validation demonstrates that PKG 2.0 excels in key tasks such as author disambiguation and biomedical entity recognition. This dataset provides valuable resources for biomedical researchers, bibliometric scholars, and those engaged in literature mining.
Loading