ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge

Zhao, Zihan; Chen, Bo; Wan, Ziping; Chen, Lu; Lin, Xuanze; Yu, Shiyang; Zhang, Situo; Ma, Da; Zhu, Zichen; Zhang, Danyang; Wang, Huayang; Dai, Zhongyang; Wen, Liyang; Chen, Xin; Yu, Kai

Computer Science > Computational Engineering, Finance, and Science

arXiv:2507.21990 (cs)

[Submitted on 29 Jul 2025 (v1), last revised 30 Jul 2025 (this version, v2)]

Title:ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge

Authors:Zihan Zhao, Bo Chen, Ziping Wan, Lu Chen, Xuanze Lin, Shiyang Yu, Situo Zhang, Da Ma, Zichen Zhu, Danyang Zhang, Huayang Wang, Zhongyang Dai, Liyang Wen, Xin Chen, Kai Yu

View PDF HTML (experimental)

Abstract:While large language models (LLMs) have achieved impressive progress, their application in scientific domains such as chemistry remains hindered by shallow domain understanding and limited reasoning capabilities. In this work, we focus on the specific field of chemistry and develop a Chemical Reasoner LLM, ChemDFM-R. We first construct a comprehensive dataset of atomized knowledge points to enhance the model's understanding of the fundamental principles and logical structure of chemistry. Then, we propose a mix-sourced distillation strategy that integrates expert-curated knowledge with general-domain reasoning skills, followed by domain-specific reinforcement learning to enhance chemical reasoning. Experiments on diverse chemical benchmarks demonstrate that ChemDFM-R achieves cutting-edge performance while providing interpretable, rationale-driven outputs. Further case studies illustrate how explicit reasoning chains significantly improve the reliability, transparency, and practical utility of the model in real-world human-AI collaboration scenarios.

Comments:	13 figures, 4 tables
Subjects:	Computational Engineering, Finance, and Science (cs.CE); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2507.21990 [cs.CE]
	(or arXiv:2507.21990v2 [cs.CE] for this version)
	https://doi.org/10.48550/arXiv.2507.21990

Submission history

From: Zihan Zhao [view email]
[v1] Tue, 29 Jul 2025 16:40:49 UTC (1,555 KB)
[v2] Wed, 30 Jul 2025 07:23:58 UTC (1,555 KB)

Computer Science > Computational Engineering, Finance, and Science

Title:ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computational Engineering, Finance, and Science

Title:ChemDFM-R: An Chemical Reasoner LLM Enhanced with Atomized Chemical Knowledge

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators