QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

Prakash, Shvetank; Cheng, Andrew; Tschand, Arya; Mazumder, Mark; Gohil, Varun; Ma, Jeffrey; Yik, Jason; Wan, Zishen; Quaye, Jessica; Alvanaki, Elisavet Lydia; Kumar, Avinash; Mazumdar, Chandrashis; Khare, Tuhin; Ingare, Alexander; Uchendu, Ikechukwu; Ghosal, Radhika; Tyagi, Abhishek; Wang, Chenyu; Garavagno, Andrea Mattia; Gu, Sarah; Guo, Alice; Hur, Grace; Carloni, Luca; Krishna, Tushar; Nayak, Ankita; Yazdanbakhsh, Amir; Reddi, Vijay Janapa

Abstract:The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced 'quark'), the first benchmark designed to facilitate the development and evaluation of LLM knowledge and reasoning capabilities specifically in computer architecture. QuArch provides a comprehensive collection of 2,671 expert-validated question-answer (QA) pairs covering various aspects of computer architecture, including processor design, memory systems, and interconnection networks. Our evaluation reveals that while frontier models possess domain-specific knowledge, they struggle with skills that require higher-order thinking in computer architecture. Frontier model accuracies vary widely (from 34% to 72%) on these advanced questions, highlighting persistent gaps in architectural reasoning across analysis, design, and implementation QAs. By holistically assessing fundamental skills, QuArch provides a foundation for building and measuring LLM capabilities that can accelerate innovation in computing systems. With over 140 contributors from 40 institutions, this benchmark represents a community effort to set the standard for architectural reasoning in LLM evaluation.

Subjects:	Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Software Engineering (cs.SE)
Cite as:	arXiv:2510.22087 [cs.AR]
	(or arXiv:2510.22087v1 [cs.AR] for this version)
	https://doi.org/10.48550/arXiv.2510.22087

Computer Science > Hardware Architecture

Title:QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators