Towards Scalable Schema Mapping using Large Language Models

Buss, Christopher; Safari, Mahdis; Termehchy, Arash; Lee, Stefan; Maier, David

Computer Science > Databases

arXiv:2505.24716 (cs)

[Submitted on 30 May 2025]

Title:Towards Scalable Schema Mapping using Large Language Models

Authors:Christopher Buss, Mahdis Safari, Arash Termehchy, Stefan Lee, David Maier

View PDF HTML (experimental)

Abstract:The growing need to integrate information from a large number of diverse sources poses significant scalability challenges for data integration systems. These systems often rely on manually written schema mappings, which are complex, source-specific, and costly to maintain as sources evolve. While recent advances suggest that large language models (LLMs) can assist in automating schema matching by leveraging both structural and natural language cues, key challenges remain. In this paper, we identify three core issues with using LLMs for schema mapping: (1) inconsistent outputs due to sensitivity to input phrasing and structure, which we propose methods to address through sampling and aggregation techniques; (2) the need for more expressive mappings (e.g., GLaV), which strain the limited context windows of LLMs; and (3) the computational cost of repeated LLM calls, which we propose to mitigate through strategies like data type prefiltering.

Subjects:	Databases (cs.DB); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2505.24716 [cs.DB]
	(or arXiv:2505.24716v1 [cs.DB] for this version)
	https://doi.org/10.48550/arXiv.2505.24716

Submission history

From: Christopher Buss [view email]
[v1] Fri, 30 May 2025 15:36:56 UTC (167 KB)

Computer Science > Databases

Title:Towards Scalable Schema Mapping using Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Databases

Title:Towards Scalable Schema Mapping using Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators