How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?

Purohit, Mirali; Muhawenayo, Gedeon; Rolf, Esther; Kerner, Hannah

Computer Science > Machine Learning

arXiv:2501.12535 (cs)

[Submitted on 21 Jan 2025]

Title:How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?

Authors:Mirali Purohit, Gedeon Muhawenayo, Esther Rolf, Hannah Kerner

View PDF HTML (experimental)

Abstract:Foundation models have made rapid advances in many domains including Earth observation, where Geospatial Foundation Models (GFMs) can help address global challenges such as climate change, agriculture, and disaster response. Previous work on GFMs focused on tailoring model architecture and pre-text tasks, and did not investigate the impact of pre-training data selection on model performance. However, recent works from other domains show that the pre-training data distribution is an important factor influencing the performance of the foundation models. With this motivation, our research explores how the geographic distribution of pre-training data affects the performance of GFMs. We evaluated several pre-training data distributions by sampling different compositions from a global data pool. Our experiments with two GFMs on downstream tasks indicate that balanced and globally representative data compositions often outperform region-specific sampling, highlighting the importance of diversity and global coverage in pre-training data. Our results suggest that the most appropriate data sampling technique may depend on the specific GFM architecture. These findings will support the development of robust GFMs by incorporating quality pre-training data distributions, ultimately improving machine learning solutions for Earth observation.

Comments:	Accepted at Good Data for Generative AI @ AAAI 2025
Subjects:	Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2501.12535 [cs.LG]
	(or arXiv:2501.12535v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2501.12535

Submission history

From: Mirali Purohit [view email]
[v1] Tue, 21 Jan 2025 22:57:09 UTC (5,602 KB)

Computer Science > Machine Learning

Title:How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators