-
LimeSoDa: A Dataset Collection for Benchmarking of Machine Learning Regressors in Digital Soil Mapping
Authors:
J. Schmidinger,
S. Vogel,
V. Barkov,
A. -D. Pham,
R. Gebbers,
H. Tavakoli,
J. Correa,
T. R. Tavares,
P. Filippi,
E. J. Jones,
V. Lukas,
E. Boenecke,
J. Ruehlmann,
I. Schroeter,
E. Kramer,
S. Paetzold,
M. Kodaira,
A. M. J. -C. Wadoux,
L. Bragazza,
K. Metzger,
J. Huang,
D. S. M. Valente,
J. L. Safanelli,
E. L. Bottega,
R. S. D. Dalmolin
, et al. (11 additional authors not shown)
Abstract:
Digital soil mapping (DSM) relies on a broad pool of statistical methods, yet determining the optimal method for a given context remains challenging and contentious. Benchmarking studies on multiple datasets are needed to reveal strengths and limitations of commonly used methods. Existing DSM studies usually rely on a single dataset with restricted access, leading to incomplete and potentially mis…
▽ More
Digital soil mapping (DSM) relies on a broad pool of statistical methods, yet determining the optimal method for a given context remains challenging and contentious. Benchmarking studies on multiple datasets are needed to reveal strengths and limitations of commonly used methods. Existing DSM studies usually rely on a single dataset with restricted access, leading to incomplete and potentially misleading conclusions. To address these issues, we introduce an open-access dataset collection called Precision Liming Soil Datasets (LimeSoDa). LimeSoDa consists of 31 field- and farm-scale datasets from various countries. Each dataset has three target soil properties: (1) soil organic matter or soil organic carbon, (2) clay content and (3) pH, alongside a set of features. Features are dataset-specific and were obtained by optical spectroscopy, proximal- and remote soil sensing. All datasets were aligned to a tabular format and are ready-to-use for modeling. We demonstrated the use of LimeSoDa for benchmarking by comparing the predictive performance of four learning algorithms across all datasets. This comparison included multiple linear regression (MLR), support vector regression (SVR), categorical boosting (CatBoost) and random forest (RF). The results showed that although no single algorithm was universally superior, certain algorithms performed better in specific contexts. MLR and SVR performed better on high-dimensional spectral datasets, likely due to better compatibility with principal components. In contrast, CatBoost and RF exhibited considerably better performances when applied to datasets with a moderate number (< 20) of features. These benchmarking results illustrate that the performance of a method is highly context-dependent. LimeSoDa therefore provides an important resource for improving the development and evaluation of statistical methods in DSM.
△ Less
Submitted 20 May, 2025; v1 submitted 27 February, 2025;
originally announced February 2025.
-
Transformer-based normative modelling for anomaly detection of early schizophrenia
Authors:
Pedro F Da Costa,
Jessica Dafflon,
Sergio Leonardo Mendes,
João Ricardo Sato,
M. Jorge Cardoso,
Robert Leech,
Emily JH Jones,
Walter H. L. Pinaya
Abstract:
Despite the impact of psychiatric disorders on clinical health, early-stage diagnosis remains a challenge. Machine learning studies have shown that classifiers tend to be overly narrow in the diagnosis prediction task. The overlap between conditions leads to high heterogeneity among participants that is not adequately captured by classification models. To address this issue, normative approaches h…
▽ More
Despite the impact of psychiatric disorders on clinical health, early-stage diagnosis remains a challenge. Machine learning studies have shown that classifiers tend to be overly narrow in the diagnosis prediction task. The overlap between conditions leads to high heterogeneity among participants that is not adequately captured by classification models. To address this issue, normative approaches have surged as an alternative method. By using a generative model to learn the distribution of healthy brain data patterns, we can identify the presence of pathologies as deviations or outliers from the distribution learned by the model. In particular, deep generative models showed great results as normative models to identify neurological lesions in the brain. However, unlike most neurological lesions, psychiatric disorders present subtle changes widespread in several brain regions, making these alterations challenging to identify. In this work, we evaluate the performance of transformer-based normative models to detect subtle brain changes expressed in adolescents and young adults. We trained our model on 3D MRI scans of neurotypical individuals (N=1,765). Then, we obtained the likelihood of neurotypical controls and psychiatric patients with early-stage schizophrenia from an independent dataset (N=93) from the Human Connectome Project. Using the predicted likelihood of the scans as a proxy for a normative score, we obtained an AUROC of 0.82 when assessing the difference between controls and individuals with early-stage schizophrenia. Our approach surpassed recent normative methods based on brain age and Gaussian Process, showing the promising use of deep generative models to help in individualised analyses.
△ Less
Submitted 8 December, 2022;
originally announced December 2022.
-
Neuroadaptive electroencephalography: a proof-of-principle study in infants
Authors:
Pedro F. da Costa,
Rianne Haartsen,
Elena Throm,
Luke Mason,
Anna Gui,
Robert Leech,
Emily J. H. Jones
Abstract:
A core goal of functional neuroimaging is to study how the environment is processed in the brain. The mainstream paradigm involves concurrently measuring a broad spectrum of brain responses to a small set of environmental features preselected with reference to previous studies or a theoretical framework. As a complement, we invert this approach by allowing the investigator to record the modulation…
▽ More
A core goal of functional neuroimaging is to study how the environment is processed in the brain. The mainstream paradigm involves concurrently measuring a broad spectrum of brain responses to a small set of environmental features preselected with reference to previous studies or a theoretical framework. As a complement, we invert this approach by allowing the investigator to record the modulation of a preselected brain response by a broad spectrum of environmental features. Our approach is optimal when theoretical frameworks or previous empirical data are impoverished. By using a prespecified closed-loop design, the approach addresses fundamental challenges of reproducibility and generalisability in brain research. These conditions are particularly acute when studying the developing brain, where our theories based on adult brain function may fundamentally misrepresent the topography of infant cognition and where there are substantial practical challenges to data acquisition. Our methodology employs machine learning to map modulation of a neural feature across a space of experimental stimuli. Our method collects, processes and analyses EEG brain data in real-time; and uses a neuro-adaptive Bayesian optimisation algorithm to adjust the stimulus presented depending on the prior samples of a given participant. Unsampled stimuli can be interpolated by fitting a Gaussian process regression along the dataset. We show that our method can automatically identify the face of the infant's mother through online recording of their Nc brain response to a face continuum. We can retrieve model statistics of individualised responses for each participant, opening the door for early identification of atypical development. This approach has substantial potential in infancy research and beyond for improving power and generalisability of mapping the individual cognitive topography of brain function.
△ Less
Submitted 10 June, 2021;
originally announced June 2021.