Machine Learning: Science and Technology · 6 January 2025

Machine and deep learning performance in out-of-distribution regressions

Assaf Shmuel, Oren Glickman, Teddy Lazebnik

Affiliations
  1. Department of Computer Science, Bar Ilan University, Ramat Gan, Israel
  2. Department of Cancer Biology, Cancer Institute, University College London, London, United Kingdom

The paper at a glance

Machine learning models assume their training data represent the task well, and when new data fall outside that distribution, performance can drop unexpectedly. We measured this out-of-distribution drop for several machine learning, deep learning and AutoML models on 15 real-world regression datasets. We also propose using symbolic regression to engineer new features, which improved out-of-distribution performance by 3.70% for machine learning and 10.20% for deep learning models on average, without reducing in-distribution performance.

15real-world regression datasets in the benchmark
3.70%average out-of-distribution gain for machine learning models
10.20%average out-of-distribution gain for deep learning models

Key findings

  • We compared in-distribution and out-of-distribution performance of XGBoost, random forest, k-nearest neighbors, support vector machine, linear regression and two AutoML models.
  • The comparison used a benchmark of 15 real-world regression datasets.
  • Adding symbolic regression-derived features improved out-of-distribution performance by 3.70% for machine learning and 10.20% for deep learning models on average.
  • These features did not reduce in-distribution performance and in fact slightly improved it.
Figure 1. RMSE scores as a function of temperature Z-score. The results are based on 50 runs of the XGBoost model trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest Z-scores, based exclusively on the temperature feature. The ID test data are randomly chosen from the remaining 85% observations.
Figure 1. RMSE scores as a function of temperature Z-score. The results are based on 50 runs of the XGBoost model trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest Z-scores, based exclusively on the temperature feature. The ID test data are randomly chosen from the remaining 85% observations. See it in the paper
On this page
  1. Abstract
  2. 1. Introduction
  3. 2. Related work
  4. 2.1. Out of distribution in data-driven models
  5. 2.2. Sr
  6. 3. Methods and materials
  7. 3.1. Dataset
  8. 3.2. Experimental setup
  9. 3.3. ML and DL models
  10. 3.4. Integration of SR in ML and DL models
  11. 4. Results
  12. 4.1. ML and DL performance in OOD data
  13. 4.2. Integration of SR in ML and DL models to improve OOD performance
  14. 5. Discussion
  15. Data availability statement
  16. Funding
  17. Appendix
  18. ORCID iDs
  19. Article notes
  20. References

Abstract

Machine learning (ML) and deep learning (DL) models are gaining popularity due to their effectiveness in many computational tasks. These models are based on an intuitive, but frequently unsatisfied, assumption that the data used to train these models is well-representing the task at hand. This gives rise to the out-of-distribution (OOD) challenge which can cause an unexpected drop in the data-driven model's performance. In this study, we evaluate the performance of various ML and DL models in in-distribution (ID) versus OOD prediction. While the degradation in OOD performance is well acknowledged, to the best of our knowledge, this is one of the first studies to quantify it for various models on a large benchmark n = 15 real-world regression datasets. We extensively (runs) compare the ID versus OOD performance of XGBoost, random forest, K-nearest-neighbors, support vector machine, and linear regression models, as well as AutoML models (Tree-based Pipeline Optimization Tool and AutoKeras). In addition, to tackle this challenge, we propose to integrate a symbolic regression (SR) as a feature engineering method model with an ML or DL model to improve its performance for OOD samples. Our results show that the incorporation of SR-derived features significantly enhances the predictive capabilities of both ML and DL models with 3.70% and 10.20%, on average, of the OOD samples, respectively, without reducing ID performance and in fact improving it to a slightly lower extent. As such, this method can help produce more generalized and robust data-driven models.

1. Introduction

Achieving a high level of performance in regression and classification tasks through machine learning (ML) and deep learning (DL) models poses a fundamental computational challenge, crucial for applications across diverse scientific and engineering fields [1–5]. The effectiveness of ML-based models is contingent on diverse components governing its performance such as the nature of the problem and the available data used to train the model [6–14]. A growing body of scholarship investigates the characteristics of a dataset in data-driven tasks in multiple aspects such as noise [15, 16], concept drift [17], and out of distribution (OOD) [18, 19].

The phenomenon of OOD data and its impact on data-driven (i.g., ML and DL) models has been the subject of extensive investigation due to its frequent occurrence and its challenging nature that causes unexpected complications in ML applications [20–22]. To illustrate the challenge posed by OOD scenarios, let us consider the example of a bike rental company that relies on a data-driven model to predict bike rentals based on past data. For years, the business was only open in nice weather, and the model trained and used on these days—provided highly satisfying results. However, recently, the business owner decided to extend the working days to the hot summer days as well. Unfortunately, the business’s model performing badly, causing a lot of economic harm. A possible explanation for the model’s poor performance on the summer days is that the hot summer days show different dynamics due to factors that are not necessarily taken into consideration in the original modeling and development. Hence, resulting in OOD data that the model fails to ‘understand’. To this end, understanding and addressing the challenges posed by OOD scenarios are critical for ensuring the robustness and reliability of data-driven models across diverse applications [23, 24]. Formally, OOD refers to data instances that deviate significantly from the training data distribution of data-driven models. This definition should be taken with caution as one should be careful not to confuse OOD with the concept drift phenomenon which describes the case where the data changes over time. OOD is a fundamental issue in data-driven modeling as it reveals the weakness of data-driven models which are designed to assume that the training data they provided is ‘well-representing’ the dynamics of the task [25]. Nonetheless, as illustrated by the above example, this condition is strenuous (or even impossible) to satisfy in many practical scenarios.

In order to address this challenge, in this study, we present a novel usage of the well-established symbolic regression (SR) method as a tool to improve ML and DL model’s extrapolation capabilities, making them more robust in terms of OOD. SR is a computational technique used in ML and evolutionary algorithms to automatically discover mathematical expressions that best fit a given dataset [26–28]. Intuitively, SR models have been shown to be less expressive and accurate than ML and DL models when considering in-distribution (ID) performance (i.e. the distribution defined by the training data proposed to the data-driven model) [29, 30]. The so-called ‘under-representation’ of SR models often also results in their ability to capture the main dynamics behind a sampled dataset and therefore has the potential to generalize better for OOD samples [31–33]. By incorporating SR-derived features before applying ML and DL models, we show that ML and DL perform (statistically) similarly (or even better) on in-distributing evaluation while also outperforming OOD evaluation.

We evaluate the proposed method by applying two ‘off-the-shelf’ SR models (QLattice [34] and GPlearn [35], automatic ML Tree-based Pipeline Optimization Tool (TPOT [36]), and automatic DL models (AutoKeras (AK) [37]) on n = 15 real-world datasets from various domains and with various properties. Our analysis shows that the ML and DL performance for the ID evaluation improved by a mean of 2.85% and 11.05% (p < 0.01), respectively, compared to the same models without the SR enhancement. As for the OOD data, the inclusion of the SR-derived feature improved the performance of the ML and DL models by a mean of 3.70% and 10.20% (p < 0.01), respectively.

The remainder of this paper is organized as follows: section 2 provides an overview of related work in the field of OOD, discussing various definitions of OOD and current solutions as well as SR methods with their strengths and limitations. Afterward, section 3 presents our proposed method that includes the data used in the experiments, the SR as a tool to tackle OOD, and the experiment design conducted in this study. Next, section 4 outlines the obtained outcomes. Finally, in section 5, we discuss the implications and possible applications drawn from our results while also discussing the limitations of the study and promising future work.

OOD is a commonly found challenge of data-driven models, in general, and for ML and DL models, in particular, [38–40]. In parallel, SR has been shown as a powerful computational tool that has palatial generalization capabilities for tabular data [41]. In this section, we provide an overview of existing solutions for OOD in ML and DL settings. Subsequently, we provide an overview of SR methods with a focus on their potential as an extrapolator that can be utilized for the OOD challenge.

2.1. Out of distribution in data-driven models

The fundamental premise of data-driven models, in general, and in ML (and DL) models, in particular, is based on the assumption that data will be identically and independently distributed (i.i.d). This means that the training and test data are assumed to come from the same distribution. This assumption often fails to hold in numerous real-world scenarios [42]. In recent years, ML and DL algorithms have become ubiquitous across various domains of life and their usage ‘outside of the lab’, when deployed in real-world settings, often encounter violations of the closed-world and i.i.d assumptions [43–45]. This decline in performance is typically attributed to shifts in data distributions [46]. Currently, there is an active investigation of this phenomenon, commonly referred to ‘OOD’ [47, 48]. A few recent studies have tackled this challenge and estimated the OOD degradation in different datasets [49, 50].

In addressing OOD scenarios within ML and DL, various solutions have been explored [51]. For instance, [52] builds on top of the Risk Extrapolation mathematical framework, as a form of robust optimization over a perturbation set of extrapolated domains, to show that reducing differences in risk across training domains can reduce a model’s sensitivity to a wide range of extreme distributional shifts. Yao et al [53] proposed a simple mixup-based technique that learns invariant predictors via selective augmentation called LISA. Simply put, this method selectively interpolates samples either with the same labels but different domains or with the same domain but different labels in the case of subpopulation shifts (e.g. imbalanced data) and domain shifts. Moreover, [54] extended the task of improving the robustness to OOD by combining an OOD detection mechanism as an inherent part of the method. Namely, the authors propose a margin-based learning framework that exploits freely available unlabeled data in the wild that captures the environmental test-time OOD distributions under both covariate and semantic shifts. Taken jointly, [55] empirically showed that OOD performance is strongly correlated with ID performance for a wide range of models and distribution shifts. In particular, the authors connected the power of this connection to the Gaussian data model revealing that the further a sample from the center of the Gaussian’s defined centroid, the weaker the connection is.

Despite these advancements, challenges persist in effectively addressing OOD in data-driven models. Evaluating OOD detection methods remains a complex task due to the inherent imbalance between ID and OOD samples. Additionally, while existing techniques provide valuable insights, understanding the implications of OOD data in specific application domains, such as medical imaging, autonomous systems, and natural language processing, remains an ongoing area of research. Further advancements are needed to enhance the robustness of models in the presence of OOD data, especially in diverse and complex real-world scenarios.

2.2. Sr

SR can be addressed through diverse techniques, including brute-force search, sparse regression, DL, and genetic algorithms [56, 57]. While each method has its strengths and weaknesses, no single approach dominates the field [27].

Initially, brute-force SR models theoretically have the potential to solve any SR task by exhaustively evaluating all possible equations to find the optimal one [58]. However, in practice, applying brute-force methods often becomes impractical due to the significant computational demands, making them challenging to use even with relatively small datasets. Moreover, these models tend to overfit when dealing with large and noisy data [59], which is frequently encountered in real-world scenarios [60]. In contrast, DL SR models excel at handling noisy data, thanks to the inherent robustness of neural networks against outliers [61]. Nevertheless, empirical evidence suggests that these models have limited generalization capabilities, which restricts their applicability in many contexts [27, 62]. Sparse regression methods have gained popularity by significantly narrowing the search space through sparsity-driven optimization, enabling the discovery of concise models [63]. For example, SINDy [64] employs a Lasso linear model to uncover sparse representations of nonlinear dynamical systems underlying time-series data. The algorithm alternates between a partial least-squares fit and a thresholding step to encourage sparsity. Due to the potential of this method across various domains, it has garnered significant attention, with researchers enhancing its performance by introducing mechanisms to better handle noisy data and select optimal models across varying threshold values [65–67]. Finally, genetic algorithm SR models effectively integrate prior knowledge to constrain the function search space [68]. For instance, SR can be guided by predefined solution shapes [69–72], or by probabilistic models that sample grammar rules governing solution generation [73–77].

SR is known to overfit data less than more complex ML or DL models [30], indicating a greater potential for generalization and OOD performance. SR has been demonstrated to outperform other ML models on small datasets [30]. The concept of ‘under-representation’ in SR models often allows them to capture the primary dynamics of a sampled dataset, potentially improving their ability to generalize to OOD samples [31–33].

3. Methods and materials

3.1. Dataset

We performed the analyses on a benchmark of regression datasets from a recent benchmarking study on automatic ML regression tasks [78]. Specifically, we analyzed a total of 15 datasets from seven studies [79–85]. The only modification we applied to the datasets was the removal of categorical variables, as we focus exclusively on numerical variables since classical SR does not support categorical variables. This modification does not influence the results as these are relative to each other.

3.2. Experimental setup

We examine the ID versus OOD performance of different ML and DL models, under different definitions of OOD. To ensure the robustness of our results, we perform 50 repetitions of the experiment for each of the 15 datasets and for each of the ML or DL models.

In each iteration of each dataset, we begin by splitting the data into ID and OOD using a given OOD metric. We define OOD observations as the 15% highest OOD-metric observations, for a given OOD measure. We then split the ID data into 70% training, 15% validation data, and 15% test data. We then train an ML or DL model on the training data and evaluate its root mean squared error (RMSE) performance on both the ID test data and the OOD data.

We adopted four popular OOD definitions from those proposed in [48]. First, we use the multivariate Z-score. Formally, in the univariate case, the Z-score of an observation from the mean of a univariate normal distribution is derived by taking the mean and standard deviation of one of the input features. For example, a Z-score of 0 would indicate an observation that is equal to the mean score, and a Z-score of 1 would indicate that the observation is one standard deviation away from the mean. The multivariate generalization of the Z-score was first introduced by Mahalanobis in 1936 [86] and is referred to as the Mahalanobis distance. Considering a probability distribution Q over RN, characterized by its mean vector⃗µ = (µ1,µ2,µ3,...,µN)T and a positive-definite covariance matrix S, the Mahalanobis distance dM from a point x = (x1,x2,x3,...,xN)T to Q is defined as:

dM (⃗x,Q) = (⃗x −⃗µ)T S−1 (⃗x −⃗µ).

In the univariate case, the Mahalanobis distance is identical to the Z-score. This definition of multivariate Z-score has been widely used and is probably the most common definition for OOD in regression tasks [87–95]. Second, as a robustness test, we consider a random-feature Z-score. For this case, we randomly (in a uniformly distributed manner) choose one feature in each iteration and define the observation as either in or out of distribution based on its Z-score [96]. Third, as an additional robustness test, we also use a weighted distance inspired by the Kullback–Leibler (KL) divergence metric. KL divergence is a concept from information theory used to measure how one probability distribution diverges from a second, reference probability distribution. In the context of OOD, KL divergence is utilized to quantify the difference between the probability distributions of ID and OOD data [97]. Finally, we use an OOD metric based on the y-sparsity, as performed in the novel work of [49].

3.3. ML and DL models

For the experiment, we adopted several popular ML and DL methods, ranging from the simplest one linear regression (LR) to more advanced models (AK):

  • • TPOT [36]—an automated ML tool that uses genetic algorithms [98] to optimize ML pipelines, including data preprocessing, feature selection, and model selection. As an AutoML library, TPOT performs hyperparameter optimization intrinsically. The only limitation we used was a time limit of 10 min per individual fold run. Time limited model runs are common in tabular benchmarks to ensure a fair comparison between different models [99–101].
  • • AK [37]—an open-source automated ML library built on top of Keras, which automatically searches for the best neural network architecture and hyperparameters for a given dataset. Similar to TPOT, AK performs hyperparameter optimization intrinsically. We limited its runs to at most 200 epochs.
  • • XGBoost (XGB) [102]—an optimized gradient boosting algorithm that is highly efficient and widely used for classification and regression tasks, known for its performance and scalability. Hyperparameter optimization was performed using TPOT, with a 10 min time limit per run.
  • • Random forest (RF) [103]—an ensemble learning method that constructs a multitude of decision trees [104] during training and outputs the mode of the classes or mean prediction of the individual trees for classification or regression tasks, respectively. Hyperparameter optimization was performed using TPOT, with a 10 min time limit per run.
  • • Support vector machine (SVM) [105]—a supervised ML algorithm used for classification and regression tasks, which finds the hyperplane that best separates classes in a high-dimensional space. Hyperparameter optimization was performed using TPOT, with a 10 min time limit per run.
  • • K-nearest neighbors (KNNs) [106]—a non-parametric, instance-based learning algorithm used for classification and regression tasks, where the classification of a data point is determined by a majority vote of its KNNs in the feature space. Hyperparameter optimization was performed using TPOT, with a 10 min time limit per run.
  • • LR [107]—a statistical method used for modeling the relationship between a dependent variable and one or more independent variables by fitting a linear equation to the observed data.

3.4. Integration of SR in ML and DL models

We use the method proposed by [108] to examine a method to improve OOD performance. Namely, The authors propose using SR as a feature engineering technique to enhance ML and DL regression models. Through extensive experiments on synthetic and real-world datasets, they demonstrate that incorporating SR-derived features significantly improves model performance, with gains in RMSE of up to 11.5% on real-world datasets. The study highlights the potential of SR for improving model accuracy and interpretability while reducing the reliance on domain expertise for feature design. In this work, we extend the investigation of this method by exploring the SR-derived feature contribution to ID and OOD separately, rather than the entire dataset at once. We hypothesize that due to the SR’s tendency not to overfit the data, its contribution to ML or DL OOD performance might be even higher compared to ID performance. To examine this hypothesis, we repeat the experiment described in section 3.2 either with or without an SR-derived feature, for both TPOT and AK. AK is used as a model which estimates various DL architechtures, while TPOT evaluates various ML models, including those mentioned above (XGB, RF, KNN, and SVM). We use two different SR models to establish the robustness of this method. First, GPlearn [35] which is a popular open-source Python library that uses genetic algorithms to fit expression trees over provided data such that the expression tree’s prediction is as close to the predicted parameter while also the expression’s complexity is minimal. Second, QLATTICE [34] which is an open-source Python library that searches for expression trees containing only multiply, linear, sine, tanh, and gaussian (unlike the other two libraries that used addition, multiplication, division, subtraction, and inversion). This library first ‘guesses’ the structure of the expression tree followed by a training procedure which both allocates the right expressions to the tree and the weights of their inputs. It uses a genetic algorithm to search over the expression tree structures.

In the GPlearn model, we utilized the four basic operators: addition, subtraction, multiplication, and division. Additionally, we conducted 200 trials using the default operators in GPlearn, finding no significant variations (results not shown). The modeling was carried out over 50 generations, with the model being tested six times using various parsimony coefficients (0.005, 0.01, 0.02, 0.03, 0.04, 0.05). For each test, the parsimony coefficient that yielded the lowest RMSE on the training data was chosen. Other than these hyperparameters, we maintained all other settings at GPlearn’s default values. In the Feyn model we evaluated various complexity values (5, 10, 15, 20, 25, 30) and chose the highest performing value. Other than that, we used the default hyperparameters of the library.

4. Results

In this section, we present the results of the experiments. First, we outline the performances of various ML and DL models in ID compared to OOD scenarios. Afterwards, we examine the integration of SR-derived features into the performance of ML and DL models in ID and OOD scenarios.

4.1. ML and DL performance in OOD data

Figure 1 presents an example of the results in a single dataset. In this example, Z-scores are determined based on the temperature variable. The target feature, hourly bike rentals, is predicted by the XGB model trained on ID data and tested on both ID and OOD observations. The horizontal axis represents the Z-score and the vertical axis represents RMSE for each observation. As expected, RMSE performance is best for lower Z-scores and deteriorates for higher Z-scores.

Figure 2 and table 1 present a complementary analysis for the entire benchmark of 15 datasets. Figure 2 uses multivariate Z-score as an OOD metric; additional metrics are presented in the appendix. To present datasets of different scales in one figure, we normalize the RMSE scores by dividing them by the corresponding dataset’s mean absolute y (target feature) value. Figure 3 presents the histograms of ID and OOD errors in each model and OOD metric. The results of our experiments reveal several intriguing patterns. Firstly, we observed that the performance of various ML and DL models, including XGB, RF, and TPOT, tends to degrade as the Z-score increases, indicating a decline in model accuracy for OOD data. Table 1 provides both the difference (%) between OOD and ID data, and the slope of a LR in the RMSE versus OOD figures. Both the difference and the slope are always positive, demonstrating the robustness of this result. This finding is consistent with the expected behavior of ML models when encountering data that deviate significantly from the training distribution.

Interestingly, while all models experienced a decrease in performance for OOD samples, the extent of this degradation varied across different models. For instance, RF demonstrated a relatively better resilience to OOD data compared to other models, including XGB. This could be attributed to the inherent robustness of ensemble methods like RF, which combine predictions from multiple decision trees to reduce variance and improve generalization.

Another noteworthy observation is the similarity in the shape of the RMSE versus Z-score curves across different models, albeit with varying magnitudes of RMSE. This suggests that certain OOD samples pose a consistent challenge to all models, regardless of their underlying algorithms. Identifying the characteristics of these particularly challenging OOD samples could provide valuable insights for designing more robust ML models. We also observe that the performance trend with respect to the Z-score is not monotonic, indicating that certain OOD observations are not as challenging to predict as others.

RMSE scores as a function of temperature Z-score
Figure 1. RMSE scores as a function of temperature Z-score. The results are based on 50 runs of the XGBoost model trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest Z-scores, based exclusively on the temperature feature. The ID test data are randomly chosen from the remaining 85% observations.

Figure 3 shows the difference between ID and OOD performance as histograms. While some OOD data align with the distribution of ID samples, the OOD errors generally exhibit a much larger tail, indicating that some OOD predictions had significantly larger errors. In some instances, the OOD errors extended beyond merely forming a tail; they were also centered around higher error values. This phenomenon could stem from the integration of multiple datasets, where, in certain datasets, OOD errors deviated markedly from ID errors and were centralized around higher error values.

Lastly, our findings reveal that although the AK model typically exhibited lesser performance compared to other ML models, in line with prior studies, it surpassed even the top-performing ML models in certain high OOD scenarios. This unexpected outcome merits further investigation.

4.2. Integration of SR in ML and DL models to improve OOD performance

In this section, we examine the contribution of integrating an SR-derived feature before the application of ML or DL models [108]. As demonstrated in figure 2, SR does not perform as well as other ML models (such as XGB and RF) in ID observations. However, its performance in OOD observations is only slightly inferior compared to these models, suggesting that features derived from SR could be beneficial for ML models in OOD scenarios.

In table 2 we further evaluate the robustness of this method using two different SR models, and four different measures of OOD. We find that the improvement in performance is robust and holds in all 24 configurations (two different SR models, three different OOD metrics, and two different ML and DL models, either in or out of distribution). Furthermore, in most cases (9 of 12) the relative improvement in OOD is larger than the relative improvement in ID, consistent with our hypothesis that the SR-derived feature contributes more in this type of data.

The integration of SR-derived features into ML and DL models further emphasizes the potential of SR as a tool for enhancing model performance in OOD scenarios. The improvement in both ID and OOD performance suggests that SR-derived features can capture underlying patterns in the data that are not easily detected by conventional ML models. This could be particularly useful in applications where interpretability and generalization to novel scenarios are crucial.

Overall, our results highlight the importance of considering OOD performance in the evaluation of ML models and suggest that incorporating SR-derived features can be an effective strategy to improve model performance, robustness, and interpretability.

Normalized RMSE (NRMSE) scores as a function of multivariate Z-score
Figure 2. Normalized RMSE (NRMSE) scores as a function of multivariate Z-score. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest multivariate Z-scores, based exclusively on the temperature feature. The ID test data are randomly chosen from the remaining 85% observations. For convenience, the x and y axes are limited to the 99th percentile of the data. The full figure is presented in the appendix.
Table 1. The table presents the median relative performance of OOD compared to ID of the various models, as well as the slope of a linear fit for the RMSE versus OOD figures. Higher diff (%) values and higher slopes indicate inferior OOD performance.
OOD TypeMultivariate-ZRandom feature Z-scoreKL-divergenceSparse-y
ModelDiff (%)SlopeDiff (%)SlopeDiff (%)SlopeDiff (%)Slope
LR1450.025490.026410.000 2391100.10
XGB690.013330.016470.000 1401730.16
RF890.016270.017410.000 1481410.19
SVM630.02280.015210.000 233870.14
KNN780.018420.021550.000 1431720.19
TPOT770.014380.013360.000 2061390.15
AK670.102420.054440.000 5861310.12
Feyn850.020410.021490.000 1181440.12

5. Discussion

In this study, we explored the performances of various ML and DL models in regression task tests on both ID and OOD samples. To ensure the robustness of our results, we performed extensive modeling with various models (XGB, RF, KNN, SVM, LR, TPOT, AK, gplearn, QLattice), a relatively large benchmark of n = 15 real-world datasets, and multiple repetitions (n = 50 each), resulting in over 40000 model runs. We summarized our results by displaying and comparing the obtained RMSE values for each model as a function of multivariate Z-scores, random feature Z-score, and KL-Divergence.

Based on the obtained results, we found that XGB and RF are not only the best-performing models, but they also maintain a relatively high performance in OOD samples. Also, the DL models performed surprisingly well in OOD data, despite their relatively lower performance in ID data. This observation can be attributed to the DL models’ ability to transform the feature space effectively. Even when the original feature space suggests sparsity or outliers, these models can create a learned feature space where such sparse samples are denser, thus not as sparse in the transformed context. While the various models obtained substantially different RMSE scores in both ID and OOD samples, the shape of the RMSE versus Z-scores figure remained remarkably similar between all models, with different stretch factors in the vertical axis. As such, we conclude that some OOD samples are harder to predict than others, making it difficult for all models to perform well in their prediction. Finally, we note that this shape is not monotonous, meaning that in some cases there are samples that are harder to predict than others, although their OOD-metric is lower. This property raises a question about the current OOD definitions commonly used.

Histograms of Normalized RMSE scores for ID and OOD observations using various OOD metrics, for each model
Figure 3. Histograms of Normalized RMSE scores for ID and OOD observations using various OOD metrics, for each model. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest multivariate Z-scores, based exclusively on the temperature feature. The ID test data are randomly chosen from the remaining 85% observations. We limit the X-axis to 1 to enhance readability (the full range is presented in figure A.9).
Table 2. The table summarizes the relative contribution of the SR-derived feature to the TPOT and AK models, separated into ID and OOD performance using four different OOD metrics.
TPOT
Feyngplearn
OOD typeID (%)OOD (%)ID (%)OOD (%)
Multivariate Z-score2.564.741.922.80
Random feature Z-score3.154.602.745.22
KL-divergence3.501.253.233.60
Sparse-y3.223.225.591.92
AK
Feyngplearn
OOD typeID (%)OOD (%)ID (%)OOD (%)
Multivariate Z-score14.3210.348.899.82
Random feature Z-score12.3113.709.139.33
KL-divergence12.424.719.3013.14
Sparse-y11.0916.087.0917.35

To improve the OOD performance of these models while preserving the ID performance, we have introduced a novel approach to predictions in OOD, based on the SR method. Namely, we add additional features to the models by solving an SR task between the input and target variables, before repeating the same task using the ML (DL) model. This outcome undeniably showcases the capacity of features derived from SR to markedly improve the OOD predictive performance of data-driven models. These results indisputably highlight how features obtained through SR have the potential to significantly enhance the predictive capabilities of data-driven models as these provide these models with a generalized yet under-representing representation of the data that these models can use to improve their generalization capabilities. Our analysis indicates that for ID evaluation, the performance of ML and DL models enhanced by a mean of 2.85% and 11.05%, respectively (p < 0.01), with the addition of the SR enchantment. Regarding the OOD data, incorporating the SR-derived feature resulted in a performance improvement of 3.70% and 10.20%, respectively, for ML and DL models (p < 0.01).

Moreover, as illustrated in table 2, the proposed SR method can statistically improve the OOD performance with different definitions of OOD. Furthermore, table 2 supports the fact that this outcome is preserved on a large number of datasets from different domains and diverse properties. As demonstrated from the table, the SR-derived feature improves the ML and DL models while the SR model by itself produces an inferior RMSE score on the same OOD samples. Thus, the combination of the two methods—SR and ML/DL, obtained the highest performance. In addition, the inclusion of SR-derived features can enhance the interpretability of ML models and help prevent overfitting.

While our study underscores the potential of combining SR and ML (DL) models to improve their OOD performance, there are several limitations to this study. First, the suitability of SR may differ based on dataset characteristics, prompting inquiries into its adaptability across diverse domains and the computational properties inherent in the dataset [26, 27]. Second, this study employed two SR methods - (QLattice [34] and GPlearn [35]). Further exploration of other SR models could yield slightly different results [27]. This raises the computational question of finding the best SR model based on dataset characteristics, and optimizing its effectiveness for the OOD task [7]. Finally, the study highlights the importance of estimating confidence intervals in ML models, especially in OOD predictions where larger estimation errors are expected.

Data availability statement

All data that support the findings of this study are included within the article (and any supplementary files).

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Appendix

Table A.1 presents the descriptive statistics of the datasets used in the study. Figures A.1, A.3, A.5, and A.7 break down the information presented in figure 2 into separate datasets. Figures A.2, A.4, A.6, and A.8 do the same and present the entire range, including outliers. Tables A.2–A.5 summarize the results for the multivariate Z-score, random feature Z-score, KL-Divergence, and y-sparsity, respectively.

Table A.1. Descriptive statistics of datasets.
Dataset# Rows# Columns# Categorical
columns
Average feature
entropy
Kurtosis
Matbench3122309.2786372.6246
Su 1122864.69433.1692
Su 2136654.8412−1.3164
Koya 1 (Rup28)1101184.67491951.502
Koya 1 (Cte)1101186.43521420.308
Koya 2 (Split28)1101178.99951160.730
Koya 2 (Poisson28nu)1101188.1476720.4195
Koya 2 (Elast28)1101175.85951016.892
Koya 2 (Comp28)1101179.7617−2046.0
Huang (FS)1141094.46993.1935
Huang (CS)1141094.46563.1935
Guo 1 (Ys)63 16228310.4080861.2404
Guo 2 (Ts)63 16228310.4088861.2404
Guo 3 (El)63 16228310.4089861.2404
Bachir112424.6315−0.6539
Normalized RMSE scores as a function of multivariate Z-score
Figure A.1. Normalized RMSE scores as a function of multivariate Z-score. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest multivariate Z-scores. The ID test data are randomly chosen from the remaining 85% observations. For convenience, the x and y axes are limited to the 99th percentile of the data.
Normalized RMSE scores as a function of multivariate Z-score
Figure A.2. Normalized RMSE scores as a function of multivariate Z-score. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest multivariate Z-scores. The ID test data are randomly chosen from the remaining 85% observations. This figure is similar to figure A.1, but presents the full range of the data including outliers.
Normalized RMSE scores as a function of random feature Z-score
Figure A.3. Normalized RMSE scores as a function of random feature Z-score. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest OOD score. The ID test data are randomly chosen from the remaining 85% observations. For convenience, the x and y axes are limited to the 99th percentile of the data.
Normalized RMSE scores as a function of random feature Z-score
Figure A.4. Normalized RMSE scores as a function of random feature Z-score. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest OOD score. The ID test data are randomly chosen from the remaining 85% observations. This figure is similar to figure A.3, but presents the full range of the data including outliers.
Normalized RMSE scores as a function of KL-Divergence
Figure A.5. Normalized RMSE scores as a function of KL-Divergence. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest OOD score. The ID test data are randomly chosen from the remaining 85% observations. For convenience, the x and y axes are limited to the 99th percentile of the data.
Normalized RMSE scores as a function of KL-Divergence
Figure A.6. Normalized RMSE scores as a function of KL-Divergence. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest OOD score. The ID test data are randomly chosen from the remaining 85% observations. This figure is similar to figure A.5, but presents the full range of the data including outliers.
Normalized RMSE scores as a function of the sparse-y OOD metric
Figure A.7. Normalized RMSE scores as a function of the sparse-y OOD metric. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest OOD score. The ID test data are randomly chosen from the remaining 85% observations. For convenience, the x and y axes are limited to the 99th percentile of the data.
Normalized RMSE scores as a function of the the sparse-y OOD metric
Figure A.8. Normalized RMSE scores as a function of the the sparse-y OOD metric. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest OOD score. The ID test data are randomly chosen from the remaining 85% observations. This figure is similar to figure A.7, but presents the full range of the data including outliers.
Table A.2. The table presents the RMSE of OOD and ID performances of the various models in each of the datasets, split by the multivariate Z-score.
LRXGB
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——91285.026——910.038
Combined (median)——1450.025——690.013
Koya1 (rup28)57.27764.46712.550.00356.61464.53914.0−0.0
Koya2 (comp28)450.86443.442−1.650.01543.184496.164−8.660.009
Koya2 (elast28)467 825.449488 019.5514.320.007481 544.811492 333.5212.240.004
Koya2 (poisson28nu)0.0140.02473.990.0290.0160.02236.790.012
Koya2 (split28)36.27648.39933.420.01842.51643.963.390.016
Koya1 (cte)3.40×10−78.34×10−7145.350.024.24×10−76.81×10−760.450.017
Matbench215.466556.608158.330.06130.749222.89970.480.013
Bachir11.4349.372−18.03−0.0075.77810.74986.040.144
Guo (ys)35.983138.874285.950.02728.21476.711171.890.01
Guo (ts)26.71783.622212.990.01521.24273.1244.120.009
Guo (el)4.64313.16183.420.0253.8465.55344.380.003
Su21.3892.3871.370.0291.262.12568.680.034
Su11.0585.239395.00.1391.2574.341245.420.091
Huang (FS)1.473992.83367 284.2334.81.4144.478216.710.137
Huang (CS)12.018189.56368 087.2540.21510.9923.029109.550.068
RFSVM
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——920.04——11920.738
Combined (median)——890.016——630.022
Koya1 (rup28)56.75162.349.850.053.82262.73316.560.001
Koya2 (comp28)478.083532.84411.450.02485.09568.99317.30.023
Koya2 (elast28)459 492485 2285.60.004522 922517 440−1.050.004
Koya2 (poisson28nu)0.0150.02142.250.0120.0140.02576.840.036
Koya2 (split28)37.35441.62311.430.01937.05341.26211.360.013
Koya1 (cte)3.69×10−77.37×10−799.460.0169×10−710×10−71.850.01
Matbench111.912220.71697.220.011262.813279.1346.210.0
Bachir6.24412.08293.490.17612.41318.59549.810.164
Guo (ys)27.98969.454148.150.00872.059158.075119.370.018
Guo (ts)21.10166.757216.360.00764.672206.387219.130.022
Guo (el)3.8275.2637.460.0036.08511.06781.880.012
Su21.3332.30472.830.0421.6962.75962.680.044
Su11.2894.433243.970.0941.0674.832352.980.123
Huang (FS)1.3924.253205.490.1281.84924.221209.551.13
Huang (CS)10.1619.17688.740.05312.138191315 6669.469
(Continued.)
Table A.2. (Continued.)
KNNTPOT
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——1030.04——1050.043
Combined (median)——780.018——770.014
Koya1 (rup28)64.87672.72712.10.00959.29372.13221.650.003
Koya2 (comp28)550.218656.37819.290.019512.173520.5911.640.014
Koya2 (elast28)499 154533 0576.790.005504 126480 196−4.75−0.002
Koya2 (poisson28nu)0.0160.02243.50.0140.0150.02358.030.022
Koya2 (split28)38.91240.1673.230.01339.72343.77210.190.018
Koya1 (cte)3.79×10−77.06×10−786.180.0144.01×10−77.78×10−793.90.014
Matbench131.858234.48277.830.01108.089199.44284.520.01
Bachir9.48610.0335.760.0666.08211.1983.970.155
Guo (ys)53.175144.964172.620.01828.5574.545161.10.009
Guo (ts)46.839185.478295.990.02521.21672.282240.690.008
Guo (el)5.08510.26101.780.0123.8775.47541.240.004
Su21.3792.32268.390.0511.3172.08358.110.039
Su11.3935.916324.840.1461.1264.869332.250.116
Huang (FS)1.5234.181174.50.1081.3475.577314.080.164
Huang (CS)11.5329.706157.660.08812.59622.35677.480.064
AutoKerasFeyn
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——432252.958——1540.048
Combined (median)——670.102——850.020
Koya1 (rup28)243.99246.8011.150.00359.34871.72420.850.005
Koya2 (comp28)4112.3934097.182−0.370.002482.535496.082.810.015
Koya2 (elast28)4332 499.9714605 584.7886.30.043504 360.666577 864.85414.570.014
Koya2 (poisson28nu)0.0770.13575.060.1850.0160.02239.310.019
Koya2 (split28)118.265132.76312.260.00443.41551.54618.730.02
Koya1 (cte)0.0810.12351.663789.9310.00.078.370.016
Matbench248.652902.724263.050.102192.619355.73184.680.017
Bachir8.91513.91156.040.1666.67315.051125.560.245
Guo (ys)35.82888.869148.050.01134.37484.267145.150.01
Guo (ts)26.97268.479153.890.00626.29391.404247.630.022
Guo (el)4.6277.71166.630.0054.65213.891198.620.028
Su23.0854.96560.930.111.0981.90973.850.038
Su11.1767.181510.410.1381.1258.045614.930.099
Huang (FS)1.74656.6663145.482.071.4564.251192.010.14
Huang (CS)15.713317.8891923.061.59613.17773.301456.290.031
Table A.3. The table presents the RMSE of OOD and ID performances of the various models in each of the datasets, split by the random feature Z-score.
LRXGB
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——2072.031——470.030
Combined (median)——490.026——330.016
Koya1 (rup28)50.63454.9218.470.00956.71275.56933.250.001
Koya2 (comp28)423.815363.421−14.25−0.003522.017441.558−15.41−0.012
Koya2 (elast28)456 254.615409 750.38−10.190.007514 543.056520 162.5381.09−0.003
Koya2 (poisson28nu)0.0170.015−11.390.0050.0170.0187.750.009
Koya2 (split28)33.84630.834−8.9−0.00139.67637.84−4.630.004
Koya1 (cte)0.00.049.10.0130.00.010.740.0
Matbench249.828633.7153.650.059145.133267.13884.070.028
Bachir10.39513.55130.360.0325.76212.963124.980.048
Guo (ys)41.09473.74379.450.04929.17248.28865.530.018
Guo (ts)30.12946.03552.790.02622.25937.65969.190.016
Guo (el)4.7239.21395.090.0733.7974.83327.270.008
Su21.5661.6092.750.0151.3031.60322.970.035
Su12.083.14651.220.1691.6393.624121.180.149
Huang (FS)10.387114.221999.636.2491.5693.1399.530.103
Huang (CS)203.2463511.0581627.4923.76910.34616.70461.460.052
RFSVM
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——440.036——2690.656
Combined (median)——270.017——80.015
Koya1 (rup28)53.31270.08531.460.01749.54253.3677.720.008
Koya2 (comp28)459.346476.0543.640.003432.901378.661−12.53−0.001
Koya2 (elast28)483 078.17429 914.285−11.010.003530 562.261492 178.993−7.230.008
Koya2 (poisson28nu)0.0150.0161.57−0.0010.0160.015−9.14−0.005
Koya2 (split28)34.13131.002−9.170.00434.64731.477−9.150.005
Koya1 (cte)0.00.05.580.0070.00.0−0.74−0.005
Matbench133.318247.57185.70.027243.135319.64531.470.021
Bachir5.60912.969131.230.0910.20616.83164.910.089
Guo (ys)28.90244.11752.650.01771.19794.733.010.024
Guo (ts)21.85445.62108.750.02560.44984.27839.420.027
Guo (el)3.7884.73124.890.0096.5587.73217.90.015
Su21.391.61215.980.0261.8131.8582.520.006
Su11.5022.97698.20.1492.0732.1292.730.07
Huang (FS)1.5033.006100.00.0972.53815.202498.940.616
Huang (CS)11.11514.11226.960.06218.307635.3213370.38.967

(Continued.)

Table A.3. (Continued.)
KNNTPOT
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——430.032——580.027
Combined (median)——420.021——380.013
Koya1 (rup28)55.96379.37741.840.01856.48673.2429.660.003
Koya2 (comp28)565.678617.1339.10.002505.586476.506−5.750.003
Koya2 (elast28)496 947.967509 092.6082.440.014511 894.47473 018.137−7.59−0.008
Koya2 (poisson28nu)0.0160.016−0.41−0.0050.0160.01915.670.007
Koya2 (split28)37.86336.041−4.810.00643.24140.577−6.160.004
Koya1 (cte)0.00.00.340.0070.00.038.480.006
Matbench136.126293.502115.610.042117.784244.567107.640.022
Bachir8.71913.60756.060.0186.04212.457106.170.058
Guo (ys)58.63991.69556.370.02829.46553.59281.880.013
Guo (ts)59.106100.70470.380.03722.26559.898169.020.017
Guo (el)5.3017.30837.850.0213.8554.87826.510.009
Su21.4671.8123.410.0341.3651.69324.080.029
Su11.7953.599100.460.121.5383.954157.140.123
Huang (FS)1.5982.69268.530.0721.6383.0586.130.067
Huang (CS)10.24316.51361.210.06511.27416.55946.880.055
AutoKerasFeyn
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——1 × 1071 × 104——610.038
Combined (median)——420.054——410.021
Koya1 (rup28)225.177233.6223.75−0.00455.54278.1140.630.009
Koya2 (comp28)4228.5364123.032-2.5−0.0042459.972438.038−82.190.001
Koya2 (elast28)4340 893.7954422 056.9291.870.013508 630.122494 836.315−2.710.008
Koya2 (poisson28nu)0.0810.11339.480.0950.0170.0216.380.005
Koya2 (split28)115.087129.21812.280.00446.40244.637−3.8−0.0
Koya1 (cte)0.0750.10641.662353.0940.00.029.050.01
Matbench249.45948 875 981.0191 × 1088146.092192.066399.734108.120.03
Bachir8.22615.08683.410.065.88115.878170.00.093
Guo (ys)37.37871.791.820.01437.3460.0260.740.021
Guo (ts)28.57455.41493.930.00929.04661.953113.290.032
Guo (el)4.6366.24134.610.0094.715.54417.70.008
Su23.2394.17228.790.0541.3021.58621.840.047
Su11.8095.299192.940.2381.8594.741154.990.19
Huang (FS)2.6161 × 1071 × 109210 373.1871.7163.11381.430.094
Huang (CS)24.6351 × 1071 × 10872 544.16212.6837.312194.250.029
Table A.4. The table presents the RMSE of OOD and ID performances of the various models in each of the datasets, split by the KL-divergence.
LRXGB
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined
(mean)
——88240.063 922——650.056 752
Combined
(median)
——410.000 239——470.000 140
Koya1 (rup28)53.59681.89452.80.000 23955.41689.79162.030.000 336
Koya2 (comp28)494.541286.23−42.12−0.000135533.225368.64−30.87−8.6 × 10−5
Koya2 (elast28)479 476.808549 122.09714.530.000 275483 565.253591 643.13622.350.000 227
Koya20.0160.0160.01−0.0001330.0170.01912.69−1.7 × 10−5
(poisson28nu)
Koya2 (split28)35.65547.91634.390.000 25537.49954.92946.480.000 345
Koya1 (cte)0.00.0187.010.000 2240.00.075.670.00 014
Matbench215.207389.30380.90.587 942141.428332.284134.950.847 855
Bachir9.04218.646106.220.000 8584.73918.067281.280.001 162
Guo (ys)42.13150.34719.5−3 × 10−630.04729.851−0.65−6 × 10−6
Guo (ts)30.46838.68826.989 ×10−622.86622.479−1.692 ×10−6
Guo (el)4.6775.1810.751.2×10−53.8114.20210.261.9×10−5
Su21.4982.10540.590.000 3661.1822.25190.460.001 079
Su12.2093.19644.67−1.9 × 10−51.8482.46533.36−4.6 × 10−5
Huang (FS)1.706565.04933 029.390.089 6371.5434.417186.180.00 024
Huang (CS)12.95912 810.71298 756.580.279 30711.54616.93546.672.4×10−5
RFSVM
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined——660.047 045——14580.038 432
(mean)
Combined
(median)
——410.000 148——210.000 233
Koya1 (rup28)54.33695.45175.670.00 03957.12283.21745.680.000 233
Koya2 (comp28)511.465334.74−34.55−9.1 × 10−5501.633313.232−37.56−0.000121
Koya2 (elast28)471 463.068607 963.50428.950.000 297483 256.245584 269.6620.90.000 318
Koya2
(poisson28nu)
0.0160.01816.35−2.8 × 10−50.0160.016−1.52−0.000138
Koya2 (split28)37.64352.9940.770.000 32238.19847.92925.470.000 246
Koya1 (cte)0.00.0115.90.000 1480.00.0−1.24−0.000176
Matbench129.779299.027130.410.701 805236.187470.42599.170.510 919
Bachir4.4615.834255.040.001 1169.14918.444101.60.000 905
Guo (ys)29.86130.0780.732 ×10−673.02170.613−3.3−5.7 × 10−5
Guo (ts)22.67922.212−2.062 ×10−658.20957.093−1.92−2.6 × 10−5
Guo (el)3.8114.210.222.4×10−56.4937.0198.095.4×10−5
Su21.3552.69398.710.001 4671.652.48650.730.000 761
Su11.6692.76965.96−1.2 × 10−51.771.8122.41−4.7 × 10−5
Huang (FS)1.4094.432214.640.000 2061.89436.5561829.870.00 525
Huang (CS)11.2738.811−21.852.9×10−513.2182621.19619 730.950.058 356

(Continued.)

Table A.4. (Continued.)
KNNTPOT
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined——680.061 217——1600.059 058
(mean)
Combined
(median)
meanmean550.000 143meanmean360.000 206
Koya1 (rup28)62.68882.76432.030.000 18756.879102.46580.140.0004
Koya2 (comp28)578.198382.879-33.78-0.000 165526.534373.79−29.01−6.4 × 10−5
Koya2 (elast28)480 946.395747 342.21255.390.000 489504 471.706626 451.26424.180.000 408
Koya2
(poisson28nu)
0.0150.0233.05−4 × 10−50.0170.0189.02−6 × 10−5
Koya2 (split28)39.26855.23740.670.000 33640.10754.68236.340.000 356
Koya1 (cte)0.00.0100.290.000 1520.00.0142.820.000 206
Matbench147.468376.143155.070.914 724135.697338.709149.610.876 538
Bachir7.43618.88153.910.000 7854.87517.912267.390.001 234
Guo (ys)62.46479.69827.599.1×10−530.24830.711.534 ×10−6
Guo (ts)63.61480.39126.377.2×10−523.18223.5751.75 ×10−6
Guo (el)5.4177.15532.070.000 1433.8544.30111.62.7×10−5
Su21.4512.65282.730.001 2811.073.332211.410.003 348
Su11.973.65785.612.6×10−56.5972.351−64.36−0.000114
Huang (FS)1.5494.095164.420.000 1411.5564.254173.460.000 119
Huang (CS)9.53615.02757.583.4×10−511.383169.7651391.380.003 462
AutoKerasFeyn
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined——3731.912 439——1150.261 695
(mean)
Combined
(median)
meanmean440.000 586meanmean490.000 118
Koya1 (rup28)219.613196.231−10.65−0.00030257.59292.5560.70.000 359
Koya2 (comp28)4159.873904.728−6.13−0.000453494.887422.105−14.713.2×10−5
Koya2 (elast28)4287 209.1814671 163.1518.960.000 586523 663.112674 246.60428.760.000 413
Koya20.0790.15292.70.002 4890.0170.01912.27−1 × 10−6
(poisson28nu)
Koya2 (split28)115.916105.641−8.86−0.00028944.12160.23436.520.000 455
Koya1 (cte)0.080.11543.7227.39810.00.0119.920.000 118
Matbench259.583576.712122.171.26 358210.78815.999287.133.918 635
Bachir6.77622.286228.930.001 3964.61334.818654.730.0023
Guo (ys)39.11743.15710.331.7×10−539.2742.8549.13−1 × 10−6
Guo (ts)29.97832.1177.144 ×10−631.02831.3491.031 ×10−6
Guo (el)4.6135.23513.483.4×10−54.6755.1289.72.9×10−5
Su22.7447.545174.940.006 8531.043.387225.620.002 808
Su11.725.064194.440.000 1161.8242.71248.727 ×10−6
Huang (FS)1.72454.5463064.20.008 2131.6084.477178.390.000 286
Huang (CS)16.803295.7121659.860.006 24813.51822.49366.39−1.1 × 10−5
Table A.5. The table presents the RMSE of OOD and ID performances of the various models in each of the datasets, split by the y-sparsity metric.
LRXGB
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——116.160.13——211.830.14
Combined (median)——110.290.10——173.280.16
Koya1 (rup28)51.92369.49633.840.02256.5453.491−5.390.004
Koya2 (comp28)416.543674.13161.840.04500.504595.64919.010.025
Koya2 (elast28)458 105.225716 480.22256.40.075500 054.953725 899.1745.160.07
Koya2 (poisson28nu)0.0170.02446.950.0510.0170.02545.580.06
Koya2 (split28)30.84444.81145.280.03834.48435.793.790.024
Koya1 (cte)0.00.0110.290.030.00.0173.290.046
Matbench166.598521.366212.950.179118.406541.399357.240.212
Bachir9.87715.50857.010.1476.22912.812105.680.161
Guo (ys)33.25591.541175.270.12624.131108.242348.570.199
Guo (ts)24.31967.716178.450.05718.897140.13641.560.205
Guo (el)4.0028.431110.670.1023.3687.983137.010.113
Su21.3033.037133.110.1071.143.616217.360.157
Su128.72423.908−16.770.5631.1925.455357.660.34
Huang (FS)1.3264.897269.470.251.1554.711307.890.242
Huang (CS)8.99933.09267.70.2286.6434.731423.060.273
RFSVM
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——228.330.15——119.320.13
Combined (median)——140.740.19——86.990.14
Koya1 (rup28)55.33961.82911.730.0255.05772.74332.120.022
Koya2 (comp28)455.284662.70545.560.044393.333711.19180.810.048
Koya2 (elast28)468 814.593712 572.25551.990.078551 326.687783 887.96142.180.071
Koya2 (poisson28nu)0.0160.02554.840.0620.0170.02547.760.056
Koya2 (split28)33.77742.84826.860.04232.0844.18337.730.037
Koya1 (cte)0.00.0138.240.0310.00.0−7.21−0.016
Matbench119.012551.028363.00.215171.203597.154248.80.208
Bachir5.88213.987137.790.18810.6118.5274.550.202
Guo (ys)23.925109.193356.40.263.354173.399173.70.273
Guo (ts)18.825141.84653.490.20558.791212.55261.540.27
Guo (el)3.3488.061140.740.1175.43212.784135.330.178
Su21.2154.328256.250.1951.3973.212129.90.114
Su11.1565.432369.970.3421.8973.54786.990.142
Huang (FS)1.0944.417303.90.2252.0034.697134.530.194
Huang (CS)5.67534.863514.320.288.62635.473311.220.257

(Continued.)

Table A.5. (Continued.)
KNNTPOT
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——185.330.17——209.120.14
Combined (median)——171.790.19——139.040.15
Koya1 (rup28)60.29159.101−1.970.01356.57767.50619.320.022
Koya2 (comp28)557.169815.12246.30.047433.806673.17155.180.037
Koya2 (elast28)488 046.395690 395.19341.460.067488 660.942636 603.45830.280.057
Koya2 (poisson28nu)0.0170.02336.570.050.0170.02445.050.056
Koya2 (split28)36.62441.06112.110.03535.07341.7318.980.032
Koya1 (cte)0.00.080.510.0220.00.088.190.026
Matbench115.833537.424363.970.21108.942537.54393.420.212
Bachir8.11222.049171.790.296.6912.76590.80.154
Guo (ys)45.871157.53243.420.26524.51110.131349.330.201
Guo (ts)41.451204.39393.090.27619.216142.283640.440.205
Guo (el)4.53211.122145.410.1563.4018.131139.040.115
Su21.3634.401222.80.1941.2333.792207.660.162
Su11.3376.35375.110.3891.1965.199334.510.321
Huang (FS)1.3484.45230.190.2241.2844.505250.830.227
Huang (CS)7.6339.617419.250.2895.96834.244473.830.269
AutoKerasFeyn
IDOODDiff (%)SlopeIDOODDiff (%)Slope
Combined (mean)——127.94-61.78——179.060.13
Combined (median)meanmean131.450.12meanmean144.140.12
Koya1 (rup28)346.3276.136−20.26−0.0846.7732.073373.720.022
Koya2 (comp28)4871.8935383.75710.510.0811.2574.696273.430.037
Koya2 (elast28)4353 292.7894835 949.20511.090.0941.1294.633310.150.08
Koya2 (poisson28nu)0.0990.059−40.16−0.1091.0653.104191.280.058
Koya2 (split28)149.58119.275−20.26−0.05630.689110.225259.160.046
Koya1 (cte)0.0860.0937.66−928.81224.795142.926476.430.042
Matbench231.382583.454152.160.1923.9469.233133.970.193
Bachir8.71313.11750.560.1310.00.0144.140.085
Guo (ys)30.7120.782293.430.20856.14372.58629.290.193
Guo (ts)24.481103.0320.730.121435.098682.90556.950.201
Guo (el)3.9719.19131.460.122525 721.637775 508.47647.510.127
Su22.4717.966222.430.3410.0160.02553.290.117
Su11.356.885409.920.38235.38954.76154.740.292
Huang (FS)2.8747.591164.150.307150.481543.659261.280.255
Huang (CS)10.84735.328225.710.2366.9528.38720.640.253
Full histograms of Normalized RMSE scores for ID and OOD observations using various OOD metrics, for each model
Figure A.9. Full histograms of Normalized RMSE scores for ID and OOD observations using various OOD metrics, for each model. The results are based on 50 runs of the different models for each of the 15 datasets in the benchmark, trained on ID data, and tested on both ID and OOD observations. The OOD observations are defined as those with the 15% highest multivariate Z-scores, based exclusively on the temperature feature. The ID test data are randomly chosen from the remaining 85% observations. The horizontal axis is presented using log-scale.

ORCID iDs

Assaf Shmuel https://orcid.org/0000-0002-1794-9381 Oren Glickman https://orcid.org/0009-0000-5158-7372 Teddy Lazebnik https://orcid.org/0000-0002-7851-8147

Article notes

Publication history
Received 10 June 2024 · Accepted 19 December 2024 · Published 6 January 2025
Keywords
  • data-driven model generalization
  • out of distribution
  • feature engineering
  • symbolic regression
  • machine learning robustness

References

  1. Kutz J N 2017 Deep learning in fluid dynamics J. Fluid Mech. 814 1–4 doi:10.1017/jfm.2016.803
  2. Reichstein M et al 2019 Deep learning and process understanding for data-driven earth system science Nature 566 195–204 doi:10.1038/s41586-019-0912-1
  3. Alzubaidi L, Zhang J, Humaidi A J, Al-Dujaili A, Duan Y, Al-Shamma O, Santamaría J, Fadhel M A, Al-Amidie M and Farhan L 2021 Review of deep learning: concepts, cnn architectures, challenges, applications, future directions J. Big Data 8 1–74 doi:10.1186/s40537-021-00444-8
  4. Raissi M and Karniadakis G E 2018 Hidden physics models: machine learning of nonlinear partial differential equations J. Comput. Phys. 357 125–41 doi:10.1016/j.jcp.2017.11.039
  5. Virgolin M, Wang Z, Alderliesten T and Bosman P A N 2020 Machine learning for the prediction of pseudorealistic pediatric abdominal phantoms for radiation dose reconstruction J. Med. Imaging 7 046501 doi:10.1117/1.JMI.7.4.046501
  6. Zhong J, Hu X, Zhang J and Gu M 2005 Comparison of performance between different selection strategies on simple genetic algorithms Int. Conf. on Computational Intelligence for Modelling, Control and Automation and Int. Conf. on Intelligent Agents, web Technologies and Internet Commerce (CIMCA-IAWTIC’06) vol 2 (IEEE) pp 1115–21
  7. Lazebnik T and Rosenfeld A 2023 Fspl: A meta-learning approach for a filter and embedded feature selection pipeline Int. J. Appl. Math. Comput. Sci. 33 103–115 doi:10.34768/amcs-2023-0009
  8. Shami L and Lazebnik T 2022 Economic aspects of the detection of new strains in a multi-strain epidemiological-mathematical model Chaos Solitons Fractals 165 112823 doi:10.1016/j.chaos.2022.112823
  9. He X, Zhao K and Chu X 2021 Automl: A survey of the state-of-the-art Knowl.-Based Syst. 212 106622 doi:10.1016/j.knosys.2020.106622
  10. Huber M F 2021 A survey on the explainability of supervised machine learning J. Artif. Intell. Res. 70 28 doi:10.1613/jair.1.12228
  11. Marcinkevics R and Vogt J E 2023 Interpretability and explainability: a machine learning zoo mini-tour (arXiv:2012.01805) link
  12. Li T, Zhong J, Liu J, Wu W and Zhang C 2018 Ease. ml: towards multi-tenant resource sharing for machine learning workloads Proc. VLDB Endowment vol 11 pp 607–20
  13. Heaton J 2016 An empirical analysis of feature engineering for predictive modeling SoutheastCon 2016 pp 1–6
  14. Khurana U, Turaga D, Samulowitz H and Parthasrathy S 2016 Cognito: automated feature engineering for supervised learning 2016 IEEE 16th Int. Conf. on Data Mining Workshops (ICDMW) pp 1304–7
  15. Lu X, Ming L, Liu W and Li H-X 2018 Probabilistic regularized extreme learning machine for robust modeling of noise data IEEE Trans. Cybern. 48 2368–77 doi:10.1109/TCYB.2017.2738060
  16. Dalessandro B 2013 Bring the noise: embracing randomness is the key to scaling up machine learning algorithms Big Data 1 110–2 doi:10.1089/big.2013.0010
  17. Gama J, Zliobaite I, Bifet A, Pechenizkiy M and Bouchachia A 2014 A survey on concept drift adaptation ACM Comput. Surv. 46 1–37 doi:10.1145/2523813
  18. Yao H, Wang Y, Li S, Zhang L, Liang W, Zou J and Finn C 2022 Improving out-of-distribution robustness via selective augmentation Proc. 39th Int. Conf. on Machine Learning (Proc. of Machine Learning Research) vol 162 pp 25407–37
  19. Krueger D, Caballero E, Jacobsen J-H, Zhang A, Binas J, Zhang D, Priol R L and Courville A 2021 Out-of-distribution generalization via risk extrapolation (rex) Proc. 38th Int. Conf. on Machine Learning vol 139 (PMLR) pp 5815–26
  20. Fort S, Ren J and Lakshminarayanan B 2021 Exploring the limits of out-of-distribution detection Advances in Neural Information Processing Systems vol 34, ed M Ranzato, A Beygelzimer, Y Dauphin, P S Liang and J W Vaughan pp 7068–81
  21. Hendrycks D et al 2021 The many faces of robustness: a critical analysis of out-of-distribution generalization Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV) pp 8340–9
  22. Miller J P, Taori R, Raghunathan A, Sagawa S, Koh P W, Shankar V, Liang P, Carmon Y and Schmidt L 2021 Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization Proc. 38th Int. Conf. on Machine Learning vol 139 pp 7721–35
  23. Hsu Y-C, Shen Y, Jin H and Kira Z 2020 Generalized odin: detecting out-of-distribution image without learning from out-of-distribution data Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR)
  24. Bengio Y et al 2011 Deep learners benefit more from out-of-distribution examples Proc. 14th Int. Conf. on Artificial Intelligence and Statistics vol 15 (PMLR) pp 164–72
  25. Jordan M I and Mitchell T M 2015 Machine learning: trends, perspectives and prospects Science 349 255–60 doi:10.1126/science.aaa8415
  26. Chen Q and Xue B 2022 Generalisation in Genetic Programming for Symbolic Regression: Challenges and Future Directions (Springer) pp 281–302
  27. Zegklitz J and Posik P 2021 Benchmarking state-of-the-art symbolic regression algorithms Genet. Program. Evol. Mach. 22 5–33 doi:10.1007/s10710-020-09387-0
  28. Keren L S, Liberzon A and Lazebnik T 2023 A computational framework for physics-informed symbolic regression with straightforward integration of domain knowledge Sci. Rep. 13 1249 doi:10.1038/s41598-023-28328-2
  29. Biggio L, Bendinelli T, Neitz A, Lucchi A and Parascandolo G 2021 Neural symbolic regression that scales Proc. 38th Int. Conf. on Machine Learning vol 139 pp 936–45
  30. Wilstrup C and Kasak J 2021 Symbolic regression outperforms other models for small data sets (arXiv:2103.15147) link
  31. Udrescu S-M and Tegmark M 2020 AI Feynman: a physics-inspired method for symbolic regression Sci. Adv. 6 eaay2631 doi:10.1126/sciadv.aay2631
  32. Stijven S, Vladislavleva E, Kordon A, Willem L and Kotanchek M E 2016 Prime-time: symbolic regression takes its place in the real world Genetic Programming Theory and Practice Xiii (Genetic and Evolutionary Computation)
  33. Mahouti P, Gunes F, Belen M A and Demirel S 2021 Symbolic regression for derivation of an accurate analytical formulation using “big data”: an application example Appl. Comput. Electromagn. Soc. J. 32 372–80 doi:10.1016/j.jhydrol.2014.09.060
  34. Brolos K R, Machado M V, Cave C, Kasak J, Stentoft-Hansen V, Batanero V G, Jelen T and Wilstrup C 2021 An approach to symbolic regression using feyn (arXiv:2104.05417) link
  35. Sathia V, Ganesh V and Nanditale S R T 2021 Accelerating genetic programming using GPUs (arXiv:2110.11226) link
  36. Olson R S and Moore J H 2016 Tpot: A tree-based pipeline optimization tool for automating machine learning JMLR: Workshop and Conf. Proc. vol 64 pp 66–74
  37. Jin H, Song Q and Hu X 2019 Auto-keras: An efficient neural architecture search system Proc. 25th ACM SIGKDD Int. Conf. on Knowledge Discovery & Data Mining (Association for Computing Machinery) pp 1946–56
  38. Arjovsky M 2020 Out of distribution generalization in machine learning PhD Thesis New York University
  39. Caro M C, Huang H-Y, Ezzell N, Gibbs J, Sornborger A T, Cincio L, Coles P J and Holmes Z 2023 Out-of-distribution generalization for learning quantum dynamics Nat. Commun. 14 3751 doi:10.1038/s41467-023-39381-w
  40. Veturi Y A et al 2022 Syntheye: investigating the impact of synthetic data on ai-assisted gene diagnosis of inherited retinal disease Ophthalmol. Sci. 3 100258 doi:10.1016/j.xops.2022.100258
  41. Birky D, Garbrecht K, Emery J, Alleman C, Bomarito G and Hochhalter J 2023 Generalizing the gurson model using symbolic regression and transfer learning to relax inherent assumptions Modelling Simul. Mater. Sci. Eng. 31 085005 doi:10.1088/1361-651X/acfe28
  42. Dundar B, Krishnapuram M, Bi J and Rao R B 2007 Learning classifiers when the training data is not IID IJCAI 2007 756–61
  43. Krongauz D and Lazebnik T 2022 Collective evolution learning model for vision-based collective motion with collision avoidance PLoS One 18 e0270318
  44. Afsar M M, Crump T and Far B 2022 Reinforcement learning based recommender systems: a survey ACM Comput. Surv. 55 1–38
  45. Vilalta R, Giraud-Carrier C and Brazdil P 2010 Meta-Learning - Concepts and Techniques (Springer) pp 717–31
  46. Ghassemi N and Fazl-Ersi E 2022 A comprehensive review of trends, applications and challenges in out-of-distribution detection (arXiv:2209.12935) link
  47. Yang J, Zhou K, Li Y and Liu Z 2022 Generalized out-of-distribution detection: a survey Int. J. Comput. Vis. 132 5635–62
  48. Kirchheim K, Filax M and Ortmeier F 2022 Pytorch-ood: a library for out-of-distribution detection based on pytorch Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops pp 4351–60
  49. Omee S S, Fu N, Dong R, Hu M and Hu J 2024 Structure-based out-of-distribution (ood) materials property prediction: a benchmark study npj Comput. Mater. 10 144 doi:10.1038/s41524-024-01316-4
  50. Li K, Rubungo A N, Lei X, Persaud D, Choudhary K, DeCost B, Dieng A B and Hattrick-Simpers J 2024 Probing out-of-distribution generalization in machine learning for materials (arXiv:2406.06489) link
  51. Liu J, Shen Z, He Y, Zhang X, Xu H, Yu R and Cui P 2023 Towards out-of-distribution generalization: a survey (arXiv:2108.13624) link
  52. Krueger D, Caballero E, Jacobsen J-H, Zhang A, Binas J, Zhang D, Priol R L and Courville A 2021 Out-of-distribution generalization via risk extrapolation (rex) Proc. 38th Int. Conf. on Machine Learning ed M Meila and T Zhang (PMLR) pp 5815–26
  53. Yao H, Wang Y, Li S, Zhang L, Liang W, Zou J and Finn C 2022 Improving out-of-distribution robustness via selective augmentation Proc. 39th Int. Conf. on Machine Learning vol 162, ed K Chaudhuri, S Jegelka, L Song, C Szepesvari, G Niu and S Sabato (PMLR) pp 25407–37
  54. Bai H, Canal G, Du X, Kwon J, Nowak R D and Li Y 2023 Feed two birds with one scone: exploiting wild data for both out-of-distribution generalization and detection Proc. 40th Int. Conf. on Machine Learning vol 202, ed A Krause, E Brunskill, K Cho, B Engelhardt, S Sabato and J Scarlett (PMLR) pp 1454–71
  55. Miller J P, Taori R, Raghunathan A, Sagawa S, Koh P W, Shankar V, Liang P, Carmon Y and Schmidt L 2021 Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization Proc. 38th Int. Conf. on Machine Learning (Proc. of Machine Learning Research) vol 139, ed M Meila and T Zhang (PMLR) pp 7721–35
  56. La Cava W, Orzechowski P, Burlacu B, de França F O, Virgolin M, Jin Y, Kommenda M and Moore J H 2021 Contemporary symbolic regression methods and their relative performance (arXiv:2107.14351) link
  57. Wang Y, Wagner N and James M R 2019 Symbolic regression in materials science MRS Commun. 9 793–805 doi:10.1557/mrc.2019.85
  58. Heule M J H and Kullmann O 2017 The science of brute force Commun. ACM 60 70–79 doi:10.1145/3107239
  59. Riolo R 2013 Genetic Programming Theory and Practice X (Springer)
  60. Miller B L et al 1995 Genetic algorithms, tournament selection and the effects of noise Complex Syst. 9 193–212
  61. Orzechowski P, La Cava W and Moore J H 2018 Where are we now? A large benchmark study of recent symbolic regression methods GECCO18: Proc. Genetic and Evolutionary Computation Conf.
  62. Petersen B K, Larma M L, Mundhenk T N, Santiago C P, Kim S K and Kim J T 2019 Deep symbolic regression: recovering mathematical expressions from data via risk-seeking policy gradients (arXiv:1912.04871) link
  63. Quade M, Abel M, Nathanutz J and Brunton S L 2018 Sparse identification of nonlinear dynamics for rapid model recovery Chaos 28 063116 doi:10.1063/1.5027470
  64. Brunton S L, Proctor J L and Kutz J N 2016 Discovering governing equations from data by sparse identification of nonlinear dynamical systems Proc. Natl Acad. Sci. 113 3932–7 doi:10.1073/pnas.1517384113
  65. Kaiser E, Kutz J N and Brunton S L 2018 Sparse identification of nonlinear dynamics for model predictive control in the low-data limit Proc. R. Soc. A 474 20180335 doi:10.1098/rspa.2018.0335
  66. Mangan N M, Kutz J N, Brunton S L and Proctor j L 2017 Model selection for dynamical systems via sparse regression and information criteria Proc. R. Soc. A 473 20170009 doi:10.1098/rspa.2017.0009
  67. Kaptanoglu A A, de Silva B M, Fasel U, Kaheman K, Callaham J L, Delahunt C B, Champion K, Loiseau J-C, Kutz J N and Brunton S L 2021 Pysindy: a comprehensive python package for robust sparse system identification (arXiv:2111.08481) link
  68. Kronberger G, Olivetti de Franca F, Burlacu B, Haider C and Kommenda M 2022 Shape-constrained symbolic regression-improving extrapolation with prior knowledge Evol. Comput. 30 75–98 doi:10.1162/evco_a_00294
  69. Salustowicz R and Schmidhuber J 1997 Probabilistic incremental program evolution Evol. Comput. 5 123–41 doi:10.1162/evco.1997.5.2.123
  70. Sastry K and Goldberg D E 2003 Probabilistic model building and competent genetic programming Genetic Programming Theory and Practice (Springer) pp 205–20
  71. Yanai K and Iba H 2003 Estimation of distribution programming based on Bayesian network The 2003 Congress on Evolutionary Computation vol 3 (IEEE) pp 1618–25
  72. Hemberg E, Veeramachaneni K, McDermott J, Berzan C and O’Reilly U-M 2012 An investigation of local patterns for estimation of distribution genetic programming Proc. 14th Annual Conf. on Genetic and Evolutionary Computation pp 767–74
  73. Shan Y, McKay R I, Baxter R, Abbass H, Essam D and Nguyen H X 2004 Grammar model-based program evolution Proc. 2004 Congress on Evolutionary Computation vol 1 (IEEE) pp 478–85
  74. Bosman P A N and de Jong E D 2004 Learning probabilistic tree grammars for genetic programming Int. Conf. on Parallel Problem Solving From Nature (Springer) pp 192–201
  75. Wong P-K, Lo L-Y, Wong M-L and Leung K-S 2014 Grammar-based genetic programming with Bayesian network 2014 IEEE Congress on Evolutionary Computation (IEEE) pp 739–46
  76. Sotto L F D P and de Melo V V 2017 A probabilistic linear genetic programming with stochastic context-free grammar for solving symbolic regression problems Proc. Genetic and Evolutionary Computation Conf. pp 1017–24
  77. Stephens T 2016 Genetic Programming in Python, with a scikit-learn inspired API: gplearn (available at: https://gplearn.readthedocs. io/en/stable/intro. html) link
  78. Conrad F, Malzer M, Schwarzenberger M, Wiemer H and Ihlenfeldt S 2022 Benchmarking AutoML for regression tasks on small tabular data in materials design Sci. Rep. 12 19350 doi:10.1038/s41598-022-23327-1
  79. Huang J S, Liew J X and Liew K M 2021 Data-driven machine learning approach for exploring and assessing mechanical properties of carbon nanotube-reinforced cement composites Compos. Struct. 267 113917 doi:10.1016/j.compstruct.2021.113917
  80. Su M, Zhong Q, Peng H and Li S 2021 Selected machine learning approaches for predicting the interfacial bond strength between FRPs and concrete Constr. Build. Mater. 270 121456 doi:10.1016/j.conbuildmat.2020.121456
  81. Atici U 2011 Prediction of the strength of mineral admixture concrete using multivariable regression analysis and an artificial neural network Expert Syst. Appl. 38 9609–18 doi:10.1016/j.eswa.2011.01.156
  82. Guo S, Yu J, Liu X, Wang C and Jiang Q 2019 A predicting model for properties of steel using the industrial big data based on machine learning Comput. Mater. Sci. 160 95–104 doi:10.1016/j.commatsci.2018.12.056
  83. Koya B P, Aneja S, Gupta R and Valeo C 2022 Comparative analysis of different machine learning algorithms to predict mechanical properties of concrete Mech. Adv. Mater. Struct. 29 4032–43 doi:10.1080/15376494.2021.1917021
  84. Dunn A, Wang Q, Ganose A, Dopp D and Jain A 2020 Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm Comput. Mater. 6 138 doi:10.1038/s41524-020-00406-3
  85. Bachir R, Mohammed A M S and Habib T 2018 Using artificial neural networks approach to estimate compressive strength for rubberized concrete Period. Polytech. Civil Eng. 62 858–65
  86. Mahalanobis P C 2018 On the generalized distance in statistics Sankhya 80 S1–S7
  87. Bendale A and Boult T 2015 Towards open world recognition Proc. IEEE Conference on Computer Vision and Pattern Recognition pp 1893–902
  88. Lee K, Lee K, Lee H and Shin J 2018 A simple unified framework for detecting out-of-distribution samples and adversarial attacks Advances in Neural Information Processing Systems p 31
  89. Ren J, Fort S, Liu J, Roy A G, Padhy S and Lakshminarayanan B 2021 A simple fix to mahalanobis distance for improving near-ood detection (arXiv:2106.09022) link
  90. Sehwag V, Chiang M and Mittal P 2021 Ssd: A unified framework for self-supervised outlier detection (arXiv:2103.12051) doi:10.48550/arXiv.2103.12051
  91. Taylor P N, Moreira da Silva N, Blamire A, Wang Y and Forsyth R 2020 Early deviation from normal structural connectivity: a novel intrinsic severity score for mild TBI Neurology 94 e1021–6 doi:10.1212/WNL.0000000000008902
  92. Mahony C R and Cannon A J 2018 Wetter summers can intensify departures from natural variability in a warming climate Nat. Commun. 9 783 doi:10.1038/s41467-018-03132-z
  93. Çetin U and Tasgin M 2020 Anomaly detection with multivariate k-sigma score using monte carlo 2020 5th Int. Conf. on Computer Science and Engineering (UBMK) (IEEE) pp 94–98
  94. Kim G 2000 Multivariate outliers and decompositions of Mahalanobis distance Commun. Stat. - Theory Methods 29 1511–26 doi:10.1080/03610920008832559
  95. Mayrhofer M and Filzmoser P 2023 Multivariate outlier explanations using shapley values and Mahalanobis distances Econ. Stat. 6–13 doi:10.1016/j.ecosta.2023.04.003
  96. Sastry C M and Oore S 2020 Detecting out-of-distribution examples with gram matrices Proc. 37th Int. Conf. on Machine Learning (Proc. of Machine Learning Research) vol 119 (PMLR) pp 8491–501
  97. Zhang Y, Pan J, Liu W, Chen Z, Li K, Wang J, Liu Z and Wei H 2023 Kullback-Leibler divergence-based out-of-distribution detection with flow-based generative models IEEE Trans. Knowl. Data Eng. 36 1–14 doi:10.1109/TKDE.2023.3309853
  98. Holland J H 1992 Genetic algorithms Sci. Am. 267 66–73 doi:10.1038/scientificamerican0792-66
  99. Erickson N, Mueller J, Shirkov A, Zhang H, Larroy P, Li M and Smola A 2020 Autogluon-tabular: robust and accurate automl for structured data (arXiv:2003.06505) link
  100. Hollmann N, Müller S, Eggensperger K, and Hutter F 2022 TabPFN: a transformer that solves small tabular classification problems in a second (arXiv:2207.01848) link
  101. Hoo S B, Müller S, Salinas D and Hutter F 2024 The tabular foundation model TabPFN outperforms specialized time series forecasting models based on simple features NeurIPS 2024 Third table Representation Learning Workshop
  102. Chen T and Guestrin C 2016 XGBoost: a scalable tree boosting system Proc. 22nd ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining pp 785–94
  103. Rokach L 2016 Decision forest: twenty years of research Inf. Fusion 27 111–25 doi:10.1016/j.inffus.2015.06.005
  104. Swain P H and Hauska H 1977 The decision tree classifier: design and potential IEEE Trans. Geosci. Electron. 15 142–7 doi:10.1109/TGE.1977.6498972
  105. Pedregosa F et al 2011 Scikit-learn: machine learning in python J. Mach. Learn. Res. 12 2825–30
  106. Zang B, Huang R, Wang L, Chen J, Tian F and Wei X 2016 An improved KNN algorithm based on minority class distribution for imbalanced dataset 2016 Int. Computer Symp. (ICS) pp 696–700
  107. Shami L and Lazebnik T 2023 Implementing machine learning methods in estimating the size of the non-observed economy Comput. Econ. 63 1459–76 doi:10.1007/s10614-023-10369-4
  108. Shmuel A, Glickman O and Lazebnik T 2024 Symbolic regression as a feature engineering method for machine and deep learning regression tasks Mach. Learn.: Sci. Technol. 5 025065 doi:10.1088/2632-2153/ad513a

This page reproduces the article Shmuel et al. (2025), Machine Learning: Science and Technology, doi:10.1088/2632-2153/ada221, with the permission of the publisher. Text, tables and figures were extracted from the PDF and the layout adapted for the web; the PDF is the version of record.

Cite this paper

APA

Shmuel, A., Glickman, O., & Lazebnik, T. (2025). Machine and deep learning performance in out-of-distribution regressions. Machine Learning: Science and Technology, 5, 045078. https://doi.org/10.1088/2632-2153/ada221

BibTeX

@article{shmuel2025machine,
  title = {Machine and deep learning performance in out-of-distribution regressions},
  author = {Shmuel, Assaf and Glickman, Oren and Lazebnik, Teddy},
  journal = {Machine Learning: Science and Technology},
  volume = {5},
  pages = {045078},
  year = {2025},
  doi = {10.1088/2632-2153/ada221}
}