Skip to main content

Algorithms

H2O-3 Secure provides the following supervised and unsupervised machine learning algorithms.

Supervised learning​

AlgorithmDescription
AdaBoostCombines many weak learners into a single classifier by increasing the weight of misclassified points on each iteration. Solves binary classification problems only.
ANOVA GLMCalculates Type III sum of squares to show how much each predictor or interaction contributes to a model.
Cox Proportional Hazards (CoxPH)Models time-to-event data using a hazard function built from a baseline hazard and a risk score.
Decision treeBuilds a single tree of tests on numeric features to classify or predict a binary target.
Deep Learning (Neural Networks)Trains a multi-layer feedforward neural network with stochastic gradient descent and back-propagation.
Distributed Random Forest (DRF)Builds a forest of classification or regression trees on row and column subsets and averages their predictions.
Generalized additive models (GAM)Extends a generalized linear model with smooth functions of predictor variables.
Gradient Boosting Machine (GBM)Builds regression trees sequentially, with each tree improving on the approximation of the previous one.
Generalized linear model (GLM)Estimates regression models for outcomes that follow exponential distributions, including Gaussian, Poisson, binomial, and gamma.
Hierarchical Generalized Linear Model (HGLM)Extends GLM with cluster-level effects for data collected in groups, such as students within schools.
Isotonic RegressionFits a non-decreasing free-form line to an ordered sequence of observations for single-variable regression.
ModelSelectionSelects the best predictor subset for a GLM model using one of four search modes.
Naïve Bayes classifierClassifies data by applying Bayes' theorem under an assumption of independence between predictors.
RuleFitFits a tree ensemble, builds a rule set from it, then fits a sparse linear model to the resulting rule and feature set.
Stacked ensemblesFinds the optimal combination of a collection of prediction algorithms by training a second-level model on their predictions.
Support Vector Machine (SVM)Builds a model that assigns new examples to one of two categories. Solves binary classification problems only.
GAM thin plate regression splineExtends GAM smoothers to work with one or more predictors using thin plate regression splines.
Distributed uplift random forest (Uplift DRF)Trains uplift trees to model the incremental impact of a treatment using treatment and control group assignment. Supports binomial classification only.
XGBoostBuilds decision trees sequentially, with each new tree correcting the deficiencies of the previous model.

Unsupervised learning​

AlgorithmDescription
AggregatorReduces a numerical or categorical dataset to fewer rows by clustering dense regions into exemplars, while keeping outliers as outliers.
Extended Isolation ForestGeneralizes Isolation Forest by randomizing the branch-cut slope, which reduces the bias introduced by axis-aligned splits.
Generalized low rank models (GLRM)Reduces the dimensionality of a dataset by decomposing it into two smaller numeric matrices.
Isolation ForestDetects anomalies by isolating observations through random feature and split-value selection, which produces shorter paths for outliers.
K-Means clusteringPartitions observations into groups so that points within a group resemble each other more than they resemble points in other groups.
Principal Component Analysis (PCA)Transforms a set of possibly collinear features into a new set of uncorrelated features.
TF-IDFMeasures how important a word is to a document within a collection of documents.
Word2vecLearns vector representations of words from a text corpus.
  • Supported data types: Which data types each algorithm accepts, and how to handle timestamp columns.
  • Early stopping: Stop a model build or grid search once it meets a stopping condition.
  • Permutation variable importance: Measure how much a feature affects prediction error by permuting it and scoring the model again.
  • Quantiles: Retrieve and display quantiles for parsed data.
  • Target encoding: Replace a categorical value with the mean of the target variable.

Feedback