Algorithms
H2O-3 Secure provides the following supervised and unsupervised machine learning algorithms.
Supervised learning
| Algorithm | Description |
|---|---|
| AdaBoost | Combines many weak learners into a single classifier by increasing the weight of misclassified points on each iteration. Solves binary classification problems only. |
| ANOVA GLM | Calculates Type III sum of squares to show how much each predictor or interaction contributes to a model. |
| Cox Proportional Hazards (CoxPH) | Models time-to-event data using a hazard function built from a baseline hazard and a risk score. |
| Decision tree | Builds a single tree of tests on numeric features to classify or predict a binary target. |
| Deep Learning (Neural Networks) | Trains a multi-layer feedforward neural network with stochastic gradient descent and back-propagation. |
| Distributed Random Forest (DRF) | Builds a forest of classification or regression trees on row and column subsets and averages their predictions. |
| Generalized additive models (GAM) | Extends a generalized linear model with smooth functions of predictor variables. |
| Gradient Boosting Machine (GBM) | Builds regression trees sequentially, with each tree improving on the approximation of the previous one. |
| Generalized linear model (GLM) | Estimates regression models for outcomes that follow exponential distributions, including Gaussian, Poisson, binomial, and gamma. |
| Hierarchical Generalized Linear Model (HGLM) | Extends GLM with cluster-level effects for data collected in groups, such as students within schools. |
| Isotonic Regression | Fits a non-decreasing free-form line to an ordered sequence of observations for single-variable regression. |
| ModelSelection | Selects the best predictor subset for a GLM model using one of four search modes. |
| Naïve Bayes classifier | Classifies data by applying Bayes' theorem under an assumption of independence between predictors. |
| RuleFit | Fits a tree ensemble, builds a rule set from it, then fits a sparse linear model to the resulting rule and feature set. |
| Stacked ensembles | Finds the optimal combination of a collection of prediction algorithms by training a second-level model on their predictions. |
| Support Vector Machine (SVM) | Builds a model that assigns new examples to one of two categories. Solves binary classification problems only. |
| GAM thin plate regression spline | Extends GAM smoothers to work with one or more predictors using thin plate regression splines. |
| Distributed uplift random forest (Uplift DRF) | Trains uplift trees to model the incremental impact of a treatment using treatment and control group assignment. Supports binomial classification only. |
| XGBoost | Builds decision trees sequentially, with each new tree correcting the deficiencies of the previous model. |
Unsupervised learning
| Algorithm | Description |
|---|---|
| Aggregator | Reduces a numerical or categorical dataset to fewer rows by clustering dense regions into exemplars, while keeping outliers as outliers. |
| Extended Isolation Forest | Generalizes Isolation Forest by randomizing the branch-cut slope, which reduces the bias introduced by axis-aligned splits. |
| Generalized low rank models (GLRM) | Reduces the dimensionality of a dataset by decomposing it into two smaller numeric matrices. |
| Isolation Forest | Detects anomalies by isolating observations through random feature and split-value selection, which produces shorter paths for outliers. |
| K-Means clustering | Partitions observations into groups so that points within a group resemble each other more than they resemble points in other groups. |
| Principal Component Analysis (PCA) | Transforms a set of possibly collinear features into a new set of uncorrelated features. |
| TF-IDF | Measures how important a word is to a document within a collection of documents. |
| Word2vec | Learns vector representations of words from a text corpus. |
Related topics
- Supported data types: Which data types each algorithm accepts, and how to handle timestamp columns.
- Early stopping: Stop a model build or grid search once it meets a stopping condition.
- Permutation variable importance: Measure how much a feature affects prediction error by permuting it and scoring the model again.
- Quantiles: Retrieve and display quantiles for parsed data.
- Target encoding: Replace a categorical value with the mean of the target variable.
Feedback
- Submit and view feedback for this page
- Send feedback about H2O-3 Secure to cloud-feedback@h2o.ai