Skip to main content

`link`

  • Available in: GLM, GAM
  • Hyperparameter: no

Description​

GLM and GAM problems consist of three main components:

  • A random component ff for the dependent variable yy: The density function f(y;θ,ϕ)f(y;\theta,\phi) has a probability distribution from the exponential family parametrized by θ\theta and ϕ\phi. This removes the restriction on the distribution of the error and allows for non-homogeneity of the variance with respect to the mean vector.
  • A systematic component (linear model) η\eta: η=Xβ\eta = X\beta, where XX is the matrix of all observation vectors xix_i.
  • A link function gg: E(y)=μ=g−1(η)E(y) = \mu = {g^-1}(\eta) relates the expected value of the response μ\mu to the linear component η\eta. The link function can be any monotonic differentiable function. This relaxes the constraints on the additivity of the covariates, and it allows the response to belong to a restricted range of values depending on the chosen transformation gg.

Accordingly, in order to specify a GLM or GAM problem, you must choose a family function ff, link function gg, and any parameters needed to train the model.

H2O's GLM and GAM support the following link functions: Family_Default, Identity, Logit, Log, Inverse, Tweedie, or Ologit.

The following table describes the allowed Family/Link combinations.

FamilyFamily_DefaultIdentityLogitLogInverseTweedieOlogit
BinomialXX
Fractional BinomialXX
QuasibinomialXX
MultinomialX
OrdinalXX
GaussianXXXX
PoissonXXX
GammaXXXX
TweedieXX
Negative BinomialXXX
AUTOX***X*X**X*X*

For AUTO:

  • X*: the data is numeric (Real or Int) (family determined as gaussian)
  • X**: the data is Enum with cardinality = 2 (family determined as binomial)
  • X***: the data is Enum with cardinality > 2 (family determined as multinomial)

Refer to the Links section for more information.

Example​

library(h2o)
h2o.init()

# import the iris dataset:
# this dataset is used to classify the type of iris plant
# the original dataset can be found at https://archive.ics.uci.edu/ml/datasets/Iris
iris <- h2o.importFile("http://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris_wheader.csv")

# convert response column to a factor
iris['class'] <- as.factor(iris['class'])

# set the predictor names and the response column name
predictors <- colnames(iris)[-length(iris)]
response <- 'class'

# split into train and validation
iris_splits <- h2o.splitFrame(data = iris, ratios = 0.8)
train <- iris_splits[[1]]
valid <- iris_splits[[2]]

# try using the `link` parameter:
iris_glm <- h2o.glm(x = predictors, y = response, family = 'multinomial', link = 'family_default',
training_frame = train, validation_frame = valid)

# print the logloss for the validation data
print(h2o.logloss(iris_glm, valid = TRUE))

Feedback