Skip to main content

`treatment_column`

  • Available in: Uplift DRF
  • Hyperparameter: no

Description​

Use this option to specify a treatment column. The column specifies information about group dividing. The groups should be randomly selected before the experiment begins and should have similar sizes.

The data being used must be categorical and have two categories:

  • 0 means the observation is in control group
  • 1 means the observation is in treatment group

Uplift DRF currently supports only one treatment and one control group.

Notes:

  • The treatment column cannot be the same as the response column y.

Example​

library(h2o)
h2o.init()

# Import the uplift dataset into H2O:
data <- h2o.importFile("https://s3.amazonaws.com/h2o-public-test-data/smalldata/uplift/criteo_uplift_13k.csv")

# Set the predictors, response, and treatment column:
# set the predictors
predictors <- c("f1", "f2", "f3", "f4", "f5", "f6","f7", "f8")
# set the response as a factor
data$conversion <- as.factor(data$conversion)
# set the treatment column as a factor
data$treatment <- as.factor(data$treatment)

# Split the dataset into a train and valid set:
data_split <- h2o.splitFrame(data = data, ratios = 0.8, seed = 1234)
train <- data_split[[1]]
valid <- data_split[[2]]

# Build and train the model:
uplift.model <- h2o.upliftRandomForest(training_frame = train,
validation_frame=valid,
x=predictors,
y="conversion",
ntrees=10,
max_depth=5,
treatment_column="treatment",
uplift_metric="KL",
min_rows=10,
nbins=1000,
seed=1234,
auuc_type="qini")
# Eval performance:
perf <- h2o.performance(uplift.model)

# Generate predictions on a validation set (if necessary):
predict <- h2o.predict(uplift.model, newdata = valid)

Feedback