Skip to main content

Split datasets into training/testing/validating

This example shows how to split a single dataset into two datasets, one used for training and the other used for testing.

note

When splitting frames, H2O-3 Secure does not give an exact split. It's designed to be efficient on big data using a probabilistic splitting method rather than an exact split.

For example, when specifying a 0.75/0.25 split, H2O-3 Secure will produce a test/train split with an expected value of 0.75/0.25 rather than exactly 0.75/0.25. On small datasets, the sizes of the resulting splits will deviate from the expected value more than on big data, where they are close to exact.

import h2o
from h2o.estimators.glm import H2OGeneralizedLinearEstimator
h2o.init()

# Import the prostate dataset
prostate = "http://h2o-public-test-data.s3.amazonaws.com/smalldata/prostate/prostate.csv"
prostate_df = h2o.import_file(path=prostate)

# Split the data into Train/Test/Validation with Train having 70% and test and validation 15% each
train,test,valid = prostate_df.split_frame(ratios=[.7, .15])

# Generate a GLM model using the training dataset
glm_classifier = H2OGeneralizedLinearEstimator(family="binomial", nfolds=10, alpha=0.5)
glm_classifier.train(y="CAPSULE", x=["AGE", "RACE", "PSA", "DCAPS"], training_frame=train)

# Predict using the GLM model and the testing dataset
predict = glm_classifier.predict(test)

# View a summary of the prediction
predict.head()
predict p0 p1
--------- -------- --------
1 0.366189 0.633811
1 0.351269 0.648731
1 0.69012 0.30988
0 0.762335 0.237665
1 0.680127 0.319873
1 0.687736 0.312264
1 0.676753 0.323247
1 0.685876 0.314124
1 0.707027 0.292973
0 0.74706 0.25294

[10 rows x 3 columns]

Feedback