Using H2O-3 Secure on Spark (Sparkling Water)
Running on Spark through Sparkling Water is part of H2O-3 Secure: the supported, production-grade path for production deployments. Contact enterprise@h2o.ai.
Sparkling Water runs H2O-3 Secure machine learning algorithms on an Apache Spark cluster, so you can prepare data with Spark and train models with H2O-3 Secure. It converts data between Spark DataFrames, RDDs, and H2OFrames.
Sparkling Water provides a separate build for each supported Spark version, and builds are not interchangeable.
Requirements
Before you start, make sure you have the following:
- A Spark installation on your cluster.
- A Spark version that Sparkling Water supports. For the current list, see the Sparkling Water documentation.
Download Sparkling Water
To get the build that matches your version of Spark, contact enterprise@h2o.ai for the download link.
Start Sparkling Water
After you download the build, start Sparkling Water and an H2O-3 Secure cluster inside Spark:
-
Unzip the archive, set
SPARK_HOME, then start the Sparkling Water shell:unzip sparkling-water-<version>.zipcd sparkling-water-<version>export SPARK_HOME=<path-to-spark>bin/sparkling-shell -
In the shell, start the H2O-3 Secure cluster:
import ai.h2o.sparkling._val h2oContext = H2OContext.getOrCreate()
For configuration and deployment options, see the Sparkling Water documentation.
Language interfaces
Sparkling Water provides Python and R interfaces in addition to Scala:
- PySparkling is the Python interface. Contact enterprise@h2o.ai for the download link.
- RSparkling is the R interface. It extends sparklyr. Use sparklyr to deploy and initialize Sparkling Water, then use the H2O-3 Secure R package to build models.
Additional resources
Next steps
To run H2O-3 Secure without Spark, see Downloading and installing and Starting H2O-3 Secure.
- Submit and view feedback for this page
- Send feedback about H2O-3 Secure to cloud-feedback@h2o.ai