Skip to main content

Using H2O-3 Secure on Spark (Sparkling Water)

H2O-3 Secure

Running on Spark through Sparkling Water is part of H2O-3 Secure: the supported, production-grade path for production deployments. Contact enterprise@h2o.ai.

Sparkling Water runs H2O-3 Secure machine learning algorithms on an Apache Spark cluster, so you can prepare data with Spark and train models with H2O-3 Secure. It converts data between Spark DataFrames, RDDs, and H2OFrames.

Sparkling Water provides a separate build for each supported Spark version, and builds are not interchangeable.

Requirements​

Before you start, make sure you have the following:

  • A Spark installation on your cluster.
  • A Spark version that Sparkling Water supports. For the current list, see the Sparkling Water documentation.

Download Sparkling Water​

To get the build that matches your version of Spark, contact enterprise@h2o.ai for the download link.

Start Sparkling Water​

After you download the build, start Sparkling Water and an H2O-3 Secure cluster inside Spark:

  1. Unzip the archive, set SPARK_HOME, then start the Sparkling Water shell:

    unzip sparkling-water-<version>.zip
    cd sparkling-water-<version>
    export SPARK_HOME=<path-to-spark>
    bin/sparkling-shell
  2. In the shell, start the H2O-3 Secure cluster:

    import ai.h2o.sparkling._
    val h2oContext = H2OContext.getOrCreate()

For configuration and deployment options, see the Sparkling Water documentation.

Language interfaces​

Sparkling Water provides Python and R interfaces in addition to Scala:

  • PySparkling is the Python interface. Contact enterprise@h2o.ai for the download link.
  • RSparkling is the R interface. It extends sparklyr. Use sparklyr to deploy and initialize Sparkling Water, then use the H2O-3 Secure R package to build models.

Additional resources​

Next steps​

To run H2O-3 Secure without Spark, see Downloading and installing and Starting H2O-3 Secure.


Feedback