Skip to main content

Telemetry

Starting with version 3.46.0.12, H2O-3 Secure can send anonymous usage telemetry to help H2O.ai prioritize features and platforms. Telemetry is opt-in and off by default: it sends nothing unless you turn it on. When enabled, it is unobtrusive by design: every send is fire-and-forget with a short timeout, so if the receiver is unreachable your code behaves exactly as if telemetry never ran. It never blocks, raises, or retries.

Telemetry covers only operational and runtime characteristics, and H2O.ai uses it only in aggregate to operate, secure, support, maintain, and improve H2O-3 Secure: to understand common deployment patterns, prioritize compatibility and support, and detect abuse or malicious activity. H2O.ai does not use it for sales prospecting, lead identification, or user surveillance.

What is sent (when enabled)​

  • One small ping when you start or connect to H2O-3 Secure (h2o.init() / h2o.connect() in Python or R, or a standalone java -jar h2o.jar / hadoop jar h2odriver.jar cluster), plus one per major action: training, scoring, MOJO and model download, upload, import, parse, AutoML, and model save/load.
  • Each ping contains the H2O-3 Secure version, the client (python / r / jvm), the operating system, an ephemeral session ID regenerated on every start, a timestamp, the algorithm name, and coarse range buckets for counts such as rows, columns, durations, and sizes. Telemetry reports these counts as ranges rather than exact figures. One exception: it sends a small cluster's node count (1–16) exactly, because 1-node versus 4-node is operationally meaningful, and buckets larger clusters.

The per-action pings come from the Python client. The R client sends a single event when a session starts, and a standalone or Hadoop JVM cluster sends one event from its leader node when the cluster forms.

What is never sent​

Code, dataset contents, prompts, model inputs or outputs, training data, prediction values, file paths or URLs, dataset or model names, column names, parameter values, hostnames, usernames, email addresses, precise location, or any other customer business data.

Source IP addresses are inherently visible to any HTTPS request. The receiver may use them transiently to derive a coarse geographic region and for network attribution, and does not persist or store them. The telemetry payload itself contains no location data.

Enable telemetry​

Telemetry is off by default. Turn it on with any of the following.

Python and R clients​

  • Per session: pass telemetry=True in Python or telemetry = TRUE in R to h2o.init() or h2o.connect().

  • Persistent: use the setter h2o.set_telemetry() to change it, and the getter h2o.telemetry_enabled() to read the current state. The setting applies immediately and persists under ~/.h2oai/telemetry, so later sessions honor it.

    Python:

    h2o.set_telemetry(True) # opt in (persisted across sessions)
    h2o.set_telemetry(False) # opt back out
    h2o.telemetry_enabled() # -> True or False

    R:

    h2o.set_telemetry(TRUE) # opt in (persisted across sessions)
    h2o.set_telemetry(FALSE) # opt back out
    h2o.telemetry_enabled() # -> TRUE or FALSE
  • Config file: add a general.telemetry key under [general] in ~/.h2oconfig:

    [general]
    telemetry = true

Standalone or Hadoop cluster (JVM)​

A cluster started directly on the JVM (java -jar h2o.jar / hadoop jar h2odriver.jar) is also off by default. Enable it by setting the disable flag to false:

java -Dsys.ai.h2o.telemetry.disabled=false -jar h2o.jar

Turn telemetry off or keep it off​

Since telemetry is off by default, you normally don't need to do anything. To turn it off after enabling it, or to force it off regardless of the preceding settings, use any of the following.

  • Clients: pass telemetry=False in Python or telemetry = FALSE in R for the current session, use h2o.set_telemetry(False) to persist it, or set general.telemetry = false in ~/.h2oconfig.
  • JVM cluster: set -Dsys.ai.h2o.telemetry.disabled=true (the default, off, also applies when the flag is unset).
  • Any environment: set DO_NOT_TRACK=1, the cross-tool standard from consoledonottrack.com. This is a hard opt-out: it always wins over every enable setting, and the Python client, the R client, and the JVM server all honor it.

The receiver also honors the standard DNT: 1 (Do Not Track) and Sec-GPC: 1 (Global Privacy Control) request headers. It drops any event that arrives with either header set, and never stores it.


Feedback