Skip to main content
Version: v1.2.11

Specify an annotation task rubric

Overview​

After creating a new annotation task, specify an annotation task rubric in the Rubric tab. An annotation task rubric refers to the labels (for example, object classes) to use when annotating a dataset. For example, after creating a new annotation task for an object detection dataset, you have to specify the object classes to use when labeling the dataset in the annotation task rubric.

Instructions​

An annotation task rubric differs based on the specified task type of the dataset used to create the annotation task.

Text annotation tasks​

  • Instructions: Specify one or more categorical target labels for a text classification task rubric.
  • Example: To specify happy and unhappy as labels, one can consider the following instructions in the Rubric tab of the annotation task:
note

To learn how to access the Rubric tab of an annotation task (or other tabs), see Access an annotation task's tabs.

  1. In the New class name box, enter happy.
  2. Click Add.
  3. Click Add class.
  4. In the New class name box, enter unhappy.
  5. Click Add.

Text classification annotation task

Image annotation tasks​

For an image classification task rubric, you need specify one or more categorical target labels in the annotation task rubric for an image classification annotation task. To learn more, see Tutorial 1B: Annotation task: Image classification

  • Instructions: Specify one or more categorical target labels.

  • Example: To create a car and coffee label, one can consider the following instructions in the Rubric tab of the annotation task:

    note

    To learn how to access the Rubric tab of an annotation task (or other tabs), see Access an annotation task's tabs.

    1. In the New class name box, enter car.
    2. Click Add.
    3. Click Add class.
    4. In the New class name box, enter coffee.
    5. Click Add.

Image classification rubric

Audio annotation tasks​

  • Instructions: Specify one or more categorical target labels.
  • Example: To specify chainsaw and clock_tick as labels, one can consider the following instructions in the Rubric tab of the annotation task:
note

To learn how to access the Rubric tab of an annotation task (or other tabs), see Access an annotation task's tabs.

  1. In the New class name box, enter chainsaw.
  2. Click Add.
  3. Click Add class.
  4. In the New class name box, enter clock_tick.
  5. Click Add.

Audio classification annotation task rubric

Large language model (LLM) parameters​

Select API endpoint type​

This LLM parameter defines the zero-shot learning model family (LLM) to utilize for the text-generative AI annotation task.

Options:

  • h2oGPT
    • This option enables h2oGPTe LLMs to be available for a text-generative AI annotation task.
  • OpenAI
    • This option enables OpenAI LLMs in your OpenAI account (API key) to be available for a text-generative AI annotation task. To connect to your OpenAI LLMs, see OpenAI API settings

LLM model name​

This LLM parameter defines the zero-shot learning model name (LLM) to utilize for a text-generative AI annotation task.

Max response tokens​

This LLM parameter defines the maximum number of tokens for a response; a low number can result in short responses, which might limit the responses.

Temperature​

This LLM parameter defines the randomness of predictions by scaling the logits. Higher temperature values increase creativity on the part of the model while producing more diverse outputs. In other words, the temperature makes the distribution more random.

Repetition penalty​

This LLM parameter defines the penalty value of tokens frequently reappearing in the text (response). For example, a token that has already appeared ten times can be penalized more than a token that has appeared only two times. A 1.0 value means no penalty.

tip

This setting can be helpful when attempting to reduce the model's tendency to generate verbatim/identical text.

Prompt template​

Select example prompt​

This setting defines the input format for the selected model. There are several options, including the option to create your own custom input format (custom).

Options:

  • summarize
    • Asks the LLM to write a brief summary of the text column.
  • translate
    • Asks the LLM to translate the text column into German.
  • product_review
    • Asks the LLM to rate the product review in the text column and return a satisfaction score, the product name, and the main issue with the product.
  • evaluate
    • Asks the LLM to rate how good an existing summary is and to suggest improvements. Unlike the other example prompts, evaluate refers to two columns: the annotation task's text column and an extra column named summary. To learn more, see Prompt placeholders.
  • custom
    • Inserts a placeholder for the text column only, so that you can write your own prompt.

Prompt placeholders​

A prompt template refers to a dataset column by placing the column name in curly braces (for example, {text}). When you select an example prompt, H2O Label Genie inserts the name of the annotation task's text column into the prompt for you.

The prompt must contain a placeholder for every column listed in the Selected columns section of the Rubric tab. The first list in that section shows the text column of the annotation task. The second list, which is only available when the dataset contains columns other than the text column, lets you select extra columns to include in the prompt.

H2O Label Genie does not correct the prompt for you. Fix it before annotating, or the LLM request will fail and the annotation task will report a failed background task.

caution

The evaluate example prompt contains a {summary} placeholder in addition to the text column placeholder, and no H2O Label Genie dataset contains a summary column by default. As a result, selecting evaluate displays an Invalid prompt error until you do one of the following:

  • In the Selected columns section, use the second list to select an extra column named summary.
  • Click the edit icon (pencil button) next to the prompt text box, and then replace {summary} with a placeholder for a column that your dataset contains. Make sure that the column is also selected in the Selected columns section.

Feedback