Skip to main content
Version: v1.7.5

Agent tool server

The agent tool server is an OpenAI-compatible endpoint that runs the code you send it and returns the output. This path turns the LLM off, so nothing plans, reasons about, or rewrites your request. The server takes the code block out of your message, runs it in the agent execution environment, and returns what the code produced.

Use the agent tool server when you want the agent container's execution environment without the agent itself. Benchmark harnesses, evaluation suites, and services that generate their own code all need somewhere to run it.

info

The agent containers serve the agent tool server as a deployment-level API. The Enterprise h2oGPTe web interface does not expose it, and no Enterprise h2oGPTe component calls it. Requests go straight to the agent container, so they bypass workspace permissions, per-user quotas, chat history, and the audit trail. This page covers /v1/chat/completions.

caution

Treat the agent tool server as a privileged internal endpoint. Any caller with a valid key runs arbitrary code inside the agent container, with that container's filesystem and environment. Restrict access to trusted services, and keep the tool port inside the deployment network.

Agent chats compared to the agent tool server​

Both paths use the same /v1/chat/completions endpoint and the same execution environment. What changes is how much of the agent runs.

BehaviorAgent chat (use_agent)Agent tool server (use_agent_tool)
LLM turnsThe model plans, writes code, reads the output, and iteratesNone. The request turns the LLM off
Tool selectionThe agent picks tools and decides when to run codeNone. Your code is the request
InputA question or task written in natural languageA fenced code block
OutputA final answer composed by the modelThe raw output of your code

Where the agent tool server runs​

Enterprise h2oGPTe runs the agent tool server alongside the agent server on both agent containers. Each server is its own process, so a request that sets use_agent_tool must go to the tool port, and a request that sets use_agent must go to the agent port.

ServiceAgent server portAgent tool server port
h2ogpt-agent-shared50045005
h2ogpt-agent-isolated50065007

The shared container runs requests from every caller side by side in one execution environment. In the isolated container, each server process accepts one request at a time. The agent server and agent tool server can still accept requests concurrently, and they share the container's filesystem and processes. The isolated service name does not guarantee isolation between callers of these two servers.

caution

When raw-code execution requires isolation, use a dedicated container and prevent concurrent ordinary-agent traffic to it. Do not rely on the per-server request limit as a boundary between mutually untrusted callers.

Each container enables the server through docker-compose.h2ogpt.yml:

H2OGPT_AGENT_TOOL_SERVER: "true"
H2OGPT_AGENT_TOOL_SERVER_PORT: "5005" # 5007 on the isolated container

The agent containers do not publish the tool ports to the host, so call the server from inside the deployment network with the service name, as in http://h2ogpt-agent-shared:5005/v1. On the host, port 5005 belongs to the function OpenAI server, and a request sent there fails with AgentTool is not enabled on this server.

The ports in this table are the Docker Compose defaults. On Kubernetes, the address is the agent Service name with the port the h2ogpt chart sets. The monolithic deployment does not run the agent tool server at all, because its function OpenAI server already binds 5005. To enable the server there, set a distinct H2OGPT_AGENT_TOOL_SERVER_PORT.

Each server in the isolated container handles one request at a time. It rejects a second concurrent request to that server with 503 and a server_busy error rather than queueing it, so retry the request. This limit is not shared between the agent and tool servers. The shared container sets no concurrency limit. A deployment can also run more than one isolated container behind a single service name, so consecutive requests are not guaranteed to reach the same container.

To confirm the server is up, send an unauthenticated GET /health to the tool port. The endpoint returns 200 when the server is ready to accept requests.

Authenticate​

The agent tool server sits behind the same key check as the agent server on the same container. Send a key from H2OGPT_H2OGPT_API_KEYS as a bearer token.

caution

The agent container does not accept the Enterprise h2oGPTe API key you create in the web interface. A key from the APIs page authenticates you to Enterprise h2oGPTe, not to the agent container.

Run code​

Send a chat completion request to the agent tool server with use_agent_tool set to true. The message content is the code you want to run, written as a fenced code block.

from openai import OpenAI

client = OpenAI(
base_url="http://h2ogpt-agent-shared:5005/v1",
api_key="<API_KEY>",
)

response = client.chat.completions.create(
model="<MODEL>",
messages=[
{
"role": "user",
"content": "```python\nprint(6 * 7)\n```",
}
],
extra_body={"use_agent_tool": True},
)

print(response.choices[0].message.content)
print(response.usage.summary)

The OpenAI schema requires model, so send a model name the deployment serves. Call GET /v1/models on the same base URL to list the valid names. The server runs no LLM turn on this path, so the name you pick does not change the output.

When you stream, the code output arrives as content deltas, and the server also sends one or more chunks that carry a usage object instead of content. The execution summary arrives on one of these chunks. Keep the last summary you see rather than assuming the first usage chunk has it. You do not need to set stream_options; the server includes usage on these chunks automatically.

Each code block gets 120 seconds to run, and a request gets 3000 seconds in total. Pass autogen_timeout or autogen_total_timeout to raise either limit for a long-running job.

Write the code block​

The server writes each fenced code block in the last user message to a file in the request's working directory, then runs the blocks in the order they appear.

  • Tag the fence with a language. Use python for Python and sh (or bash) for shell code. The file extension in a # filename: header takes priority over the fence tag when the two disagree.
  • Omit the execution directive. The server adds # execution: true to any fenced block that has no # execution: line, so your code runs as sent. General code tools covers the same directives for agent chats, where you write them yourself and can also target a custom tool environment with # tool:.
  • Name the file with # filename:. Give a block a name when a later block in the same request imports it or runs it.
  • Skip execution with # execution: false. The server writes the file and returns without running it, which is useful for a helper module that a later block in the same request imports.

Write the helper and the block that uses it in the same message. A later request is not guaranteed to see the files an earlier one wrote.

The following block writes report.py and runs it:

# filename: report.py
import platform

print(platform.python_version())

Read the response​

The response follows the OpenAI chat completion shape. Three fields carry the result:

  • choices[0].message.content holds the execution transcript, including the code output.
  • usage.summary holds the execution result, which is the exit code and the captured output. A successful run starts with exitcode: 0 (execution succeeded).
  • When you stream, usage.file_ids_names lists any files the code produced, as a list of single-entry objects that map a file ID to its filename, for example {"file-8f2a1c...": "report.py"}. Take the key of each entry as the file ID, then download it from the same base URL with GET /v1/files/{file_id}/content, or client.files.content(file_id) in the Python SDK.

summary and file_ids_names are extra fields on the usage object. They resolve at runtime, but a type checker flags them, so read them through usage.model_extra["summary"] and usage.model_extra["file_ids_names"] in typed code.

Code that exits non-zero is not an API error. The request returns 200, and the failure shows up in usage.summary as a non-zero exit code with the error output from the code.

Troubleshooting​

SymptomStatusCauseFix
AgentTool is not enabled on this server.400The request reached a server that is not the agent tool server.Send the request to the tool port on the deployment network: 5005 on the shared container, 5007 on the isolated container.
Agent is not enabled on this server.400The request set use_agent instead of use_agent_tool.Set use_agent_tool to true, or send agent requests to the agent server port.
Server is currently processing the maximum number of concurrent requests503A second request reached the same server in the isolated container while its first request was still running.Retry after that request finishes. The agent and tool servers have separate limits.
The response contains no code output.200The message had no fenced code block, or the only block carried # execution: false.Wrap the code in a fence with a language tag, and remove # execution: false if you want the code to run.

Next steps​

  • Learn about General code tools for extending agent chats with your own Python code
  • Review the Agents overview for the full agent path this server turns off
  • See APIs to create the Enterprise h2oGPTe API key used everywhere else in the product

Feedback