Agent tool server
The agent tool server is an OpenAI-compatible endpoint that runs the code you send it and returns the output. This path turns the LLM off, so nothing plans, reasons about, or rewrites your request. The server takes the code block out of your message, runs it in the agent execution environment, and returns what the code produced.
Use the agent tool server when you want the agent container's execution environment without the agent itself. Benchmark harnesses, evaluation suites, and services that generate their own code all need somewhere to run it.
The agent containers serve the agent tool server as a deployment-level API. The Enterprise h2oGPTe web interface does not expose it, and no Enterprise h2oGPTe component calls it. Requests go straight to the agent container, so they bypass workspace permissions, per-user quotas, chat history, and the audit trail. This page covers /v1/chat/completions.
Treat the agent tool server as a privileged internal endpoint. Any caller with a valid key runs arbitrary code inside the agent container, with that container's filesystem and environment. Restrict access to trusted services, and keep the tool port inside the deployment network.
Agent chats compared to the agent tool server​
Both paths use the same /v1/chat/completions endpoint and the same execution environment. What changes is how much of the agent runs.
| Behavior | Agent chat (use_agent) | Agent tool server (use_agent_tool) |
|---|---|---|
| LLM turns | The model plans, writes code, reads the output, and iterates | None. The request turns the LLM off |
| Tool selection | The agent picks tools and decides when to run code | None. Your code is the request |
| Input | A question or task written in natural language | A fenced code block |
| Output | A final answer composed by the model | The raw output of your code |
Where the agent tool server runs​
Enterprise h2oGPTe runs the agent tool server alongside the agent server on both agent containers. Each server is its own process, so a request that sets use_agent_tool must go to the tool port, and a request that sets use_agent must go to the agent port.
| Service | Agent server port | Agent tool server port |
|---|---|---|
h2ogpt-agent-shared | 5004 | 5005 |
h2ogpt-agent-isolated | 5006 | 5007 |
The shared container runs requests from every caller side by side in one execution environment. In the isolated container, each server process accepts one request at a time. The agent server and agent tool server can still accept requests concurrently, and they share the container's filesystem and processes. The isolated service name does not guarantee isolation between callers of these two servers.
When raw-code execution requires isolation, use a dedicated container and prevent concurrent ordinary-agent traffic to it. Do not rely on the per-server request limit as a boundary between mutually untrusted callers.
Each container enables the server through docker-compose.h2ogpt.yml:
H2OGPT_AGENT_TOOL_SERVER: "true"
H2OGPT_AGENT_TOOL_SERVER_PORT: "5005" # 5007 on the isolated container
The agent containers do not publish the tool ports to the host, so call the server from inside the deployment network with the service name, as in http://h2ogpt-agent-shared:5005/v1. On the host, port 5005 belongs to the function OpenAI server, and a request sent there fails with AgentTool is not enabled on this server.
The ports in this table are the Docker Compose defaults. On Kubernetes, the address is the agent Service name with the port the h2ogpt chart sets. The monolithic deployment does not run the agent tool server at all, because its function OpenAI server already binds 5005. To enable the server there, set a distinct H2OGPT_AGENT_TOOL_SERVER_PORT.
Each server in the isolated container handles one request at a time. It rejects a second concurrent request to that server with 503 and a server_busy error rather than queueing it, so retry the request. This limit is not shared between the agent and tool servers. The shared container sets no concurrency limit. A deployment can also run more than one isolated container behind a single service name, so consecutive requests are not guaranteed to reach the same container.
To confirm the server is up, send an unauthenticated GET /health to the tool port. The endpoint returns 200 when the server is ready to accept requests.
Authenticate​
The agent tool server sits behind the same key check as the agent server on the same container. Send a key from H2OGPT_H2OGPT_API_KEYS as a bearer token.
The agent container does not accept the Enterprise h2oGPTe API key you create in the web interface. A key from the APIs page authenticates you to Enterprise h2oGPTe, not to the agent container.
Run code​
Send a chat completion request to the agent tool server with use_agent_tool set to true. The message content is the code you want to run, written as a fenced code block.
- Python
- Python (streaming)
- curl
from openai import OpenAI
client = OpenAI(
base_url="http://h2ogpt-agent-shared:5005/v1",
api_key="<API_KEY>",
)
response = client.chat.completions.create(
model="<MODEL>",
messages=[
{
"role": "user",
"content": "```python\nprint(6 * 7)\n```",
}
],
extra_body={"use_agent_tool": True},
)
print(response.choices[0].message.content)
print(response.usage.summary)
from openai import OpenAI
client = OpenAI(
base_url="http://h2ogpt-agent-shared:5005/v1",
api_key="<API_KEY>",
)
stream = client.chat.completions.create(
model="<MODEL>",
messages=[
{
"role": "user",
"content": "```python\nprint(6 * 7)\n```",
}
],
stream=True,
extra_body={"use_agent_tool": True},
)
summary = None
for chunk in stream:
if chunk.usage is not None:
summary = getattr(chunk.usage, "summary", None) or summary
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
if summary:
print(f"\n{summary}")
curl http://h2ogpt-agent-shared:5005/v1/chat/completions \
-H "Authorization: Bearer <API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "<MODEL>",
"messages": [
{"role": "user", "content": "```python\nprint(6 * 7)\n```"}
],
"use_agent_tool": true
}'
The OpenAI schema requires model, so send a model name the deployment serves. Call GET /v1/models on the same base URL to list the valid names. The server runs no LLM turn on this path, so the name you pick does not change the output.
When you stream, the code output arrives as content deltas, and the server also sends one or more chunks that carry a usage object instead of content. The execution summary arrives on one of these chunks. Keep the last summary you see rather than assuming the first usage chunk has it. You do not need to set stream_options; the server includes usage on these chunks automatically.
Each code block gets 120 seconds to run, and a request gets 3000 seconds in total. Pass autogen_timeout or autogen_total_timeout to raise either limit for a long-running job.
Write the code block​
The server writes each fenced code block in the last user message to a file in the request's working directory, then runs the blocks in the order they appear.
- Tag the fence with a language. Use
pythonfor Python andsh(orbash) for shell code. The file extension in a# filename:header takes priority over the fence tag when the two disagree. - Omit the execution directive. The server adds
# execution: trueto any fenced block that has no# execution:line, so your code runs as sent. General code tools covers the same directives for agent chats, where you write them yourself and can also target a custom tool environment with# tool:. - Name the file with
# filename:. Give a block a name when a later block in the same request imports it or runs it. - Skip execution with
# execution: false. The server writes the file and returns without running it, which is useful for a helper module that a later block in the same request imports.
Write the helper and the block that uses it in the same message. A later request is not guaranteed to see the files an earlier one wrote.
The following block writes report.py and runs it:
# filename: report.py
import platform
print(platform.python_version())
Read the response​
The response follows the OpenAI chat completion shape. Three fields carry the result:
choices[0].message.contentholds the execution transcript, including the code output.usage.summaryholds the execution result, which is the exit code and the captured output. A successful run starts withexitcode: 0 (execution succeeded).- When you stream,
usage.file_ids_nameslists any files the code produced, as a list of single-entry objects that map a file ID to its filename, for example{"file-8f2a1c...": "report.py"}. Take the key of each entry as the file ID, then download it from the same base URL withGET /v1/files/{file_id}/content, orclient.files.content(file_id)in the Python SDK.
summary and file_ids_names are extra fields on the usage object. They resolve at runtime, but a type checker flags them, so read them through usage.model_extra["summary"] and usage.model_extra["file_ids_names"] in typed code.
Code that exits non-zero is not an API error. The request returns 200, and the failure shows up in usage.summary as a non-zero exit code with the error output from the code.
Troubleshooting​
| Symptom | Status | Cause | Fix |
|---|---|---|---|
AgentTool is not enabled on this server. | 400 | The request reached a server that is not the agent tool server. | Send the request to the tool port on the deployment network: 5005 on the shared container, 5007 on the isolated container. |
Agent is not enabled on this server. | 400 | The request set use_agent instead of use_agent_tool. | Set use_agent_tool to true, or send agent requests to the agent server port. |
Server is currently processing the maximum number of concurrent requests | 503 | A second request reached the same server in the isolated container while its first request was still running. | Retry after that request finishes. The agent and tool servers have separate limits. |
| The response contains no code output. | 200 | The message had no fenced code block, or the only block carried # execution: false. | Wrap the code in a fence with a language tag, and remove # execution: false if you want the code to run. |
Next steps​
- Learn about General code tools for extending agent chats with your own Python code
- Review the Agents overview for the full agent path this server turns off
- See APIs to create the Enterprise h2oGPTe API key used everywhere else in the product
- Submit and view feedback for this page
- Send feedback about Enterprise h2oGPTe to cloud-feedback@h2o.ai