Use memory blocks in chat
Attach a memory block to a chat to give the LLM or agent persistent context across sessions. Depending on the access mode, the LLM or agent can also update the memory block during the conversation.
Attach a memory block to a chat​
- Open the Customize chat panel.
- Select the Configuration tab.
- From the Memory Block dropdown, select a memory block.

You can also attach a memory block through llm_args when sending a query. Reference the memory block by ID or name.
By ID:
with client.connect(chat_session_id) as session:
reply = session.query(
message="What is our project reference number?",
llm_args={"memory_block_id": "your-memory-block-uuid"},
timeout=120,
)
By name:
with client.connect(chat_session_id) as session:
reply = session.query(
message="What is our project reference number?",
llm_args={"memory_block_name": "Project Knowledge"},
timeout=120,
)
Name lookup matches only memory blocks owned by the current user with that exact name. To use a shared or public memory block, pass its memory_block_id instead.
Use a memory block with an agent​
Include both the memory block reference and use_agent: True in llm_args:
with client.connect(chat_session_id) as session:
reply = session.query(
message="Analyze our Q1 sales data and save key findings.",
llm_args={
"memory_block_id": "your-memory-block-uuid",
"use_agent": True,
"max_time": 90,
},
timeout=180,
)
Use a memory block with an AI Assistant​
When creating an AI Assistant, select a memory block from the Memory Block dropdown. All chat sessions with that assistant use the selected memory block automatically without passing it in llm_args.
Injection modes​
| Mode | Value | Behavior | Best for |
|---|---|---|---|
| System prompt | system_prompt (default) | Wraps content in <agent_memory name="..."> XML tags and appends it to the system prompt. | Persistent background context. |
| User instruction | user_instruction | Wraps content in <agent_memory name="..."> XML tags and prepends it to the user's message. | When memory should take precedence over system prompt instructions. |
| Agent file | agent_file | Writes content to an AGENTS.md file in the agent's working directory. The agent reads and updates this file directly. | Agent chats where the agent manages memory structure. |
| Agent tool | agent_tool | Adds nothing to the prompt and creates no file. The assistant calls Read memory to fetch the content when it needs it, and Update memory to save changes, either replacing the content or appending to it. | AI Assistants that should decide when to read and write memory, especially for large memory blocks you don't want injected into every prompt. |
Agent file and Agent tool describe mechanisms that exist only during an agent run. In a non-agent (LLM only) chat, there is no working directory and no memory tools, so Enterprise h2oGPTe falls back to the System prompt behavior: the content is added to the system prompt, and the LLM writes back with <memory_update> tags. The access mode still applies, so a Write only block has no content added to the prompt.
Agent tool mode only works with AI Assistants, not regular agent chats. A regular agent chat can't reach an agent_tool block — use Agent file or System prompt mode there instead.
Access modes​
The access mode determines whether the LLM or agent can read the memory, write to it, or both.
| Mode | Value | LLM chats | Agent chats (agent_file) | Best for |
|---|---|---|---|---|
| Read & Write | read_write (default) | Content injected; LLM uses <memory_update> tags to save new information. | AGENTS.md created with content; agent reads and updates it. | General-purpose memory that accumulates knowledge. |
| Read only | read | Content injected; <memory_update> tags ignored. | AGENTS.md created as read-only. | Stable reference data (style guides, compliance rules). |
| Write only | write | Content not injected; LLM can write with <memory_update> tags. | Header-only AGENTS.md created for the agent to populate. | Note-taking without influence from previous content. |
An AI Assistant with a connected memory block gets the Read memory and Update memory tools according to the block's access mode — see Tool visibility is controlled by access mode. This applies regardless of injection mode, but in Agent tool mode these tools are the assistant's only way to reach the memory.
How LLMs update memory​
In non-agent LLM chats with write or read-write access, the LLM wraps new information in <memory_update> XML tags:
<memory_update>Customer confirmed budget of $50,000 for Q2.</memory_update>
Enterprise h2oGPTe extracts the content from these tags, appends it to the existing memory block, and strips the tags from the visible response.
If the LLM places its entire response inside <memory_update> tags, the visible reply appears empty. The memory block still updates correctly.
How agents update memory with agent_file​
Enterprise h2oGPTe writes the memory block content to an AGENTS.md file in the agent's working directory before execution. The agent reads and modifies this file during its run. After execution, Enterprise h2oGPTe saves the final AGENTS.md content back to the memory block.
In agent mode, AGENTS.md content replaces the memory block entirely (no append). The agent must preserve any existing information it needs to keep.
How agent tool mode works​
With agent_tool injection mode, the memory block's content is never pre-injected into the system prompt, the user message, or an AGENTS.md file. Instead, the assistant reads and updates memory explicitly, on its own schedule, by calling two tools:
| Tool palette label | Function called | What it does |
|---|---|---|
| Read memory | read_memory_block() | Fetches the current content of the connected memory block. |
| Update memory | update_memory_block() | Saves new or changed content back to the memory block. |
Because nothing is injected up front, the assistant must decide to call Read memory before it can use existing memory content. Unlike agent_file, there is no automatic sync at the end of the run.
Nothing is saved unless the assistant calls Update memory.
Tool visibility is controlled by access mode​
Which of the two tools appear in the assistant's tool palette is determined by the connected memory block's access mode, not by the assistant's own tool settings:
| Access mode | Tools shown in the palette |
|---|---|
read | Only Read memory |
write | Only Update memory |
read_write | Both Read memory and Update memory |
| No memory block connected | Neither tool |
For a memory tool that the access mode allows, your Allow, Approve, or Deny setting on that tab still applies. For example, set Update memory to Approve to review each memory write before the assistant saves it.
In the assistant's Custom tool tab, tools the access mode doesn't allow show a lock icon with a tooltip and can't be changed (locked to Deny). Tools the access mode does allow appear as normal rows you can still edit. Edit the connected memory block's access mode to change which tools are available, instead of changing the tool permission directly.
Content truncation​
The max_content_length field controls how much content the memory block stores. When content exceeds this limit, Enterprise h2oGPTe truncates it and keeps the most recent content. Set to 0 to turn off truncation. Default: 10,000 characters.
Larger memory blocks consume more of the model's context window. Choose a limit that balances context richness with prompt size.
- Submit and view feedback for this page
- Send feedback about Enterprise h2oGPTe to cloud-feedback@h2o.ai