Query a collection
Two AI assistant tool actions run a retrieval-augmented generation (RAG) query against a document collection on an assistant's behalf: create_chat_session and query_collection. Both take a collection_id and a question, then return an answer grounded in that collection.
The tool action create_chat_session shares its name with a Python client method that creates a persistent chat session. They're unrelated. The tool action runs one question and returns immediately; it doesn't create a session.
Compare the two actions​
The following table compares the two actions:
| Attribute | create_chat_session | query_collection |
|---|---|---|
| Parameters | collection_id, message, llm (optional), system_prompt (optional) | collection_id, question, llm (optional) |
| Returns | The answer plus a list of references to the retrieved chunks | The answer only, with no references |
| Error signal | Sets error to true in the returned dict | None. An error message is indistinguishable from an answer |
| Tool permissions panel | Listed as Chat session (RAG). Set it to Allow, Approve, or Deny like any other tool action. | Not listed. You can't set an Approve or Deny policy for it, so it runs whenever the assistant's code calls it. It's unavailable only when you replace the assistant's generated code with a Custom tool definition, or when the assistant's creation spec lists allowed tool actions and omits it. |
The platform generates both actions as part of the assistant's tool code. If you supply your own code in Custom tool definition on the Custom tool tab, it replaces the generated code, and neither action is available unless you define it yourself. An assistant created from a spec that lists its allowed tool actions gets only the actions on that list.
Call create_chat_session​
Call the action with a collection ID and a question:
create_chat_session(
collection_id="3f9c2b6a-8e41-4c9d-9a2b-7d6f1e5c4a30",
message="What is the escalation SLA for a P1 incident?",
)
The response is a dict:
{
"response": "P1 incidents must be escalated to the on-call lead within 15 minutes of being opened.",
"collection_id": "3f9c2b6a-8e41-4c9d-9a2b-7d6f1e5c4a30",
"references": [
{
"document_id": "b7a3f210-5c8e-4d1a-9f60-2e7b1c4a9d33",
"document_name": "incident-response-runbook.pdf",
"chunk_id": 42,
"pages": "12-13",
"score": 0.87,
"content": "P1 incidents require escalation to the on-call lead within 15 minutes of being opened."
}
]
}
On failure, the dict also has error set to true. Check error before you use response as an answer. On failure, response holds a user-facing error message, not a grounded answer.
Don't use references to detect failure. A successful call returns an empty list when retrieval finds no matching chunks, and a call that fails after retrieval can still return references.
llm overrides the model used to generate the answer. system_prompt sets the system prompt for this one call. Without it, the call uses the default prompt for collection-grounded retrieval, not the assistant's own system prompt.
Call query_collection​
Call the action with a collection ID and a question:
query_collection(
collection_id="3f9c2b6a-8e41-4c9d-9a2b-7d6f1e5c4a30",
question="What is the escalation SLA for a P1 incident?",
)
The response is the answer text only:
"P1 incidents must be escalated to the on-call lead within 15 minutes of being opened."
Access to the collection​
Neither action requires you to link the collection to the assistant. Both take an explicit collection_id. Access follows the same rules as anywhere else in Enterprise h2oGPTe. You own the collection, it's shared with you directly, it's shared with one of your groups, it's public, or you're an API-key user granted access to it.
Both actions also require access to the assistant itself, checked separately from collection access.
Linking a collection on the assistant's Linked Collections tab doesn't grant access. It's what the assistant's list_connected_collections action returns instead. create_chat_session and query_collection can query any collection you have access to, including unlinked ones.
What controls retrieval​
These actions don't use the assistant's own Expert Settings. Settings such as Temperature, Output Token Limit, Top K Chunks, Number of neighboring chunks to include for Summary RAG, and Include Chat Conversation History apply only to the assistant's own conversation.
Each call carries no conversation history, so include everything the question needs in message or question.
Instead, both actions retrieve using your deployment's default retrieval configuration, which an administrator sets, and a fixed generation configuration you can't change from the assistant's settings. If your administrator sets the deployment's default retrieval mode to one that skips retrieval, these calls return an answer that isn't grounded in the collection, and references is empty.
Errors and limits​
If the collection doesn't exist, you can't access it, collection_id isn't a valid UUID, or the call exceeds its wait, the call raises an error instead of returning a message. Wrap the call if your assistant's code needs to recover from that.
If a guardrail blocks the query or the generated answer, both actions return a message saying content guardrails blocked the request, instead of the underlying guardrail error.
If the collection's embedding model is no longer supported, both actions return a message telling you to re-import the documents into a new collection.
Any other retrieval failure returns a generic message asking you to retry or contact your administrator.
Limits​
Both actions limit message and question to 50,000 characters and reject an empty value. A longer value raises an error. query_collection waits less time for a response than create_chat_session, so a long-running query is more likely to fail there. Use create_chat_session for questions that need a long answer.
Related topics​
- Submit and view feedback for this page
- Send feedback about Enterprise h2oGPTe to cloud-feedback@h2o.ai