Use a knowledge graph for collection retrieval
Graph RAG (also written GraphRAG) builds a knowledge graph from entities and relationships in collection documents. You can use Graph RAG to retrieve context across documents when standard retrieval returns isolated chunks.
Overview​
Graph RAG uses two retrieval signals:
- Standard hybrid search (vector and lexical).
- Graph-based traversal across connected entities.
Use Graph RAG when you need reasoning across documents:
- Your answer depends on facts from multiple documents.
- Your data has indirect relationships across projects, vendors, and customers.
- Your collection is large and standard RAG misses connected context.
How Graph RAG works​
Graph RAG runs in two phases.
Phase 1: Graph building​
When you build a knowledge graph, Enterprise h2oGPTe:
- Sends each document chunk to an LLM for entity extraction.
- Merges entities and relationships into one graph.
- Embeds entity descriptions for retrieval.
- Stores the graph in object storage (MinIO or S3).
Phase 2: Graph-augmented retrieval​
When you query with Graph RAG, Enterprise h2oGPTe:
- Runs standard hybrid search to retrieve top chunks.
- Runs a graph query to find entities related to your question.
- Searches the collection for chunks that mention those entities, then merges those chunks into the results.
- Adds graph analysis to the LLM context.
- Generates a response from chunk context and graph context.
Build a knowledge graph​
Build prerequisites​
- You have a Collection with ingested documents.
- Your deployment has at least one LLM for entity extraction.
Build the graph​
- Open the target Collection page.
- Open the action menu (three-dot icon) next to Add Documents in the Documents section.
- Click Build Knowledge Graph.
- Select the LLM for entity extraction.
- Start the build.
You can also build the graph from the Chat side panel. Open the action menu next to Add Documents in the Documents section of the side panel.
If you do not see Build Knowledge Graph in the action menu, verify your role permissions. Make sure at least one extraction LLM is configured and available. Ask your administrator to enable Knowledge Graph for your deployment version.
Track progress in the notification tray:
- Number of chunks processed.
- Selected LLM.
- Elapsed time.
Entity extraction is the most time-consuming step. Build times increase linearly with the number of chunks. A collection with about 80 chunks typically takes 1 to 2 minutes.
Graph status​
After the build, the Collection page and the Chat side panel show the graph status next to the Documents heading:
- Graph Ready (green): The graph is built and current.
- Graph Outdated (yellow): New documents were added after the last build.
- Graph Building (blue): A build is in progress.
- Graph Failed (red): The build failed.
Rebuild or update the graph​
- Update Knowledge Graph: Processes only new documents since the last build.
- Rebuild Knowledge Graph: Rebuilds the graph from scratch.
Use Update Knowledge Graph after you add documents. Use Rebuild Knowledge Graph after you delete or heavily modify documents.
Query with Graph RAG​
Query prerequisites​
- The selected Collection has a completed knowledge graph.
- You can open a chat session with that collection.
Run a Graph RAG query​
- Open a Chat session with a Collection that has a built knowledge graph.
- In the Chat settings, set Generation Approach to Graph RAG.
- Ask your question.
Graph RAG only appears in the Generation Approach dropdown after a knowledge graph has been built for the collection. If you do not see it, build the graph first from the action menu on the Collection page or the Chat side panel.
Once a collection's graph reaches Graph Ready, the Automatic generation approach selects Graph RAG for that collection. Because Automatic is the default, chats left on it start using graph-augmented retrieval as soon as the build finishes, without anyone selecting Graph RAG. To keep Automatic on standard RAG even when a graph exists, set graph_rag_auto_select to false in the chat's rag_config.
If a query returns No knowledge graph found for this collection, the graph may have been deleted or the status is not ready. Return to the collection and rebuild the graph.
Graph RAG returns that error only when the collection has no ready graph. If the graph exists but produces no matching entities for a question, or the stored graph cannot be read, the query falls back to standard RAG and still returns an answer. To confirm that the graph contributed, check that the reply offers View KG grounding, or look for Falling back to standard RAG in the chat worker logs.
Graph RAG compared with standard RAG​
Use Graph RAG for questions that require reasoning across documents. Use standard RAG for factual questions from a single document.
| Aspect | Standard RAG | Graph RAG |
|---|---|---|
| Best for | Factual questions answered by one or a few chunks | Cross-document reasoning, indirect relationships |
| Retrieval method | Vector and lexical search over chunks | Hybrid search plus graph-based entity expansion |
| Setup | None beyond document ingestion | Requires building a knowledge graph first |
| Query latency | Lower — chunks are read from the vector store | Higher — downloads and extracts the graph per query |
| Example question | "What is the retention policy?" | "What is the connection between Project X and customer Y?" |
Admin configuration​
Deployment settings​
Administrators can configure Graph RAG with environment variables. Despite the H2OGPTE_CORE_ prefix, each variable is read by the service that runs its phase: the graph-build settings by the crawl service, the retrieval settings (H2OGPTE_CORE_GRAPH_RAG_MAX_ENTITY_SEARCHES and H2OGPTE_CORE_GRAPH_RAG_CHUNKS_PER_ENTITY) by the chat service, and H2OGPTE_CORE_GRAPH_RAG_COST_CONTROLS by both crawl and core.
In Docker Compose, these services all load the same .env file, so a single entry covers them. In Kubernetes, set the variable on crawl.extraEnv, chat.extraEnv, and core.extraEnv in your Helm values — the chart has no dedicated Graph RAG fields, so extraEnv is the only path, and a variable set only on core does not reach the build or query paths. All of these are read at process start, so restart the affected services after you change one.
The following environment variables are available:
H2OGPTE_CORE_GRAPH_RAG_LLM(default:auto): Selects the extraction LLM.autoprefers the cheapest non-reasoning model under the cost threshold. If no non-reasoning or affordable models are available, it falls back to any available model.H2OGPTE_CORE_GRAPH_RAG_COST_CONTROLS(default:{"max_cost_per_million_tokens": 5, "model": null}): Sets automatic model limits and allowlists.H2OGPTE_CORE_GRAPH_RAG_MAX_PARALLEL_INSERT(default:40): Controls chunk parallelism during graph build.H2OGPTE_CORE_GRAPH_RAG_LLM_MAX_ASYNC(default:80): Sets the maximum number of concurrent LLM calls.H2OGPTE_CORE_GRAPH_RAG_ENTITY_EXTRACT_MAX_GLEANING(default:0): Adds extra extraction passes per chunk.0disables extra passes.H2OGPTE_CORE_GRAPH_RAG_FORCE_LLM_SUMMARY_ON_MERGE(default:999999): Sets the fragment count that triggers merge-time LLM summarization. When an entity or relationship accumulates more description fragments than this value, Enterprise h2oGPTe summarizes them with an LLM call. Below the threshold, fragments are joined with a separator. Keep the default to turn merge-time summarization off.H2OGPTE_CORE_GRAPH_RAG_SUMMARY_MAX_TOKENS(default:500): Sets the maximum number of tokens in LLM-generated entity and relationship summaries during graph merge. This limit applies only whenH2OGPTE_CORE_GRAPH_RAG_FORCE_LLM_SUMMARY_ON_MERGEtriggers a summary.H2OGPTE_CORE_GRAPH_RAG_MAX_ENTITY_SEARCHES(default:15): Sets the maximum number of entity searches per query. Each search runs a keyword (full-text) lookup on the entity name against the collection's chunk store, so higher values can improve recall for entities named verbatim in the text but increase query latency. The searches run sequentially, and the effective ceiling is the number of entities the graph query returns for the question — raising this value above that number has no effect.H2OGPTE_CORE_GRAPH_RAG_CHUNKS_PER_ENTITY(default:5): Sets the maximum number of document chunks returned per entity search. Higher values can surface more supporting context but leave less room for other retrieved chunks.
Restricting available models​
To limit model choices for graph building, set model in H2OGPTE_CORE_GRAPH_RAG_COST_CONTROLS:
H2OGPTE_CORE_GRAPH_RAG_COST_CONTROLS: '{"max_cost_per_million_tokens": 5, "model": ["meta-llama/Llama-3.1-8B-Instruct", "gpt-4.1-nano"]}'
When you set model to a list, users only see those models in the graph build dropdown.
Use Graph RAG with the Python client​
Build a graph with the client​
Before you start, make sure you have a valid address and api_key, and the collection_id of the target collection.
Replace <YOUR_DOMAIN> with your Enterprise h2oGPTe hostname, <API_KEY> with an API key, and <COLLECTION_ID> with the ID of the collection you want to build the graph for:
from h2ogpte import H2OGPTE
client = H2OGPTE(address="https://<YOUR_DOMAIN>", api_key="<API_KEY>")
# Build the graph
job = client.build_collection_graph(
collection_id="<COLLECTION_ID>",
llm="meta-llama/Llama-3.1-8B-Instruct", # optional, defaults to auto
force_rebuild=False, # True to rebuild from scratch
timeout=600,
)
if job.errors:
raise RuntimeError(f"Graph build failed: {job.errors}")
print(f"Graph built in {job.duration}") # for example: Graph built in 00:01:47
The call returns when the job finishes, whether it succeeded or failed, so check job.errors before you treat the build as complete.
Parameter guidance:
- Set
llmto select a specific extraction model. - Set
force_rebuild=Trueto rebuild from scratch. - Set
timeoutto bound how long the client waits for a stalled build. The client enforces it only while a build reports no new progress, so a build that keeps progressing is not interrupted. The default is 86400 seconds.
Check graph status​
Use this call to confirm readiness before querying:
status = client.get_collection_graph_status(collection_id="<COLLECTION_ID>")
print(f"Status: {status.status}") # 'none', 'building', 'ready', 'failed'
print(f"Built at: {status.built_at}") # datetime, or None if never built
print(f"Outdated: {status.outdated}") # True if new docs added since build
Send a Graph RAG query​
Use rag_config={"rag_type": "graph_rag"} to force graph-based retrieval:
chat_session_id = client.create_chat_session(collection_id="<COLLECTION_ID>")
with client.connect(chat_session_id) as session:
reply = session.query(
"What is the connection between Project Atlas and customer retention?",
rag_config={"rag_type": "graph_rag"},
)
print(reply.content)
Known limitations​
- Build time scales with document count: Large collections can take longer to build.
- Entity merge ambiguity: Unrelated topics can produce incorrect merges for shared names such as "GPU" or "API."
- Model quality: Smaller extraction models can miss entities or relationships.
- Incremental update scope: Updates process new documents only; modifications and deletions require a rebuild.
- Query concurrency: Graph RAG retrieval is serialized within each chat worker process. Concurrent Graph RAG queries, including those on different collections, are processed one at a time.
- Per-query graph load: Each Graph RAG query downloads and extracts the collection's graph from object storage, so query latency grows with graph size and is higher than for standard RAG.
Related topics​
- Key terms - Review the definition of Graph RAG and other Enterprise h2oGPTe terms
- Chat settings - Compare Graph RAG with the other generation approaches
- Collections usage overview - Select a retrieval strategy for your use case
- Submit and view feedback for this page
- Send feedback about Enterprise h2oGPTe to cloud-feedback@h2o.ai