Skip to main content
Version: v1.7.5

v1.7.5

Released Sep 17, 2026.

This release adds AI Assistants, agent skills, and an LLM registry, rewrites the guardrail categories and gives Prompt Guard a new backend, and fixes scheduling and Graph RAG issues that could affect existing collections.

Breaking changes​

  • Several models are no longer in the default model list: gemini-2.5-pro and gemini-2.5-flash, which Google now limits to users who already used them, and gpt-4o-mini, gpt-4.1, gpt-4.1-nano, o1-mini, o4-mini, Qwen/Qwen3-235B-A22B-FP8, and nvidia/Llama-3.3-Nemotron-Super-49B-v1.5 are no longer in the default model list. If you use one of them, switch to a supported model before you upgrade, or have an administrator add it back with Models > Add Model if its provider still serves it.

New features​

Agents and tools

  • AI Assistants: AI Assistants are persistent AI workers that wake on a trigger (a new message, a forum post, or a scheduled task), do their work using agent tools, and go back to waiting. Between runs, an assistant keeps its document collections, memory blocks, and files, so each run builds on the last. Assistants collaborate through forums and direct messages, and can require human approval before taking sensitive actions.
  • Agent skills and the h2oGPTe Library: Package reusable agent capabilities as skills, and browse a shared library of skills instead of rebuilding the same tool logic for every agent.
  • Custom tool file management: Upload or replace the underlying files for remote MCP tools and other custom tools directly from the UI.
  • Accurate collection document counts: An assistant's agent can now tell your collection's documents apart from its own working files, so it correctly answers questions like "how many documents are in your collection?"
  • Collection-grounded RAG for assistant tool actions: An assistant's generated tool code can now run a RAG query against a document collection through create_chat_session and query_collection, and return an answer grounded in that collection.
  • Session files in the chat side panel: The chat's Files tab now lists every file an agent saved during the session, not only files attached to a reply.
  • Snowflake OAuth sign-in: A new predefined Snowflake Database (SSO) tool queries Snowflake as the signed-in user, using the access token from your organization's identity provider. An administrator must set up identity-provider token forwarding and a Snowflake external OAuth integration.
  • Slack connector: An administrator registers the Slack app and chooses its scopes. The tool acts as you by default, and acts as the workspace bot, only in channels the bot was invited to, when the app grants bot scopes only. Connectors must be enabled for your deployment.
  • CodeAPI persistent MCP tools: CodeAPI tools are available as persistent MCP tools, off by default until an administrator turns them on.

Models

  • Supported model updates: Added moonshotai/Kimi-K3 via OpenRouter and moonshotai/Kimi-K2.5 via Azure. Accuracy figures for newly added models are estimates rather than benchmark results. See Known issues.
  • LLM registry: Administrators can manage runtime LLMs from a dedicated administrator UI, control access with role-based permissions, and switch which LLM serves a given model slot. Editing an existing runtime LLM's configuration is unavailable in this release. See Known issues.

Security and guardrails

  • Guardrail categories rewritten: All 14 guardrail categories have been rewritten. Verdicts on existing content can change after upgrading. See Known issues.
  • Prompt Guard uses your guardrails LLM: When jailbreak detection (Prompt Guard) is turned on, it now uses your guardrails LLM instead of a separate classifier model, and it blocks a prompt when it can't read the verdict. Jailbreak detection stays off by default.

Chat

  • Suggested follow-up chips: Chat responses can show suggested follow-up questions as clickable chips. Administrators control availability with the user_suggested_follow_ups setting: off disables chips for everyone, on lets each user set their own preference in Preferences. Off by default.
  • Turn-completion notifications: While an agent run is in progress, select the bell on its response (Notify me by email when this run finishes) to get an email when the run ends. It requires system email set up by an administrator, or your own Email notifications (Gmail) setup in the profile menu.
  • Chat session Auto-Archive: Chat sessions now follow the same lifecycle as collections. A chat moves from Active to Expiring during a warning window, then to Archived, and is permanently deleted after a grace period. Transitions apply hourly.
  • Chat lifecycle settings: chat_session_inactivity_days (default -1, off) sets the days of inactivity that make a chat eligible for archiving. chat_expiration_limit_days (default 30) sets the warning window, the grace period after archiving, and the latest expiry date a user can set.
  • Chats inside collections: For a chat inside a collection, the earlier of the chat and collection deadlines applies, for both expiry dates and inactivity.
  • Preserve individual chats: Users can turn off auto-archiving for their own chats from the Auto-Archive tab, gated by the new Preserve chat sessions from global expiration (h2ogpte/chat/keep_forever) permission. Users can also set a fixed expiry date on a chat.
  • Per-collection exemption: Administrators can exempt a collection from the global chat inactivity policy on the Manage collections page. Expiry dates set on individual chats still apply.
  • Archived chat administration: A new Archived Chats filter lists archived sessions for recovery. The chats table has sortable Status, Expiry Date, and Inactivity Interval columns, and POST /admin/chats/cleanup runs an on-demand cleanup with an optional force flag.
  • Preserved chats survive collection cleanup: When automatic cleanup removes an archived collection, its preserved chats stay as standalone chats. Deleting a collection explicitly still removes every chat in it.

Documents and collections

  • Precise collection expiry: Set a collection to expire at an exact date, time, and timezone instead of a date alone, and h2oGPTe archives it within a minute of that moment. Available in the UI, the REST API, and the SDK.
  • Stale description indicator: Collections now flag when an auto-generated description is stale, so you know when it's worth regenerating.
  • Source chat for agent-generated files: The Documents page now shows which chat session produced each agent-generated file, with a link to open that chat, plus Filter by chat and Group by chat options to bring one session's files together. GET /documents returns the source session and accepts the same filter and grouping parameters.
  • Tag filtering on GET /tags: GET /tags now returns every tag by default, including tags not linked to any document. Pass linked_only=true to return only tags attached to at least one document, and user_scope=true to limit results to tags you can access. The two parameters compose.

Scheduled tasks

  • Adaptive /schedule and /create-assistant commands: These chat commands now confirm based on how much they had to assume. A fully specified request is created right away, a request that relied on a default, such as an assumed time zone, asks you to reply yes or no in the chat, and a vague request opens the pre-filled form.
  • Scheduled task email delivery status: A scheduled task's Task Runs list and the Activity tab now show an Email Status column with Email Sent, Email Pending, or Email Failed for each run's notification.

Administration

  • Redesigned Jobs page: Job history now persists across restarts instead of clearing each time, with a filterable administrator UI, full error logs, and access to job data through the API and SDK.
  • Redesigned Usage Insights: The Usage and Quota page (titled Usage & Feedback) is now Usage Insights, with four tabs instead of five: Usage, which combines the former Dashboard and Token Usage tabs and opens first, Observability, Guardrails for administrators, and Feedback.

Authentication and access

  • Per-role and per-user AI Assistants gating: Administrators can turn AI Assistants on or off for specific roles or specific users, instead of only for the whole deployment.
  • Restrict the IdP Groups share tab: A dedicated permission lets administrators hide the IdP Groups tab from the share dialog.

Improvements​

Agents and tools

  • System prompts generate on upload: Uploading code to a General Code tool now generates its system prompt automatically. You no longer need to click the sparkle button.

Scheduled tasks

  • Simplified scheduled tasks UI: The scheduled tasks UI no longer shows redundant controls.
  • Activity filters: You can now filter the scheduled tasks Activity view.
  • Session distinction: Chats that a scheduled task started now show an Auto-scheduled icon on the Chats page, in a collection's chat list, and in Usage Insights.

Administration

  • Visible storage migration progress: When a storage migration is running, administrators see live progress and a clear failure state instead of a blocked sign-in with no explanation.
  • Consistent unlimited value: The per-user concurrent ingestion and chat limits now use -1 for unlimited, consistent with other limit settings. Setting either to 0 still works the same way.
  • Runtime settings can't be overridden: You can no longer mark a runtime setting, such as Output Token Limit, as overridable per user or role. Such overrides never took effect.
  • Broader default audit coverage: The audit trail now covers new administrative and mutating actions by default, without needing per-action configuration. User creation, sessions, and impersonation are not yet covered. See Known issues.

Documents and collections

  • Ingestion skips unsupported content: Ingestion now skips unsupported file types and unreadable archives instead of failing the whole job.
  • Automatic image rotation: Ingestion now uses a dedicated model to detect rotated scanned or photographed pages, and falls back to the previous OCR-based check when the model isn't confident.

Python client​

  • Adds timestamps for each segment of a transcribed audio or video document.
  • Adds PATCH support to update an existing extractor without deleting and recreating it.

Bug fixes​

Models

  • Fixed the OpenAI-compatible endpoints rejecting a base model name in the model field.

RAG

  • Fixed Graph RAG cutting off long extraction responses and dropping chunks whose extraction took longer than two minutes, which could leave a knowledge graph with few or no relationships.

Scheduled tasks

  • Fixed a one-time task scheduled for a local time running in UTC, so a task set for 9 AM in New York ran hours early.
  • Fixed a one-time task that was changed to a repeating schedule stopping after its first run.

User interface

  • Fixed Arabic and Spanish translation issues across the UI.
  • Fixed several accessibility issues, including a skip-to-content link, unique page titles per view, and improved keyboard focus handling.

Known issues​

  • AI Assistants UI limitations remain. Missing list pagination, duplicate queries, an auto-refresh that isn't gated, incorrect chat ordering, editing while a run is in progress, silent failures that show the wrong status badge, and settings that get overwritten are all still open for AI Assistants in this release.
  • Assistant notification email is set at creation only: The edit dialog doesn't include Notification email, so set it when you create the assistant.
  • No transcript export. The SDK can retrieve audio timestamps, but there's no UI to export a transcript in this release.
  • User creation, sessions, and impersonation are not audited. These actions don't yet reach the audit trail in 1.7.5, and related audit events don't record which method produced them. Both are compliance gaps to plan around until they're addressed.
  • Existing knowledge graphs aren't rebuilt automatically. Graphs built before this release keep their previous relationships, and adding documents updates only the new documents. To rebuild a graph with this release's fixes, select Rebuild Knowledge Graph in the collection's actions menu, or call build_collection_graph() with force_rebuild=True.
  • Guardrail verdicts can change after upgrading. The 14 rewritten guardrail categories, and the new Prompt Guard backend if you turn on jailbreak detection, can classify the same content differently than before the upgrade.
  • Editing a runtime LLM is unavailable. You can't edit an existing runtime LLM's configuration in this release; the capability returns in the next release.
  • Non-administrators see raw error text for their own jobs. Job error details aren't simplified for non-administrator users yet.
  • New model accuracy figures are estimates. Accuracy figures published for newly added models are estimates, not the result of h2oGPTe benchmark runs.
  • Existing once-only tasks keep their old execution limit. Scheduled tasks created before this release aren't migrated to the new once-task execution limit.
  • Interval schedules drift. An "every N hours" schedule counts from when the previous run finished, not from a fixed clock, so it drifts by the run's duration.
  • Gmail sender address isn't encrypted: h2oGPTe encrypts the app password you save in Email notifications (Gmail), but stores the sender address in plain text.

Feedback