Skip to main content
Version: v1.7.5

Ingestion skips and automatic image rotation

Overview​

When you add documents to a Collection, Enterprise h2oGPTe processes each file independently. If one file in a batch can't be ingested — because its type isn't supported or it's an archive that can't be read — h2oGPTe skips that file and reports it, while the rest of the batch continues to ingest normally. Separately, when you upload an image file, h2oGPTe automatically detects and corrects its rotation before running OCR, so sideways or upside-down scans and photos are read correctly without any action on your part.

note

This page describes the default ingestion modes (Standard and Lite). In Agent Only mode, files are passed through as-is: unsupported types are not rejected, archives are not expanded, and automatic image rotation does not run. See Ingestion modes.

Unsupported file type skips​

If you upload a file whose type h2oGPTe doesn't support, that file is skipped. It does not stop or fail the rest of the upload batch.

When your upload finishes, a "Some files were not imported" toast appears with a summary like "N file(s) could not be imported. Open notifications for details." Open the notification tray, then open that job to see the specific reason for each skipped file, for example:

File 'notes.xyz' has an unsupported file type (.xyz) and was skipped.

note

For the full list of file types h2oGPTe supports, see Supported file types for a Collection.

Unreadable archive skips​

If you upload an archive (for example, a .zip file) as part of a Collection, h2oGPTe extracts it and ingests its contents. An archive is skipped, and reported the same way as above, in two cases:

  • The archive can't be read, for example because it's corrupted or truncated:

    File 'documents.zip' could not be read as an archive and was skipped.

  • The archive is readable but contains nothing to ingest, for example because it's empty or only contains files your operating system adds automatically (such as a __MACOSX folder or ._* resource-fork files on macOS):

    Archive 'documents.zip' yielded no ingestible files (empty, could not be decoded, or only OS metadata such as _MACOSX or . forks).

An archive is skipped as a whole: if any entry inside it can't be extracted, the whole archive is skipped, including entries that would otherwise have been ingested successfully. OS metadata files like __MACOSX and ._* are filtered out automatically and never counted as a failure.

Automatic image rotation​

Photographed or scanned pages are often rotated 90°, 180°, or 270° from upright, which can make OCR text extraction inaccurate. When you upload an image file, h2oGPTe automatically detects its rotation and straightens it as part of OCR, so you don't need to pre-rotate images yourself. Automatic rotation requires OCR. If the OCR model option is set to Disable in the Collection's upload settings, images keep their original orientation.

Rotation detection first tries a machine learning model trained specifically for this task. If the model's confidence is below an internal threshold, h2oGPTe falls back to an alternative method that tries a detected skew angle along with 90°, 180°, and 270°, and picks the one that gives the best OCR result. Either way, rotation correction never blocks or fails ingestion of the document — at worst, a page keeps its original orientation.

note

Automatic rotation applies to image file uploads (for example, photos and scans) converted to PDF during ingestion. It does not currently apply to files that are already in PDF format.

Next steps​


Feedback