Academy → Hermes FeaturesOfficial documentation · Arabic guidance

Document Extraction

استخراج النصوص من المستندات

Intermediate3 min readLesson 61 question✓ 2026-08-18
Before you read

What this page is, and what it holds.

“Document Extraction” from the official documentation, reproduced in full. About 3 minutes to read.

2sections
2code examples
1tables
0commands
504source words
The official one-line description

How readfile converts PDFs, Office documents, and notebooks to text — and what to do when a PDF is scanned images

What you will be able to do

Outcomes taken from this page, not a template.

  • Read the table and take only the row that applies to you.
Page map

Jump to the part you need.

  1. 01Supported formats
  2. 02Scanned PDFs: the coverage warning
The full official page

Nothing summarised away.

The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.

The read_file tool automatically converts common document formats to readable text, so the agent can inspect a PDF or spreadsheet the same way it reads source code.

Supported formats

A lookup table. Do not read it all; find the row that applies to you.

FormatExtensionsConverterAvailability
Jupyter notebooks.ipynbBuilt-in (stdlib)Always
Word documents.docxBuilt-in (stdlib)Always
Excel workbooks.xlsxBuilt-in (stdlib)Always
PDF.pdfOptional anydoc converterAuto-installed on first use*
Legacy Office.doc, .ppt, .xls, .pptx, and variantsOptional anydoc converterAuto-installed on first use*
OpenDocument.odt, .ods, .odpOptional anydoc converterAuto-installed on first use*
Rich text / eBooks.rtf, .epubOptional anydoc converterAuto-installed on first use*

\* The optional converter is the firecrawl-anydoc package, installed lazily where installs are permitted (security.allow_lazy_installs in config.yaml). Without it, the three stdlib formats still work; other formats fall back to the binary-file guard.

Conversion output is Markdown, paginated through read_file's normal offset/limit window. Documents over 50 MB are refused to keep tool turns bounded.

Extraction works with remote terminal backends (Docker, Modal, SSH): the file's bytes are transferred across the backend boundary and converted host-side, so a document inside a sandbox reads the same as a local one.

Scanned PDFs: the coverage warning

Explains the idea itself. Read it slowly; the later sections build on it.

PDF conversion reads the text layer only. Pages that are scanned images — common in legal documents, resale packages, signed contracts, faxes — contain no text layer and silently convert to nothing. The telltale signature is section headers with empty bodies.

When a meaningful share of pages yields no text (over 20% of the document, or 10+ pages absolute), read_file prepends a warning to the extraction. Each unreadable gap is labeled with the last text extracted before it — usually a section divider — so the agent can target only the gaps it actually needs instead of OCRing the whole document:

Text7 lines
[EXTRACTION COVERAGE WARNING: 198 of 311 pages in this PDF yielded no
text. ... Unreadable gaps, each labeled with the last text extracted
before it:
  pages 42-77 (36 pages) — after "Antigua Maintenance Corp Bylaws" (p41)
  pages 92-213 (122 pages) — after "... Covenants, Codes and Regulations" (p91)
  page 224 (1 page) — after "... Insurance Declaration Pages" (p223)
Decide which gaps you actually need — do NOT OCR or render everything. ...]

The warning lists the exact page ranges and the recovery paths:

  1. A few pages — render + vision. Convert the pages to images and read them with the vision tool:
Shell1 line
   pdftoppm -jpeg -r 150 -f 92 -l 94 document.pdf /tmp/page

Then inspect each image with vision_analyze. Zero extra dependencies (poppler is required for the detection itself).

  1. Many pages — OCR. The ocr-and-documents skill covers bulk OCR with marker-pdf (90+ languages, handles equations and tables; ~3-5 GB install).

Detection uses poppler's pdftotext for per-page text counts. If poppler is not installed, extraction still works — the coverage check is silently skipped.

Knowledge check

1 question answered by this page alone.

Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.

1. In this lesson's table, what is the “Extensions” for “OpenDocument”?