Document Extraction
استخراج النصوص من المستندات
What this page is, and what it holds.
“Document Extraction” from the official documentation, reproduced in full. About 3 minutes to read.
How readfile converts PDFs, Office documents, and notebooks to text — and what to do when a PDF is scanned images
Outcomes taken from this page, not a template.
- Read the table and take only the row that applies to you.
Jump to the part you need.
Nothing summarised away.
The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.
The read_file tool automatically converts common document formats to readable text, so the agent can inspect a PDF or spreadsheet the same way it reads source code.
Supported formats
A lookup table. Do not read it all; find the row that applies to you.
| Format | Extensions | Converter | Availability |
|---|---|---|---|
| Jupyter notebooks | .ipynb | Built-in (stdlib) | Always |
| Word documents | .docx | Built-in (stdlib) | Always |
| Excel workbooks | .xlsx | Built-in (stdlib) | Always |
.pdf | Optional anydoc converter | Auto-installed on first use* | |
| Legacy Office | .doc, .ppt, .xls, .pptx, and variants | Optional anydoc converter | Auto-installed on first use* |
| OpenDocument | .odt, .ods, .odp | Optional anydoc converter | Auto-installed on first use* |
| Rich text / eBooks | .rtf, .epub | Optional anydoc converter | Auto-installed on first use* |
\* The optional converter is the firecrawl-anydoc package, installed lazily where installs are permitted (security.allow_lazy_installs in config.yaml). Without it, the three stdlib formats still work; other formats fall back to the binary-file guard.
Conversion output is Markdown, paginated through read_file's normal offset/limit window. Documents over 50 MB are refused to keep tool turns bounded.
Extraction works with remote terminal backends (Docker, Modal, SSH): the file's bytes are transferred across the backend boundary and converted host-side, so a document inside a sandbox reads the same as a local one.
Scanned PDFs: the coverage warning
Explains the idea itself. Read it slowly; the later sections build on it.
PDF conversion reads the text layer only. Pages that are scanned images — common in legal documents, resale packages, signed contracts, faxes — contain no text layer and silently convert to nothing. The telltale signature is section headers with empty bodies.
When a meaningful share of pages yields no text (over 20% of the document, or 10+ pages absolute), read_file prepends a warning to the extraction. Each unreadable gap is labeled with the last text extracted before it — usually a section divider — so the agent can target only the gaps it actually needs instead of OCRing the whole document:
[EXTRACTION COVERAGE WARNING: 198 of 311 pages in this PDF yielded no
text. ... Unreadable gaps, each labeled with the last text extracted
before it:
pages 42-77 (36 pages) — after "Antigua Maintenance Corp Bylaws" (p41)
pages 92-213 (122 pages) — after "... Covenants, Codes and Regulations" (p91)
page 224 (1 page) — after "... Insurance Declaration Pages" (p223)
Decide which gaps you actually need — do NOT OCR or render everything. ...]The warning lists the exact page ranges and the recovery paths:
- A few pages — render + vision. Convert the pages to images and read them with the vision tool:
pdftoppm -jpeg -r 150 -f 92 -l 94 document.pdf /tmp/pageThen inspect each image with vision_analyze. Zero extra dependencies (poppler is required for the detection itself).
- Many pages — OCR. The
ocr-and-documentsskill covers bulk OCR with marker-pdf (90+ languages, handles equations and tables; ~3-5 GB install).
Detection uses poppler's pdftotext for per-page text counts. If poppler is not installed, extraction still works — the coverage check is silently skipped.
1 question answered by this page alone.
Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.