الأكاديمية ← ميزات Hermesتوثيق رسمي · إرشاد عربي

استخراج النصوص من المستندات

Document Extraction

متوسط3 دقائق قراءةالدرس 6سؤال واحد✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

صفحة «استخراج النصوص من المستندات» من التوثيق الرسمي، معروضة هنا كاملة. القراءة نحو 3 دقائق.

2أقسام
2أمثلة برمجية
1جداول
0أوامر
504كلمة من المصدر
الوصف الرسمي في سطر

How readfile converts PDFs, Office documents, and notebooks to text — and what to do when a PDF is scanned images

ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01Supported formats
  2. 02Scanned PDFs: the coverage warning
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

The read_file tool automatically converts common document formats to readable text, so the agent can inspect a PDF or spreadsheet the same way it reads source code.

Supported formats

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

FormatExtensionsConverterAvailability
Jupyter notebooks.ipynbBuilt-in (stdlib)Always
Word documents.docxBuilt-in (stdlib)Always
Excel workbooks.xlsxBuilt-in (stdlib)Always
PDF.pdfOptional anydoc converterAuto-installed on first use*
Legacy Office.doc, .ppt, .xls, .pptx, and variantsOptional anydoc converterAuto-installed on first use*
OpenDocument.odt, .ods, .odpOptional anydoc converterAuto-installed on first use*
Rich text / eBooks.rtf, .epubOptional anydoc converterAuto-installed on first use*

\* The optional converter is the firecrawl-anydoc package, installed lazily where installs are permitted (security.allow_lazy_installs in config.yaml). Without it, the three stdlib formats still work; other formats fall back to the binary-file guard.

Conversion output is Markdown, paginated through read_file's normal offset/limit window. Documents over 50 MB are refused to keep tool turns bounded.

Extraction works with remote terminal backends (Docker, Modal, SSH): the file's bytes are transferred across the backend boundary and converted host-side, so a document inside a sandbox reads the same as a local one.

Scanned PDFs: the coverage warning

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

PDF conversion reads the text layer only. Pages that are scanned images — common in legal documents, resale packages, signed contracts, faxes — contain no text layer and silently convert to nothing. The telltale signature is section headers with empty bodies.

When a meaningful share of pages yields no text (over 20% of the document, or 10+ pages absolute), read_file prepends a warning to the extraction. Each unreadable gap is labeled with the last text extracted before it — usually a section divider — so the agent can target only the gaps it actually needs instead of OCRing the whole document:

Text7 أسطر
[EXTRACTION COVERAGE WARNING: 198 of 311 pages in this PDF yielded no
text. ... Unreadable gaps, each labeled with the last text extracted
before it:
  pages 42-77 (36 pages) — after "Antigua Maintenance Corp Bylaws" (p41)
  pages 92-213 (122 pages) — after "... Covenants, Codes and Regulations" (p91)
  page 224 (1 page) — after "... Insurance Declaration Pages" (p223)
Decide which gaps you actually need — do NOT OCR or render everything. ...]

The warning lists the exact page ranges and the recovery paths:

  1. A few pages — render + vision. Convert the pages to images and read them with the vision tool:
Shellسطر واحد
   pdftoppm -jpeg -r 150 -f 92 -l 94 document.pdf /tmp/page

Then inspect each image with vision_analyze. Zero extra dependencies (poppler is required for the detection itself).

  1. Many pages — OCR. The ocr-and-documents skill covers bulk OCR with marker-pdf (90+ languages, handles equations and tables; ~3-5 GB install).

Detection uses poppler's pdftotext for per-page text counts. If poppler is not installed, extraction still works — the coverage check is silently skipped.

اختبار الفهم

سؤال واحد إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. في جدول هذا الدرس، ما «Extensions» المقابل لـ«OpenDocument»؟