Document extraction
Some documents are pictures of text: scanned books, photocopied archives, vertically typeset pages. Extraction reads those and returns text and structure you can search and edit.
Vertical Japanese and Chinese typesetting is converted to horizontal output, and multi-column pages come back in reading order rather than as one scrambled block.
When to use it
The line against format conversion is whether the text is still locked inside an image. Scans, photos and old vertically-set books belong here, because the characters have to be recognised first. Documents that are already clean and digital, where the text can be selected and only the extension is wrong, go through format conversion — faster and cheaper in credits.
Extraction only, no translation — whatever language goes in comes back out. Use file translation when you need a translation.
What it accepts and produces
It accepts PDF, images (PNG / JPG / WEBP / TIFF and similar), DOCX, EPUB and Markdown, and produces Word, PDF, TXT or Markdown. Billing is per page: 4 credits a page, and a single image counts as one page.
It runs as a background job, so you can close the page and pick the result up later under History. Finished jobs can be previewed online.
When extraction is most accurate
- Clean print, scanned straight, gives the highest recognition rate.
- Vertical Japanese, two-page spreads and other unusual reading orders are corrected automatically — no need to split pages yourself.
- Where the image is blurry or largely handwritten, checking against the original before exporting is the safer move.
Choosing among the four exports
| Export | Pick it when |
| Word | You'll keep editing, format it, or hand it to someone else to continue. |
| Markdown | It's going into a notes app, a knowledge base or your own writing setup. |
| TXT | You want plain text to feed another tool or to search through. |
| PDF | You'll just read and archive it. |
Frequently asked questions
Is vertical Japanese put back in the right order?
Yes. Vertical Japanese, two-page spreads and other unusual reading orders are corrected automatically and converted to horizontal output, with no need to split pages yourself. Multi-column pages come back in reading order rather than as one scrambled block.
How is it billed?
Per page: 4 credits a page, and a single image counts as one page.
Can it translate what it extracts?
Not on this page. When you need a translation, take the extracted Word or Markdown file into file translation and run it there.
中文版
Geyi is a proprietary AI document platform owned and operated by Nayuta Technology and Culture Company Limited.