Document to Word
Turn a PDF, a scan or a Word file into a genuinely editable Word document: digital text is extracted directly, scans and text-in-image are read with OCR, figures stay as images. Nothing is uploaded.
PDF and Word files are parsed inside your browser — nothing is uploaded.
How to use
- Choose files —— Drop in or pick PDF / Word documents, scans included — several can be queued at once. Files are read locally.
- Pick a recognition mode —— Automatic is the default: digital text is read directly and only scans go through OCR. For English-only files, switch the language to English only for a faster, more accurate run.
- Convert —— Click Convert. PDFs are parsed page by page. The first time OCR is needed the engine and language data (about 10MB) are downloaded; your browser caches them afterwards.
- Check the result —— Each row then shows a Download link plus the number of pages, characters, figures recovered and scanned pages recognised.
- Download —— Use Download for a single file, or Download all to save them one after another.
- Read the notice —— A "nothing could be recognised" warning means blank pages, plain photos or text too faint to read. Those pages are kept as images — no text is invented for them.
About Document to Word
"PDF to Word" sounds like one job. It is really three jobs with very different difficulty, so here is exactly how this tool handles each.
If your PDF is a digital document — exported from Word, a web page or a layout app, where the text itself is selectable — the file already stores that text along with its coordinates. This tool reads it, rebuilds lines and paragraphs from those coordinates, pulls out every figure together with its position, puts it back into the document as an image, and packages the result as a real .docx you can edit straight away.
If your PDF is a scan or a photo — the whole page is one bitmap and the "text" is only pixels — there is no text layer to read, so OCR is the only way through. This tool has OCR built in: it renders each scanned page to an image, recognises it into editable text, then rebuilds paragraphs and headings from the recognised positions. A page that recognises badly is not guessed at — it is kept as an image instead, so you never get a document full of invented words. Chinese, English, mixed Chinese/English and traditional Chinese are all available.
If the text exists only as an image — a screenshot pasted into Word, or a figure in a PDF that is really a picture of text — those images are recognised individually. The rule is one line: if it recognises confidently it becomes text, if it does not the image is kept as it was.
One thing worth knowing: everything happens inside your browser. Contracts, medical records and financial statements are exactly the files people refuse to upload to a stranger server — here they never leave your device, OCR included. You can disconnect from the internet and the tool still works.
A word on accuracy, honestly: printed text, sharp and at a decent resolution, recognises best. Handwriting, heavy skew, faded photocopies and text crossed by table rules will come out with errors and need proofreading. The result reports its average confidence, and when that is poor this tool keeps the image rather than pretending.
FAQ
Will the layout match the original exactly?
Not pixel for pixel, but paragraph structure, heading levels, centring and first-line indents are preserved, and figures go back in at their original place. A PDF has no concept of a "paragraph" — only fragments of text with coordinates — so paragraphs are inferred from typesetting patterns. Single-column reports, contracts and manuals come out best; two-column papers and complex tables will need a little tidying.
Are the figures redrawn?
No. Figures are the original image data lifted straight out of the document and re-inserted at their original position and size, so there is no loss of clarity and nothing gets redrawn or distorted.
How accurate is OCR on scans?
Best on printed text that is sharp and at a decent resolution. We tested a deliberately skewed 200 DPI Chinese scan: every key phrase was recognised correctly, mean confidence 92, about 1.1 seconds per page. Handwriting, heavy skew, faded photocopies and text crossed by table rules will have errors. The tool reports its confidence, and pages that recognise poorly are kept as images rather than guessed at.
Does the OCR run on a server?
No, it runs in your browser. The recognition engine (about 4MB) and the language data (about 2-3MB) are downloaded from this site the first time they are needed, then cached by your browser. Nothing about your document is sent anywhere for recognition, and the file never leaves your device. Because it all runs locally, speed depends on your machine — long documents will take longer.
Why is there a wait the first time?
Because the recognition engine and the language data have to be downloaded — about 10MB in total. It only happens when a scanned page or an image containing text is actually met, so pure digital PDFs and Word files download nothing extra. Once downloaded, your browser caches them and the next run starts almost instantly.
Is my file uploaded or stored?
No. Every step — parsing, recognition, layout rebuilding, image extraction and Word generation — runs on your own device, and the file never crosses the network. Close the tab and it is gone. You can disconnect from the internet and the tool still works.
How large a document can it handle?
Multi-page PDFs are fine — pages are processed in order with page breaks inserted. But memory use grows with length, so keep a single file under roughly 200 pages and 80MB. For very long scans, work in batches: OCR is noticeably slower than plain text extraction.