What does PDF to Word Converter do?
Convert PDF to an editable Word .docx in your browser. Rebuilds paragraphs, headings and bold text from the PDF text layer. No uploads, no sign up.
- Price
- Free
- Account
- Not required
- Processing
- Entirely in your browser; your data and files are not uploaded
- Works on
- Any modern browser on desktop, tablet or phone
- Category
- PDF Tools
How to use the pdf to word converter
- Drop a PDF onto the drop zone or click to choose one. If it is password protected, enter the password to open it locally.
- Optionally type the pages to convert, for example 1-3, 5. Leave it empty for every page.
- Choose whether each PDF page should start on a new page in Word.
- Click Convert to Word, check the word count and the preview of the first paragraphs, then download the .docx file.
Worked example
A two-page report with a title and section headings
A report whose body text is 10 pt, with a 20 pt title and 13 pt section headings, converts to a Word file where the title uses the Heading 1 style, the sections use Heading 2 and wrapped body lines are joined back into paragraphs. With "Keep page breaks" on, page 2 starts on a new page in Word. A scanned PDF of the same report has no text layer, so the tool reports that nothing can be extracted.
How it works
Text and positions are read from each page with PDF.js getTextContent and mapped to top-down page coordinates. Rotated text (such as diagonal watermarks) is skipped. Items whose baselines are within half the font size are grouped into a line and sorted left to right, with a space added where the gap is wider than 15% of the font size. Lines join into a paragraph unless the vertical gap is more than 1.4 times the usual line spacing, the font size changes by more than 15%, the line starts with a list marker, is indented, or follows a short line ending in punctuation. The body size is the most common font size by character count; paragraphs at 1.5 times that size or more become Heading 1, at 1.18 times or more Heading 2, and short fully bold lines also become Heading 2. Bold is detected from font names such as Arial-BoldMT. The .docx is written as standard Office Open XML (document, styles, relationships and properties) and zipped in your browser.
Assumptions
- The PDF has a text layer. Scanned pages are images and contain no text to extract; this tool does not do OCR.
- Text is read in rows from top to bottom. Multi-column layouts are read across the columns, so paragraphs from side-by-side columns can be mixed.
- Headings are inferred from font size and weight, so a document that uses the same size for everything will have few or no headings.
Frequently asked questions
Are my files uploaded?
No. Your PDF is processed in your browser on your own device and is never sent to our server. You can disconnect from the internet after the page loads and the tool still works.
Will the Word file look exactly like the PDF?
No. This is a text and structure conversion: you get editable paragraphs, headings and bold text in a clean Word layout. Images, tables, columns, exact fonts, colours, headers and footers are not reproduced. For an exact visual copy, keep the PDF or use PDF to JPG.
Why does it say no text was found?
The PDF is most likely a scan or photos of pages. The words are part of a picture, so there is no text to extract. Reading text from images needs OCR software, which this tool does not include.
Can I convert a password protected PDF?
Yes, if you know the password. It is used only inside your browser to open the file. Without the password the file cannot be read.
Does it work with Google Docs and LibreOffice?
Yes. The output is a standard .docx file that Microsoft Word, Google Docs, LibreOffice Writer and Apple Pages can open.
Limitations
- Files up to 100 MB are accepted, but very large or complex PDFs depend on your device memory and can be slow on phones.
- Images, tables, columns, text boxes, exact fonts, colours, headers, footers and footnote positions are not reproduced.
- Scanned or image-only PDFs produce no text because OCR is not supported.
- Words hyphenated across lines are rejoined when the next line starts with a lowercase letter, which can occasionally remove a real hyphen.
- Right-to-left and vertical scripts may come out in the wrong order.