Extract text from a PDF
Pull the text out of any PDF right in your browser. Fast for normal PDFs, with optional on-device OCR for scans. Nothing is uploaded.
Files are processed in your browser and never uploaded.
How to use Extract text from a PDF
Add your file
Drop your PDF file onto the upload area, or tap the area to choose it from your device.
Set the options
Choose Use OCR for scanned pages and Add page markers. Every setting is on screen before you add a file, so you can see what the tool will do before it runs.
Run it and save
Press Extract text, Don →. The button stays disabled until there is something to work on, and each file shows its own progress and result. Press Save .txt, Don. Each finished file keeps a name based on the one you started with, so you can tell the versions apart.
What this tool does
Extract text from PDFs in the browser
Any PDF is opened and its text pulled out, ready to copy, search or paste into something else.
Fast text-layer extraction
A normal PDF has a text layer, and reading that is near instant even on a long document.
Optional on-device OCR for scanned PDFs
A scan has no text layer, so OCR runs on your device to read the page as an image. It is slower, and it is optional for that reason.
Copy or download as a text file
Copy the result to the clipboard or download it as a .txt file, whichever suits what comes next.
Word and character count
Word and character counts are shown with the result, which is useful when you are extracting for a length limit.
No upload, no sign-up
The PDF is opened and read in your browser, and even the OCR runs locally, so nothing is uploaded.
Frequently asked questions
Are my PDFs uploaded anywhere?
No. The PDF is read and its text is pulled out inside your browser on your device, and nothing is ever sent to a server.
Does it work on scanned PDFs?
Normal digital PDFs (made from Word, Google Docs, print to PDF and so on) carry a real text layer, so their text comes out instantly. A scanned PDF is just an image of text with no text layer; turn on "Use OCR for scanned pages" and the text is read on your device. The OCR engine downloads once (a few MB) and is then cached, it is slower, and it is optimized for major languages, so accuracy is best on common scripts and clear scans.
Will the layout and formatting be kept?
It pulls out the words and line breaks, but extraction is essentially linear, so complex multi-column layouts, tables and heavy formatting can come out in an awkward reading order. You can clean up the result in the editable box before saving.
Which languages are supported?
Text-layer extraction works for any language already stored in the PDF. The optional OCR is optimized for major languages and uses English by default, so it is most accurate on common scripts and may struggle with unusual fonts or other writing systems.
Can it handle big PDFs?
Yes for the fast text-layer method. OCR is much heavier because every scanned page is analysed on your device, so on older phones or very long scanned documents it can be slow; work in smaller batches if needed.
Last updated: