PDF to Text
Choose a PDF, click Extract text, then copy the text or download it as a .txt file — read straight from the PDF on your own device.
- Free
- No account
- Runs in your browser
- Nothing uploaded
Runs entirely in your browser — your input is never uploaded, logged, or stored.Privacy policy
What is PDF to Text?
PDF to Text extracts the text from a PDF and gives it to you as plain text you can copy or download as a .txt file. It is for quoting from a report, pasting a contract clause into an email, feeding a document into a spreadsheet or another tool, or simply getting the words out of a PDF that is awkward to select. The PDF is read in your browser and never uploaded.
It works on PDFs that contain real text — anything exported from Word, Google Docs, a web page or most accounting software. Inside such a file, text is stored as pieces of characters, each placed at a position on the page; there are no lines or paragraphs as such. The tool reads those pieces with pdf.js, Mozilla's PDF library, and rebuilds the reading order from their positions: pieces whose baselines are within half a font size of each other form one line, read left to right; a horizontal gap wider than about 0.15 of the font size becomes a space; and a vertical step more than 1.3 times the page's usual line spacing becomes a blank line, which marks a paragraph. Pages are separated by a blank line, and "Mark page breaks" adds a "--- Page N ---" line before each page. Pages stored sideways are read the way you see them.
A scanned PDF is different: each page is a photograph of paper, so there is no text to extract — only pixels. The tool detects pages with no text at all and says so, naming them. For those, save the pages as images with PDF to JPG and run them through Image to Text, which uses optical character recognition (OCR). This tool does no OCR, so it never guesses at words.
Layout is approximated, not reproduced. Text in two side-by-side columns, or in table cells on the same row, is read straight across, so the columns are interleaved line by line. Hyphenated words at line ends are kept as they are. Hebrew and Arabic are stored in a PDF in visual order, so each line is put back into reading order with a simplified version of the Unicode bidirectional algorithm: right-to-left words read right to left, while numbers and English words inside them stay left to right. Complex mixed-direction text (nested quotes, brackets) may still come out in the wrong order. Some PDFs store characters in an order or encoding that does not map back to readable text (certain custom fonts do this); the output then contains the wrong characters, and only OCR can help. Headers, footers and page numbers are included, because to the PDF they are text like any other.
You can choose all pages or a range such as 1-3, 5. The result shows the word and character counts, a Copy button and a Download .txt button; the file is saved as UTF-8 and named after the PDF. PDFs up to 200 MB are accepted; ones that need a password to open are refused with a clear message. Nothing is uploaded — you can confirm in your browser's developer tools (Network tab) that no request carries your file; the site loads its own scripts and analytics, never your documents.
1. The file size is checked before reading (200 MB limit) and its first bytes for the PDF signature; pdf.js then opens it. 2. For each chosen page, pdf.js getTextContent returns positioned text pieces; each is converted to the page as displayed (rotation applied, y measured from the top), with its font size from its transform. 3. Pieces are sorted top to bottom, and grouped into one line when their baselines differ by no more than half the smaller font size; each line is sorted left to right. 4. Within a line, a space is inserted where the gap between two pieces exceeds 0.15 × the font size (unless one already has a space). 5. Between lines, a step greater than 1.3 × the page's lower-quartile line spacing (or 1.8 × the font size when the page has fewer than four lines) is a paragraph break. 6. A line containing Hebrew or Arabic letters is converted from visual to reading order (right-to-left runs reversed; in a mostly right-to-left line, the run order reversed too). 7. Pages are joined with a blank line (optionally with "--- Page N ---" markers); pages with no text are listed separately.
Worked examples
- Two paragraphs with lines 14 pt apart and a 30 pt gap between them → "The first paragraph\ncontinues here.\n\nA second paragraph\nends."
- "Hello" ending at 102 pt and "world" starting at 105 pt in 12 pt type (a 3 pt gap, more than 0.15 × 12 = 1.8 pt) → "Hello world"; pieces 0.5 pt apart are joined without a space.
- A two-column page → each output line holds the left column's line followed by the right column's: the columns interleave.
- report.pdf, 3 pages, page 2 a scan → report.txt with the text of pages 1 and 3, and a note that page 2 has no text layer.
- A scanned PDF with no text on any page → "No text layer found", with links to PDF to JPG and Image to Text.
How to use PDF to Text
- Click "Choose a PDF" or drop a PDF onto the box.
- Keep Pages on All, or choose Range… and type pages like 1-3, 5. Tick "Mark page breaks" if you want each page labelled.
- Click "Extract text".
- Click "Copy text" to put it on the clipboard, or "Download .txt" to save <name>.txt.
Common errors
- "No text layer found" — the PDF is a scan or photo of paper, so it contains images, not text. Use PDF to JPG, then Image to Text (OCR).
- Columns are mixed together — text placed side by side is read straight across the page. Copy the columns separately from a PDF reader, or edit the result.
- The output is gibberish or wrong letters — the PDF's font does not map its characters back to readable text. OCR (PDF to JPG, then Image to Text) is the only way round it.
- "… is password-protected. Open it in a PDF reader with the password, save an unprotected copy, and use that." — the PDF needs a password to open.
- "Page 7 does not exist — this PDF has 6 pages." — the range contains a number higher than the page count shown next to the file name.
- "… is actually a JPG file, not PDF" — the file is an image with a .pdf name. To read text from an image, use Image to Text.
FAQ
How do I extract text from a PDF?
Choose the PDF here and click "Extract text". The text appears in a box with a Copy button and a Download .txt button, with lines and paragraphs rebuilt from the page layout.
How do I convert a PDF to a TXT file for free?
Extract the text with this tool and click "Download .txt". You get a UTF-8 plain-text file named after your PDF, with no sign-up and no watermark.
Why can't I copy text from my PDF?
Either the PDF is a scan, so its pages are images with no text, or its text is hard to select in your reader. This tool handles the second case; for a scan, use PDF to JPG and then Image to Text (OCR).
Does this work on scanned PDFs?
No — a scanned page has no text to extract, and this tool does no OCR. It tells you which pages have no text, and links to Image to Text, which can read them from images.
Is it safe to extract text from a confidential PDF here?
Yes. The PDF is read inside your browser and never uploaded. You can check in your browser's developer tools, under Network, that no request contains your file.
Will the formatting be kept?
Only as plain text: lines, paragraphs and page breaks are kept, but fonts, bold, tables and images are not. Side-by-side columns are interleaved line by line.
Can I extract text from only some pages?
Yes. Choose Range… under Pages and type the pages you want, such as 2-4, 7. Only those pages are read.
Does it work on iPhone and Android?
Yes. Open this page in your phone's browser, choose the PDF from your files and tap "Extract text", then copy the result or download the .txt file.
Related tools
- Split PDFChoose a PDF, type the pages you want, and extract them, delete them, or split the file into separate PDFs — all on your own device.
- Image to Text ConverterDrop, pick or paste an image or screenshot, click Extract text, then copy the recognised English text or download it as a .txt file.
- PDF to JPG ConverterPick a PDF, choose the pages and resolution, and download each page as a JPG image — rendered on your own device, never uploaded.
- Compress PDFChoose a PDF, pick lossless clean-up or strong compression, and see the exact before and after size — compressed on your device, never uploaded.
Prefer AllUtil on Google
One click adds AllUtil to your Google preferences. You'll see our tools highlighted with a Preferred badge in Search and AI answers.