Trying to copy a sentence out of a PDF fails often enough to be familiar: the selection will not drag, the pasted result comes back full of stray line breaks, or nothing highlights at all. This tool pulls the text out of a PDF so you can copy it or save it as a .txt file.
What makes this awkward is that a PDF has no concept of a line. Inside the file there are only fragments of text scattered at coordinates. This tool treats a change in vertical position as a line break and reassembles something a human can read. The document is parsed in your browser and never uploaded.
How to use
- Add a PDF — Select a file or drag it in and the full text is extracted straight away. Longer documents take a moment.
- Review the result — By default a
--- Page 1 ---separator marks where each page begins so you can tell the content apart. The total character count appears underneath. - Re-extract specific pages — Under More, enter a range like
1-3, 5and press Apply to extract only those pages. Turn off Mark page boundaries if you want continuous prose to paste elsewhere. - Copy or save — Use the copy button to put it on the clipboard, or download to get a
.txtfile named after the original PDF.
Frequently asked questions
No text comes out at all.
It is almost certainly a scanned PDF. A document produced by photographing or scanning paper holds a single image per page and no character data — your eye sees letters, but the file contains only pixels.
Getting text out of that requires OCR, which is hard to do accurately in a browser alone. Check whether you can drag-select text in a PDF viewer first; if you cannot, this tool cannot extract it either.
The line breaks come out strange.
This is the fundamental limitation of a format with no line concept. The tool treats a change in the vertical coordinate of text fragments as a new line.
That works well most of the time, but two-column papers, tables and documents with footnotes can interleave columns or flatten table cells onto one line. For those, narrowing the page range and extracting in small pieces, then tidying by hand, is usually quicker.
The characters come out garbled.
The PDF either does not embed its fonts or encodes them unusually. If the file carries the information needed to draw glyphs but not the mapping that says which character each glyph represents, extraction produces nonsense.
This shows up occasionally in PDFs from older software or ones that embed only a font subset. Short of obtaining the source document and exporting it again, there is not much to be done.
Does it extract table contents?
The text comes through, but the table structure does not. A PDF has no notion of a table — only lines and characters sitting at their own coordinates.
The result is cell contents listed in sequence. If you need the data back as a table, arrange the extracted text into CSV and hand it to the CSV ↔ JSON converter.
Concepts worth knowing
How a PDF holds text
Text inside a PDF is not organised into paragraphs the way a word processor would. It is a sequence of operators like Tj saying: draw this string, in this font, at this position. Paragraphs, lines and tables are things a reader infers, not things written in the file.
So every PDF text extractor reconstructs structure by reading coordinates. Cleanly produced documents reconstruct well; heavily formatted ones come apart. Recent PDF standards do include tagging for structural information, but relatively few documents actually carry it.
PDFs marked as copy-protected
PDF supports permission flags restricting printing, copying and editing. These are not encryption — they are closer to a request that compliant viewers voluntarily honour. The content itself sits in the file in the clear.
That is why text often extracts fine from documents marked no-copy: nothing is technically blocking it. The flag still expresses the rights holder's intent, though, so what you do with the extracted text is a separate judgement from whether you can.
Cleaning up extracted text
A few passes make the output substantially more usable. Start by deleting repeated headers, footers and page numbers — a single search-and-replace usually clears them.
Next, rejoin lines broken mid-sentence. The original layout's line breaks survive extraction and leave sentences chopped up; replacing any line break not preceded by a full stop with a space fixes most of it. Finally collapse runs of blank lines. The dedupe/sort tool and the regex tester on this site make short work of all three.