Skip to content
toolsdocks

Extract text from a PDF

Pull the embedded text out of a PDF, page by page.

Runs on your device
Loading tool…

How to use

  1. Add one or more PDFs.
  2. Choose plain text or Markdown, and whether to keep line breaks, join hyphenated words and add page markers.
  3. Copy or download the text.

Worked example

The 12-page sample report becomes a TXT file with “— Page 1 —” separators and its wrapped paragraphs re-joined; as Markdown its 24 pt title becomes “#”, the 16 pt section headings “##” and the bulleted key points a list.

Supported formats and limits

InputPDF
OutputTXT, Markdown
LimitsUp to 200 MB.
Enginepdf.js text content with line, paragraph and column reconstruction

Limitations

  • Pages that are images contain no text layer; those pages are listed and can be sent to OCR.
  • Two-column pages are detected and read column by column; tables, text boxes and more complex layouts may come out in an unexpected order.
  • Headings are guessed from font size and bold text, so they are not always right.

Questions

Why is some of my PDF missing from the text?

Pages that are images, such as scans, have no text layer. They are listed so you can send them to the OCR tool.

Will the reading order be right?

Usually for single-column pages and simple two-column layouts, which are read column by column. Tables, text boxes and complex layouts can come out in an unexpected order.

How are headings found for Markdown?

They are guessed from font size and bold text. In the sample report the 24 pt title becomes a level 1 heading and 16 pt section headings level 2, but the guess is not always right.

Guides

Privacy

Runs on your device. Files and text are processed in this browser tab and are not uploaded.

See the privacy policy for how toolsdocks handles data.