Home › Guides › EPUB to text

How to get the text out of an EPUB

An EPUB is a zip of web pages. Extracting the words is straightforward once you know that.

An EPUB is an archive containing HTML files, stylesheets and images. That makes extracting the text tractable — it is already text inside, wrapped in markup.

Step-by-step

  1. Add your EPUB.
  2. Extract. Chapters come out in reading order.
  3. Copy or download the plain text.

What comes out

The readable text, in order, with markup removed. Paragraph breaks are kept because they carry meaning; styling is not, because plain text has no way to express it.

Images do not survive — plain text cannot hold them. Tables become their contents in sequence, which is readable but loses the arrangement.

Uses this is good for

DRM

Books bought from most shops carry digital rights management, and a DRM-protected EPUB cannot be read by this tool — the contents are encrypted, and removing that protection is a separate matter with legal implications that vary by country. This works with unprotected EPUBs: public domain texts, books from DRM-free publishers, and files you produced yourself.

Reading order is not file order

An EPUB is a zip archive of XHTML files, and the order they appear in the archive has nothing to do with the order they are meant to be read in. That is held separately, in the spine — a list inside the package file naming each document in sequence. A tool that simply walks the archive can produce chapters out of order, or a book that starts with the copyright page.

It also explains a common surprise: files present in the archive but absent from the spine are not part of the reading order at all. Publishers leave behind alternative covers, unused sections and stray drafts. Following the spine gets the book; unpacking the zip gets the book plus whatever else was in the folder.

What flattening to text costs

Plain text has one dimension, and a book has several, so some things necessarily collapse. Footnotes and endnotes lose their anchoring — they either land inline, interrupting the sentence that referenced them, or gather at the end detached from what they annotate. Images disappear entirely, taking any caption's meaning with them. Tables become runs of words whose column structure is gone.

Emphasis goes too, which matters more than it sounds: a book that distinguishes speech, quotation or a foreign term by italics alone loses that distinction completely. For search, word counts, reading on a plain device or feeding text to another tool, none of this is a problem. For anything where the formatting was carrying meaning, extract the text and keep the original alongside it.

Frequently asked questions

Is my book uploaded?

No. The file is unpacked in your browser.

Why will my purchased book not open?

It is almost certainly DRM-protected, which means the contents are encrypted. This tool reads unprotected EPUB files.

Are chapters kept in order?

Yes. The EPUB's own reading order is followed, so the text comes out as the book is meant to be read.

Open the text tools →