An EPUB File Is a Small Website in a ZIP Archive. Here Is What That Tells You About Ebook Conversion.

ToolHQ TeamOctober 8, 20268 min read

Open an EPUB file with an archive viewer and you will find HTML documents, CSS stylesheets, images, and an XML metadata file called the package document. The content itself is in XHTML, the same format as web pages. The styling is in CSS, the same language web browsers use. The whole package is wrapped in a ZIP container with a special first file that identifies it as an EPUB publication. The W3C, which also maintains the HTML specification, publishes the EPUB standard precisely because the two formats share the same foundation.

That technical reality changes how you think about what ebook conversion actually is. Converting an EPUB file to a MOBI file is not transforming an audio track into a different audio track, where the underlying data is fundamentally the same. It is converting a specific flavor of HTML document packaging into a different flavor of HTML document packaging. The text content stays the same. The presentation layer is re-expressed in a different container format. When both source and destination formats are HTML-based, conversion is relatively clean. When one of them is PDF, the problem is entirely different.

The ebook format landscape that exists today was not designed. It accumulated. Every major reader manufacturer and platform made independent decisions over fifteen years, and the result is that the same book can exist in multiple incompatible formats, tied to different platforms and devices. Understanding how each format originated explains why conversion between some pairs is straightforward and conversion between others loses structure.

The Kindle Format Family Tree

When Amazon acquired Mobipocket SA in 2005, it inherited both a format and an audience. Mobipocket had created its MOBI format in 2000, designed for PalmPilot-era handheld devices with 160-by-160-pixel screens and 8 MHz processors. The format used a compressed subset of HTML for content, stored in a binary container built for the constraints of late-1990s hardware.

Amazon launched the first Kindle in November 2007 using a format called AZW. AZW was the Mobipocket format with Amazon's DRM wrapper applied. A DRM-free AZW file is byte-for-byte identical to a MOBI file except for the file extension and a flag in the header. Amazon had not reinvented the format; it had licensed and locked it.

That strategy worked for the first generation of Kindle devices, but the format's origins in 1990s hardware began to show. As web content grew richer, the HTML subset in MOBI could not support CSS layouts, embedded fonts, or the formatting that textbooks and illustrated titles required. In November 2011, Amazon launched the Kindle Fire tablet alongside Kindle Format 8, known as AZW3 or KF8. AZW3 was a complete redesign based on HTML5 and CSS3, compatible with the formats modern web browsers used.

For the next decade, new Kindle devices supported AZW3 while older devices continued reading MOBI. Authors publishing to Kindle needed to produce content in Amazon's proprietary format. Ebooks purchased outside the Kindle ecosystem could not be read on Kindles without conversion, and Kindle purchases could not be read in EPUB-based apps. In 2022, Amazon reversed course. The Send to Kindle service dropped support for its own Kindle File Format in favor of EPUB. New Kindle Paperwhite and Kindle Scribe devices gained native EPUB support. After seventeen years of operating a parallel ecosystem, Amazon's reading platform began accepting the open standard the rest of the industry had used all along.

The Reflowable and the Fixed

The most significant conversion challenge in the ebook world is not between EPUB and MOBI. It is between PDF and anything reflowable.

A PDF document positions every element at specific coordinates on a page. Text appears at a known location because the PDF specification stores x and y coordinates, font sizes, and layout information for each text fragment. The document looks the same at any screen size because the layout is fixed in the file. Zoom in on a PDF and the text gets larger; the page does not reflow to fill the screen.

EPUB, MOBI, and AZW3 are all reflowable formats. Text fills the available width of the screen, whatever that width happens to be. Change the font size on your e-reader and the paragraphs reflow automatically. A book formatted for EPUB renders reasonably on a pocket phone screen and on a large tablet without modification, because the format describes content and structure, not layout coordinates.

Converting a PDF to EPUB requires the conversion software to infer information the PDF does not explicitly contain. Where one paragraph ends and another begins must be determined from proximity and formatting. Reading order must be deduced from column layouts and page flow. Headers must be identified from font size differences. Images must be separated from the text that flows around them. All of this inference is imperfect, and the quality of a PDF-to-EPUB conversion depends heavily on how well structured the source PDF is.

Consider Rafael, a graduate student in philosophy who has accumulated hundreds of academic paper PDFs. He reads on a Kobo e-reader, which handles EPUB well but struggles with PDFs from journals that use two-column layouts, small fonts, and dense footnotes. Converting those PDFs to EPUB works well for papers from certain publishers where the PDF was generated from structured source files: the text extracts cleanly and footnotes appear in a sensible location. Papers scanned from printed journals produce messier conversions, because the PDF contains no text, only images of pages, and the conversion must first run optical character recognition before it can attempt reflowing.

What Happens During Format Conversion

When you convert between EPUB and MOBI or between EPUB and AZW3, the process is relatively straightforward. The conversion tool reads the HTML content and CSS from the EPUB container, applies any necessary transformations to accommodate format-specific features or limitations, and writes the content into the target format's container structure. Metadata like title, author, and table of contents transfers across formats. Embedded images transfer. Formatting is reproduced as closely as the target format's CSS support allows.

Converting in the opposite direction, from Kindle formats to EPUB, follows the same logic. The AZW3 HTML content is read, potentially with some style adjustments for EPUB compatibility, and repackaged in an EPUB-conformant ZIP container with the appropriate metadata files. The text and images are the same; only the envelope has changed.

PDF-to-EPUB conversion is the outlier. It involves text extraction, reading order inference, paragraph reconstruction, and optional OCR if the PDF contains scanned pages rather than actual text. The result is functional EPUB but may require manual cleanup if the source document had complex layout, footnotes, multi-column text, or tables.

When to Convert and What to Expect

EPUB to MOBI is the conversion most Kindle users need when they have ebooks from libraries or other sources that do not sell in Kindle format. The conversion preserves text, images, and basic formatting. Complex CSS layouts may simplify during the process, but standard prose converts cleanly.

MOBI or AZW to EPUB is what users need when moving ebooks from the Kindle ecosystem to another reader. Since the underlying content is HTML-based in both formats, the conversion is reliable.

PDF to EPUB is appropriate for PDFs that were generated from word processors or layout applications and contain actual text, not scanned images. The conversion quality is highest for simple single-column layouts and degrades with complexity. For scanned PDFs, a separate OCR step before conversion produces better results.

When you convert ebooks using ToolHQ's ebook converter, the file is processed securely on the server and deleted immediately after you download the result. No copy is retained. For any ebook that contains personal documents or purchased content, knowing that the file is processed and discarded rather than stored matters.

Conclusion

The ebook format fragmentation that required converters to exist in the first place is largely a product of platform competition. Amazon built a hardware business on Kindle and needed to differentiate its ecosystem. The EPUB standard existed from 2007 but Amazon had no incentive to adopt it while Kindle dominated the market.

Amazon's 2022 shift to EPUB support signals that the format competition is effectively over: EPUB has become the baseline for both open-standard readers and the largest proprietary platform. Conversion tools remain necessary for the years of existing content in older formats and for users who still own pre-2022 Kindle devices. The underlying technical reason is simple: EPUB is HTML in a ZIP file, and HTML won the document format war the same way it won the web.

Frequently Asked Questions

Is an EPUB file actually a ZIP archive?

Yes. You can rename any .epub file to .zip and open it to see the contents: HTML documents, CSS stylesheets, images, and XML metadata files. The W3C's EPUB specification explicitly builds on the same HTML and CSS standards used for web pages.

Why did Amazon use MOBI instead of EPUB for the Kindle?

Amazon acquired Mobipocket in 2005, which gave them the MOBI format and an existing ebook reader ecosystem. When the first Kindle launched in 2007, Amazon used a DRM-wrapped version of MOBI called AZW rather than adopting the open EPUB standard, keeping its ebook library exclusive to Kindle devices.

Why does PDF to EPUB conversion sometimes produce poor results?

PDF stores exact coordinates for every text element. Converting to EPUB requires inferring paragraph structure, reading order, and document hierarchy from those coordinates, which works well for simple layouts but fails on two-column academic papers, scanned documents, or complex formatted tables.

Try These Free Tools