Deriving PDF to HTML, version 1.2
PDF Association newsSeptember 27, 2026
PDF Association newsSeptember 27, 2026
About PDF Association staff
The Deriving PDF to HTML TWG has just published version 1.2 of its derivation algorithm.
Version 1.1 of this document was never published, and served only as an internal intermediate transition-point. Therefore this summary compares the published 1.0 version with the just-released 1.2 revision. It is intended to highlight changes in the document’s scope and direction, not serve as a formal or exhaustive release note.
High-level changes from 1.0 to 1.2
- Shifted the document from a baseline derivation specification to a Well-Tagged PDF / PDF 2.0-oriented usage specification.
- Broadened the scope from layout extraction to semantic conversion, with greater emphasis on structured authoring, namespace-aware processing, and processor behavior.
- Added more explicit handling of pagination and page semantics, including page breaks and page labels represented as HTML navigation for document pagination.
- Strengthened alignment with HTML and accessibility standards, including ARIA roles and attributes, form semantics, and interoperable element mapping.
- Increased the detail of processor requirements and authoring expectations for producing reliable derived HTML from tagged PDF.
Expanded the coverage of content patterns such as headings, lists, links, forms, labels, notes, and interactive elements. - Increased attention to reuse, interoperability, and accessibility outcomes for mobile and HTML-based consumption.
- Clarified the accessibility boundary: derived HTML can carry semantic and ARIA meaning, but it does not automatically equal the accessibility characteristics of the original PDF.
- Refined the model around semantic structure, role mapping, and valid HTML output instead of visual reconstruction alone.
- Improved practical guidance for implementers, especially around deterministic processing, valid HTML, and predictable output behavior.
Improvements in version 1.2
- Formalized the derivation model around canonical PDF structure semantics and HTML output.
- Expanded the treatment of structure semantics and the relationship between PDF structure elements and their derived HTML counterparts.
- Added clearer examples and more explicit rules for common content patterns and edge cases.
- Improved guidance for form field processing, HTML namespace handling, and role-mapping behavior.
- Strengthened the treatment of accessibility semantics and how HTML and ARIA metadata should support meaningful output.
- Improved handling of metadata, document structure, and output conventions such as `lang`, title, viewport metadata, and pagination.
- Better preserved interactive and UI-like semantics in a way that remains meaningful in HTML.
- Extended the coverage of special cases, including headings, labels, lists, links, notes, and form-related structures.
- Improved the editorial structure and made the algorithm narrative clearer and more consistent than the earlier 1.0 baseline.
Historical context and short version
The original work on a derivation-to-HTML specification predates the formal Well-Tagged PDF (WTPDF) specification. In practice, the idea of deriving structured HTML from Tagged PDF existed before WTPDF became the reference model for authoring and processing Tagged PDF.
This matters because the initial release of the algorithm represented practical guidance written before PDF's ecosystem had a single, fully formalized interpretation of Well-Tagged PDF. At the same time, the broader direction of the work is aligned with the industry's progress toward PDF/UA-2 compliance, structured reusability, and conversion workflows that originate from authoring pipelines such as LaTeX and other structured document sources. In that sense, the derivation effort is part of the same ecosystem shift toward machine-readable, reusable, standards-based PDF content.
Version 1.2 is the first version of this algorithm that is fully aligned with the new model. The new document also explicitly addresses math semantics, including MathML and formula handling, which is a key part of reusable conversion, especially for derivation that's oriented towards accessibility needs.
In short: 1.0 was the original baseline derivation specification. The unreleased version 1.1 moved the work toward a more modern and standards-aware interpretation. The now-released version 1.2 significantly expands and clarifies the document by making the derivation logic more implementation-oriented, and more aligned with PDF 2.0, Well-Tagged PDF authoring, and HTML interoperability.
Going forward
HTML derivation is unique because it explicitly defines processor behavior—something the core PDF specification rarely addresses. To support processor developers, our upcoming work focuses on establishing clear guidance to ensure predictable content extraction and seamless data reuse.
The immediate priority is an Application Note for authors on creating PDFs designed for full reusability and repurposing in HTML environments. Subsequent work will expand into specifying a dedicated PDF reflow mode powered by HTML derivation.
Download Deriving HTML from PDF 1.2 (PDF) today!



