PDF Association logo

Discover pdfa.org

Key resources

Get involved

How do you find the right PDF technology vendor?
Use the Solution Agent to ask the entire PDF communuity!
The PDF Association celebrates its members’ public statements
of support
for ISO-standardized PDF technology.

Member Area

Diagram showing a PDF that declares PDF/A-2u conformance but, after save operations (compress/merge/split/convert), loses its structure tree so metadata remains while accessibility info is gone — conformance depends on the last tool that wrote it.

Every save rewrites a PDF. What usually survives an ordinary compress or merge is the XMP declaration, not the structure it describes: the file still claims PDF/A conformance it no longer has.

One structure layer—semantic document markup that maps a page’s visual layout to a structured document—enables both human screen readers and machine readers (LLMs/retrieval) to interpret the content.

For twenty years, the logical structure inside a tagged PDF served one audience: assistive technology. Then the main reader of the world’s PDFs stopped being human, and it needs exactly what screen readers always did.

Factur-X – One file, two native readers

HMRC has confirmed the UK’s 2029 e-invoicing mandate will run on Peppol, with PDFs excluded by definition. Weeks from now, France’s mandate goes live with a hybrid PDF/A-3 format as a first-class citizen. Same term, two philosophies.

A document claiming conformance to PDF/A’s conformance level “b” can be visually perfect and machine-unreadable at the same time. This is not a bug — it is what the standard was designed to do. Understanding the difference between rendering a glyph and encoding a character is the first step to building archives that AI can actually use.

Header illustration for the PDF Smart whitepaper "From Flat Files to AI-Ready Assets: The Strategic Role of OCR and Standardised PDFs in Enterprise Data Lakes."

Most enterprise AI initiatives are failing not because of the model — but because of the document. This whitepaper examines why flat, image-based PDFs render corporate archives invisible to RAG pipelines and LLMs, and makes the operational case for cloud-native OCR, PDF/A-2u standardisation, and zero-trust document architecture as the three non-negotiable preconditions for an AI-ready data lake.

WordPress Cookie Notice by Real Cookie Banner