PDF Association logo

Discover pdfa.org

Key resources

Get involved

How do you find the right PDF technology vendor?
Use the Solution Agent to ask the entire PDF communuity!
The PDF Association celebrates its members’ public statements
of support
for ISO-standardized PDF technology.

Member Area

Running Linux in PDF 🤪

First, video games. Then AI. Now Linux in PDF? | University of Illinois Library wins grant to develop EA-PDF implementations | Update to European Accessibility Act rules in the EN 301 549 | WCAG Evaluation Methodology (WCAG-EM) 2.0 | PDF Association comments on proposed SEC reporting rules | World Digital Preservation Day – 5 Nov 2026 | E-invoicing in France | Analyzing PDFs for AI prompt injections | More sponsored ISO standards | Google glitches the Matrix | Enjoy cartoons about PDF | PDFacade … Read more

PDF in the WildSeptember 28, 2026
screenshot of Linux in PDF
Running Linux in PDF 🤪
screenshot of Linux in PDF

First, video games. Then AI. Now Linux in PDF? | University of Illinois Library wins grant to develop EA-PDF implementations | Update to European Accessibility Act rules in the EN 301 549 | WCAG Evaluation Methodology (WCAG-EM) 2.0 | PDF Association comments on proposed SEC reporting rules | World Digital Preservation Day – 5 Nov 2026 | E-invoicing in France | Analyzing PDFs for AI prompt injections | More sponsored ISO standards | Google glitches the Matrix | Enjoy cartoons about PDF | PDFacade … Read more

PDF in the WildSeptember 28, 2026

PDF Association staff

About PDF Association staff


First, video games. Then AI. Now Linux in PDF?

Over the years we’ve reported on any number of crazy things one can accomplish inside a PDF with specific implementations; whether they are a good idea is another question entirely!

In 2024, for example, we noted that well-sedated Doom players (we can’t think of any) could now get their fix from within a PDF. Our PDFacademicBot this month is also reporting on Brazilian researchers who have built a reusable PDF game engine called “PDFGameBuilder”.

More recently, we learned that an over-caffeinated developer had the idea of operating an LLM in a PDF context (but why!?)

Running Linux on a PDF substrate is something we never thought we’d see. What’s next, z/OS in PDF?

University of Illinois Library wins grant to develop EA-PDF implementations

To support a new grant awarded to our friends at the University of Illinois, the PDF Association’s EA-PDF LWG will soon return to an active meeting schedule to consider updating the spec to review the latest updates to PDF/A-4 and to address various requests raised by implementers of EA-PDF 1.0.

Update to European Accessibility Act rules in the EN 301 549

In 2019, the PDF Association published a standard means of announcing conformance with various standards via XMP metadata. Thanks to this news regarding EN 301 549, it’s time to update your PDF Declarations on WCAG compliance to WCAG 2.2!

As an aside, we must note that it’s unfortunate that the latest official PDF of EN 301 549, while a Tagged PDF, wasn’t made to conform to PDF/UA, the ISO standard for fully accessible PDF.

WCAG Evaluation Methodology (WCAG-EM) 2.0

For those working on accessible PDF a new document from the W3C titled “WCAG Evaluation Methodology (WCAG-EM) 2.0” describes a step-by-step process on how to evaluate how well digital products conform to Web Content Accessibility Guidelines (WCAG) 2.

Although the document explicitly mentions PDF as an applicable digital product, it fails to mention PDF/UA (ISO 14289, available at no-cost) or the PDF Association’s WTPDF specification. You can post feedback on WCAG-EM to their GitHub repository.

PDF Association comments on proposed SEC reporting rules

U.S. Securities and Exchange Commission logo

The US Securities and Exchange commission (SEC), which regulates financial markets in the US, has issued a proposed Rule regarding conditions under which the entities it regulates may deliver specific types of information to recipients electronically without first obtaining their affirmative consent.

PDF is an accepted deliverable format, but the Rule could (and in our view, should) go much further to require entities to distribute documents that meet PDF/A and PDF/UA requirements.

Read our formal comment on the proposed Rule.

World Digital Preservation Day – 5 Nov 2026

World Digital Preservation Day (WDPD) is held annually on the first Thursday of November. In 2026, this falls on Thursday 5 November with the theme “Save our Stories (SOS): Our digital heritage in danger”.

This event is a great opportunity to promote the ISO 19005 family of PDF/A standards, as well as the many applications of PDF/A, including ZUGFeRD and Factur-X for e-invoicing, Order-X for electronic purchase orders, and our own EA-PDF for the long-term preservation of email.

E-invoicing in France

Factur-X logo

Factur-X is a PDF/A-3 and PDF/A-4 based hybrid invoice that is both human-readable and machine-readable (thanks to embedded XML). It opens in any PDF viewer, and thereby supports sole traders and small businesses that can't afford expensive, complex EDI solutions. These same PDF invoices can also be automatically processed by any suitable accounting software without involving a human.

This Thomson Reuters blog post confuses the message with: “Once the mandate takes effect, an emailed PDF won't meet the e-invoicing obligation..

What they meant to say is that a PDF that does not conform to the Factur-X specification won’t meet the obligations, as the same post clearly states:

Factur-X is a common choice among SMEs and micro-enterprises. This format provides compliance with little to no changes to daily work habits. Factur-X acts as a hybrid option that still lets the client read and download the e-invoice as a PDF if their workflow requires it, while also meeting machine-readable government standards.

Analyzing PDFs for AI prompt injections

A developer created a tool to find hidden prompt injections in PDF files, describing this as an arms race!

Ironically, the developer found that AI wasn’t useful in detecting prompt injection. It was enough to simply detect white text, tiny text, or text located off the visible page, all of which are relatively straightforward using conventional format analysis. As experts in PDF, we note that there are plenty of other tricks that could be used to hide prompt injections.

More sponsored ISO standards

In 2023, the PDF Association finalized an agreement with ISO and ANSI for sponsored, global no-cost access to the core ISO publications defining PDF 2.0 - restoring the 30-year tradition of providing no-cost access to the latest PDF specification. In 2025, a similar agreement for sponsored no-cost access to ISO publications related to PDF/UA was also established. While we might dispute the landmark and uniqueness claims in ISO’s recent announcement of the latest set of ISO standards to be sponsored, there is clearly a growing trend for no-cost access to important ISO technology standards.

Google glitches the Matrix

Search-engine specialists recently noticed that Google searches results were kind of, um, missing PDFs!

A few days later the glitch seems to be resolved… but we’re checking the walls to be sure they haven’t been moved.

Overleaf webinar promotes accessible PDFs for STEM

This informative webinar recording can be accessed here.

Enjoy cartoons about PDF

Our culture desk went wild for this! Check out this collection of PDF-related cartoons.

PDFacademicBot for September 2026

The PDFacademicBot brings academic research on PDF and related technologies to the industry’s attention.

Akash, M.S. (Aug. 2026) Will the Future Journal Even Be a Paper? AI, Interactive Code, and the Next Generation of Research Dissemination. https://www.researchgate.net/publication/413456930_Will_the_Future_Journal_Even_Be_a_Paper_AI_Interactive_Code_and_the_Next_Generation_of_Research_Dissemination

Peter’s comment: This paper recognizes some core foundational issues associated with traditional scholarly publications, including aspects that we’ve written about for several years, when reviewing the usual tripe of PDF haters proposing yet another file format as silver bullets, while missing the bigger picture:

“The analysis indicates that the strongest future model is not a simple replacement of PDF by software, but a layered research object in which narrative, code, data, metadata, provenance, and version history are linked and independently preservable. The paper argues that successful transition requires coordinated reforms in publication infrastructure, integrity verification, digital preservation, and research assessment. The future journal is therefore best understood not as the disappearance of the paper, but as the emergence of a verifiable digital research record in which a conventional article is only one interface to a larger scholarly object.

The problem is narrower and more consequential: the PDF is often treated as the complete research object when the underlying  evidence is computational, interactive, or continuously revised. …  A PDF can remain readable for decades with comparatively simple infrastructure.”

PDF’s existing feature set provides many, but not all, of these capabilities if only academic publishers would update their systems. Issues around “bit rot” (of URLs, data sets and software) and trust transcend all file formats. But solutions for authenticity, provenance, embedded data, rich metadata, MathML or other forms of rich semantics already exist with today’s PDF.

More on PDF…

Abramov, S. (Aug. 2026) “AcroMELD: Recovering Interactive PDF Forms with Structure-Aware Graph Set Transformers.” arXiv. https://doi.org/10.48550/arXiv.2608.22338.

Allu, U. et al. (2026) “Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion.” arXiv. https://doi.org/10.48550/arXiv.2609.24220.

Anjaneyulu, P., Priyanka, D. and Manaswini, Y. (2026) “A Hybrid Transformer–CNN–BiLSTM Framework with Comparative Benchmarking for Automated Speech-to-Structured PDF Document Generation,” pp. 10. https://journal.uob.edu.bh/server/api/core/bitstreams/9932e9db-fdb8-41f6-a109-0691410cf06b/content

Bhuvanesh, V. et al. (2026) “Inclusive Offline Multimodal Retrieval-Augmented Generation System for Accessible PDF-Based Knowledge Assistance,” International Journal of Research and Innovation in Applied Science, 11(7), pp. 1304–1321. https://ideas.repec.org/a/bjf/journl/v11y2026i7p1304-1321.html

Brown, A. et al. (Aug. 2026) “DocLayout-MM-RAG: A Layout-Aware Annotation Framework for Grounded Question Answering over Documents,” Proceedings of the 2026 ACM Symposium on Document Engineering. New York, NY, USA: Association for Computing Machinery (DocEng ’26). https://doi.org/10.1145/3820755.3821479.

Doi, N. and Tanaka, M. (July 2026) “Zero-Shot Structured Data Extraction from Timely Disclosure Documents in PDF Format with Large Language Models,” 2026 20th IIAI International Congress on Advanced Applied Informatics (IIAI-AAI), pp. 790–797. https://doi.org/10.23919/IIAI-AAICPS00095.2026.00148.

Faria, J.P. et al. (Sept. 2026) “From PDFs to Places: A Semi-Automated, Auditable Workflow to Map Study Localities for Geoheritage Inventories,” Geoheritage, 18(4), p. 222. Available at: https://doi.org/10.1007/s12371-026-01450-z.

Foppiano, L. (Aug. 2026) “Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics.” arXiv. https://doi.org/10.48550/arXiv.2608.16390.

The above article analyses the SafeDocs large-scale PDF corpus we helped create from an AI perspective by examining the distribution of text tokens in PDFs. We certainly agree with their conclusion about improving reporting effectiveness of AI systems: “corpus statistics [should] be reported in both units: documents and tokens.”

Foppiano, L., Khamassi, S. and Gupta, V. (Sept. 2026) “Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion.” arXiv. https://doi.org/10.48550/arXiv.2609.26381.

Freeman, A. and DuPre, L. (May 2026) Managing Massive PDF Remediation Projects with AI Support. University of South Carolina (2026 Southeastern Conference (SEC) AI Library Summit). Video (8 minutes). https://scholarsjunction.msstate.edu/sec-ai-2026/8.

Gadbail, H.N. et al. (Aug. 2026) “A Multi-Provider, Client-Configurable Web Application for Multi-Granularity LLM-Driven PDF Summarization,” 2026 7th International Conference On Computational Vision and Bio Inspired Computing (ICCVBIC), pp. 1423–1427. https://doi.org/10.1109/ICCVBIC71195.2026.11689471.

Jain, J., Sinojia, Y.V. and Joshi, S. (August 2026) “Layout-Aware Document Parsing and Knowledge Graph Extraction,” 2026 10th International Conference on Inventive Systems and Control (ICISC), pp. 1228–1233. https://doi.org/10.1109/ICISC69558.2026.11681334.

Jain, M., Wefel, S. and Schneider, P.H. (Aug. 2026) “Hiding in Plain Documents: PDF Steganography Techniques, Detection, and Criminal-Use Implications,” Availability, Reliability and Security. ARES 2026 International Workshops, Sweden, August 24–27, 2026, Proceedings, Part III. Berlin, Heidelberg: Springer-Verlag, pp. 183–200. https://doi.org/10.1007/978-3-032-35586-7_11.

Khadija, M.A. et al. (2026) “OCR-integrated retrieval-augmented generation for PDF- document question answering system in vocational higher education,” Journal of Soft Computing Exploration [Preprint]. https://joscex.shmpublisher.com/index.php/journal/article/download/225/68.

Looney, C.W. and Duston, C.L. (Aug. 2026) “Using Gemini and LuaLaTeX to transcribe physics videos into PDF/UA-2 and ISO 32005 math-accessible PDFs.” arXiv preprint. https://doi.org/10.48550/arXiv.2608.20733.

Meng, X. et al. (Aug. 2026) “EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering.” arXiv. Available at: https://doi.org/10.48550/arXiv.2608.21252.

Mittelbach, F. et al. (Aug. 2026) “Accessible Presentations with LATEX,” Proceedings of the 2026 ACM Symposium on Document Engineering, New York, NY, USA: Association for Computing Machinery (DocEng ’26). https://doi.org/10.1145/3820755.3821645.

Congratulations to some of our very active LaTeX Project LWG members for their work in improving accessibility of presentations with LaTeX!

Niko Isosaari (July 2026) Artificial intelligence-based method for automatic inspection of engineering drawings. Masters in Engineering thesis. Aalto University, Finland. https://aaltodoc.aalto.fi/server/api/core/bitstreams/32ff7397-0608-4637-9882-ba827b5e3c29/content.

Obadage, R.R. et al. (Sept. 2026) URL Extraction from Scholarly Documents: A Cross-Format Comparative Analysis, arXiv.org. https://doi.org/10.1145/3805696.3846023.

Wilson Ramos do Carmo Junior and Victor Travassos Sarinho (Sept. 2026) “PDFGameBuilder: A Reusable Game Engine for Executable PDF Documents.” SE4Games 2026, São Paulo, Brazil, p. 8. https://cbsoft.sbc.org.br/2026/data/papers/workshops/PDFGameBuilder%20A%20Reusable%20Game%20Engine%20for%20Executable%20PDF%20Documents.pdf.

Rigdon, G. (2026) “Rendered document indexing vs. structured source indexing: A comparative study for design-document multimodal RAG,” Issues in Information Systems, 27(1), p. 13. https://rigdonhouse.com/docs/AIML/VisionResearch_Rigdon_IIS.pdf

Romero, B.V.R. and Lapa, B.F.C. (Aug. 2026) “Open-Source Document-Image Compression and OCR Degradation: A Reproducible, Multi-Axis Evaluation,” IEEE Access. https://doi.org/10.1109/ACCESS.2026.3724788.

Schmieder, J. (2026) “TEXPDF: Stata module to provide standalone LaTeX-to-PDF compilation for Stata.” Boston College Department of Economics (Statistical Software Components). https://EconPapers.repec.org/RePEc:boc:bocode:s459880

Sigbara, B.M. et al. (Sept. 2026) “A Hybrid PDF Security Trainer Framework for Secure Document Classification and Storage in Distributed Cloud Environments,” Pan-African Journal of Research and Innovation in Computer Science and Mathematical Theory, 1(1), pp. 1–9. https://pajri.com/index.php/pajrijcsait/article/download/2/2

Soric, M. et al. (Sept. 2026) “Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents (Extended Version).” https://inria.hal.science/hal-05421324

Srikanth, S. et al. (Aug. 2026) “EvidenceFusion-MM: A Multimodal Forensic Framework for Detecting AI-Generated Text, Document Image Tampering, and PDF Structural Anomalies,” 2026 International Conference on Intelligent Multimedia, Networking, and Security (IMNS), pp. 1–6. https://doi.org/10.1109/IMNS67862.2026.11655259.

Trifan, M. et al. (July 2026) “Agentic Document Engineering: Automated Workflows for Editable Technical Artifacts,” 2026 IEEE 30th International Conference on Intelligent Engineering Systems (INES), pp. 637–642. https://doi.org/10.1109/INES69513.2026.11661118.

Wang, H. and Wang, S. (Aug. 2026) “Query-Conditioned Document-VLM Streaming for Time-to- First-Correct-Answer: Reinforcement Learning over Progressive PDF Tiles,” Global Sustainability, 1 (2026)(2), pp. 65–81. https://atlaspubs.com/index.php/GSD/article/download/23/9

Yayla, K. (Sept. 2026) “Measuring Latent Documentary Quality in Scholarly Communication: An Indicator Framework Based on PDF Validation and Technical Debt.” Research Square. Preprint: https://doi.org/10.21203/rs.3.rs-9861050/v1.

Zayed, A. (Aug. 2026) “Beyond AI Text Detection: Dr. Ahmed Zayed’s AI Fingerprint as an Integrated Multi-Engine and File-Forensic Framework for Auditable Authorship Evidence.” https://doi.org/10.5281/zenodo.22132745.


WordPress Cookie Notice by Real Cookie Banner