PDF Association logo

Discover pdfa.org

Key resources

Get involved

How do you find the right PDF technology vendor?
Use the Solution Agent to ask the entire PDF communuity!
The PDF Association celebrates its members’ public statements
of support
for ISO-standardized PDF technology.

Member Area

What does “PDF” stand for?

C’mon, AI – PDF isn’t just about text! | Supports just 2 of 4 operators? That’s not a PDF parser! | What does “PDF” stand for? | 6 steps and a lie to save a PDF? | Using PDF to inject AI prompts | Microsoft recommends archiving Publisher files to PDF | Microsoft Edge gets in on promoting PDF editing | GeoPDF is not an “open standard” | This PDF runs AI | ICC HDR and Spectral Imaging Experts’ Day announced | Yet another document format for AI | Zotero 10 expands PDF capabilities | PDF’s popularit … Read more

PDF in the WildAugust 25, 2026
Screenshot of an AI-generated dictionary definition of
What does “PDF” stand for?
Screenshot of an AI-generated dictionary definition of

C’mon, AI – PDF isn’t just about text! | Supports just 2 of 4 operators? That’s not a PDF parser! | What does “PDF” stand for? | 6 steps and a lie to save a PDF? | Using PDF to inject AI prompts | Microsoft recommends archiving Publisher files to PDF | Microsoft Edge gets in on promoting PDF editing | GeoPDF is not an “open standard” | This PDF runs AI | ICC HDR and Spectral Imaging Experts’ Day announced | Yet another document format for AI | Zotero 10 expands PDF capabilities | PDF’s popularit … Read more

PDF in the WildAugust 25, 2026

PDF Association staff

About PDF Association staff


C’mon, AI – PDF isn’t just about text!

Developer Federico Ricciuti wrote an article highlighting the importance of metadata for effective LLM understanding of input documents. Among his other observations, he points out:

For me, one of the most compelling applications is analyzing how much personally identifiable information (PII) is hidden in these large datasets. It is concerning that this data—whether partial, compressed, or transformed—can end up embedded directly within LLM weights, where auditing is notoriously difficult (even for open-weight models). Understanding what kind of information these datasets hold is crucial.

Supports just 2 of 4 operators? That’s not a PDF parser!

Assuming his documentation reflects reality, this developer’s Rust library for PDF classification (which he announced on LinkedIn) says it supports only 2 of PDF’s 4 text-showing operators (see ISO 32000, Table 107).

No wonder it runs fast. Still a #fail.

What does “PDF” stand for?

From the culture desk

Screenshot of a fake dictionary entry for "PDF".
AI-generated image.

As the Huffington Post points out about “PDF”…

“The word has become so embedded in everyday language that most people use it the way they use terms like “Wi-Fi” and “p.m.” ― automatically and without any particular thought about what the letters mean.”

The Post quotes language trends expert  Madeline Enos as noting that:

“Most acronyms stay within specialist industries, but PDF made the rare jump into everyday language. Today, people don’t just recognize the letters ― they also use ‘PDF’ as a verb,” Enos said. “It’s increasingly common to hear someone say ‘Can you PDF it?’ or ‘Just PDF that document,’ showing how technical abbreviations can become everyday words over time.”

Lisa Gitelman, a media historian and professor at New York University, told the reporter:

“I once wrote an essay laying out four reasons for its enormous success,” Gitelman said. “First, PDFs offer page images ― like facsimiles, like xeroxes, like microfilm, unlike word processing files or most proprietary eBook formats. This means they feel stable. They retain layout as well as content. Second, as page images, they are handy in publication workflows. Every book, every analog publication, has been a PDF somewhere in its life history. Third, they are easily transmissible across digital networks, even in low-bandwidth environments.” The fourth and final reason: PDFs are searchable, she added.

We couldn’t have said it better ourselves!

Our CEO is also quoted, so of course, we have to put it in PDF in the Wild!

6 steps and a lie to save a PDF?

Great question, Becca!

In fact, we see this less and less. Most modern applications that offer PDF output do so with a dedicated menu option, rather than forcing the user through a printing workflow.

More importantly, in the 21st century, print workflows are usually a terrible way to produce PDF files. This is because printers have no use for the semantic information that provides accessibility support and facilitates rich content extraction and reuse – what today’s AI systems need to thrive.

Always look for the “Save as PDF” or “Export to PDF” ahead of printing to PDF – unless you’re actually printing, of course!

Using PDF to inject AI prompts

PDF has many ways to make text invisible to human eyes, but visible to machines.

A Connecticut, USA court recently found a series of instructions in an official court filing that were designed to manipulate artificial intelligence.

Small white text is one of the simplest means of such injection and is therefore visible to the widest possible number of machines, almost regardless of how they ingest PDFs.

It’s also easy for LLMs to detect, and therefore defeat.

“…404 Media uploaded the plaintiff’s motion to OpenAI’s ChatGPT and asked it to render a decision on the case. ChatGPT ruled against the motion. When we asked it if the filing contained a prompt injection, it said that I noticed and ignored it in my analysis. It did not influence the proposed denial. Its presence also raises a credibility and professionalism concern.”

Microsoft recommends archiving Publisher files to PDF

The venerable Publisher, which originally shipped in 1991, will reach its end-of-life (EOL) this October.

Microsoft’s first recommendation: export to PDF.

Microsoft Edge gets in on promoting PDF editing

We’ve previously reported on browsers integrating PDF features… now it seems to be Edge’s turn.

Screenshot from a Microsoft Edge promotional page showing a PDF editor feature.

GeoPDF is not an “open standard”

In a recent video, GeoPDF’s developer claims that GeoPDF is an “open standard”.

The Library of Congress’s authoritative Sustainability of Digital Formats website makes clear that while GeoPDF has been publicly documented the technology is not an open standard and should not be referenced as such.

According to the Open Geospatial Consortium, “PDF Georegistration Encoding Best Practice Version 2.2” is simply best practice guidance. The document explicitly states as a warning on its front cover

“This document defines an OGC Best Practice on a particular technology or approach related to an OGC standard. This document is not an OGC Standard and may not be referred to as an OGC Standard. It is subject to change without notice. However, this document is an official position of the OGC membership on this particular technology topic.” (emphasis added)

This PDF runs AI

We’ve seen a PDF that’s the size of the universe. We’ve seen PDFs that run video games.

What’s next?

Turns out… a PDF that runs an LLM.

What’s next?

We don’t know… but if and when someone manages to pack a restaurant into a PDF, you’ll be reading about it in PDF in the Wild!

ICC HDR and Spectral Imaging Experts’ Day announced

Color and imaging specialists, including members of the Imaging Model TWG, may be interested in this upcoming ICC event in Norway. The agenda covers a range of color imaging topics with a focus on HDR.

Replace PDF with “AI native” format …

A recent paper by Liu, J. et al. (May 2026) “The Last Human-Written Paper: Agent-Native Research Artifacts.” arXiv. Preprint. https://doi.org/10.48550/arXiv.2604.24658 has been getting a lot of headlines, especially in STEM publishing communities. The paper promotes https://www.agenticresearch.sh/ by start-up Orchestra Research Institute.

These headlines are entirely misleading, as the article is not about PDF as a file format; it is about publishing practices that fail to capture the full range of information desired by AI

“We introduce the [Agent-Native Research Artifact] ARA protocol and its surrounding ecosystem as a foundation for agent-native scientific communication. Together, they address two structural failures of the PDF format: knowledge that narrative conventions discard (failed attempts, implicit configurations, unexplored branches) and specifications too underspecified to execute. ARA resolves both by restructuring a research contribution as a machine-actionable artifact, one that is navigable, complete, and verifiable without human interpretation.”

In fact, PDF does not have these “structural failures”, and as our CTO has previously described, PDF is a far more capable format for STEM publishing than its current use by STEM publishers implies.

Today’s PDF supports open data, MathML, detailed rich semantics, annotation and review capabilities, authenticity and provenance, metadata and, of course, extremely long documents - all in a standalone archival-friendly standardized format addressing “bit rot”.

Zotero 10 expands PDF capabilities

We use Zotero to manage the growing library of literature mentioning PDF – see below for this month’s PDFacademicBot. The latest Zotero 10 release adds several new features for PDF files, highlighting the importance of PDF for academic and industry publishing, including a reflow reading mode, smarter read-aloud, and advanced searching.

PDF’s popularity

Screenshot of a tweet by Forth noting that in searching the web "PDF" has overtaken all of the abrahamic religions when it comes to search-engine guessing the next word in a prompt that begins "how do I convert to".

PDFacademicBot for August 2026

The PDFacademicBot brings academic research on PDF and related technologies to the industry’s attention.

Aflisia, N. et al. (2026) “Experimental Study on the Effectiveness of Nahwu E-Module Using Flip PDF Professional in Enhancing Students’ Higher-Order Thinking Skills (HOTS),” in A.J. Obaid et al. (eds.) Emerging Technologies for Next Generation Learning and Education Spaces. Cham: Springer Nature Switzerland, pp. 17–23. https://doi.org/10.1007/978-3-032-21822-3_3.

Ferawati, D. (2026) “Islamic Religious Education Teachers’ Strategies for Implementing a Digital Learning Model in Teaching Rukhshah Material Using PDF Format at SMP Negeri 1 Bunguran Utara, Natuna Regency, Riau Islands,” Indonesian Journal of Innovation Multidisipliner Research, 4(3). https://multidisipliner.org/ijim/article/download/1867/1350

Guo, S. et al. (Aug. 2026) “From Documentation to Zero-day Vulnerabilities: LLM-Driven Fuzzing of JavaScript Engines in PDF Readers.” arXiv preprint. Accepted by ACM CCS 2026. https://doi.org/10.48550/arXiv.2608.06641.

Kim, T.-W., Jang, H.-S. and Lee, D. (2026) “Design of a Lightweight PDF-Based Personal Information Detection System to Prevent Data Leakage from Residual Files : A Comparative Perspective with Deep Learning Approaches,” INTERNATIONAL CONFERENCE ON FUTURE INFORMATION & COMMUNICATION ENGINEERING, 17(1), pp. 237–240. https://www.dbpia.co.kr/journal/articleDetail?nodeId=NODE12917798

Liu, J. et al. (May 2026) “The Last Human-Written Paper: Agent-Native Research Artifacts.” arXiv. Preprint. https://doi.org/10.48550/arXiv.2604.24658.

Madhava Rao, D. et al. (July 2026) “SKIM-PDF: A Unified IQR–Chi-Squared Adaptive Selection Pipeline for Structural PDF Malware Detection,” Indian Journal of Computer Science and Technology, pp. 920–927. https://doi.org/10.59256/indjcst.20260502100.

N. Abarna et al. (2026) “Explainable Hybrid Graph-Autoencoder Framework for Advanced PDF Malware Detection,” 2026 9th International Conference on Computing Methodologies and Communication (ICCMC), pp. 865–870. https://doi.org/10.1109/ICCMC69250.2026.11625149.

Sai Medikonda, M.L. et al. (2026) “An Intelligent PDF Question-Answering System; A Retrieval-Augmented Generation Approach,” 2026 9th International Conference on Computing Methodologies and Communication (ICCMC), pp. 1793–1800. https://doi.org/10.1109/ICCMC69250.2026.11624861.

Pedro de Teixeira Lopes Basílio Plácido (2026) Deterministic Information Extraction from Business Documents. Master of Science in Computer Science and Engineering. Universidade do Porto. https://repositorio-aberto.up.pt/bitstream/10216/175609/2/786417.pdf.

Rasyid, S.A. et al. (2026) “Evaluating Vision-Language Models vs. OCR Pipelines for Spatial Grounding in ISO/IEC 17025 Audits,” 2026 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT), pp. 617–622. https://doi.org/10.1109/IAICT71158.2026.11620745.

Sanaz Aghajamali, Saeid Metvaei, and Zhen Lei (August 2026) “AI-ASSISTED PDF EXTRACTION AND RULE-BASED PRODUCTION TIME ESTIMATION IN STRUCTURAL STEEL ASSEMBLY USING HISTORICAL PRODUCTION DATA.” Transforming Construction with Off-site Methods and Technologies (TCOT 2026), Newcastle upon Tyne, UK, p. 10. https://journals.northumbria.ac.uk/index.php/tcot/article/download/1914/2024.

Sharma, D. and Devare, M. (2026) “Beyond Post-Hoc Interpretability: Active SHAP-Guided Dynamic Ensemble Routing For PDF Malware Detection,” International Journal of Artificial Intelligence and Machine Learning, 6(6), pp. 668–679. https://svedbergopen.com/index.php/ijaiml/article/download/740/577.

Sharma, R. and Gautam, D. (2026) “AI-Based Document Analysis and Question Answering System,” Revolutionary Advances in Computing and Electronics: An International Journal, 2(2), pp. 17–32. https://doi.org/10.65890/race.v2i2.192.

Tremblay Taillon, N. and Langlais, P. (May 2026) “eSciBench: An Extensible Scientific PDF Extraction Benchmark,” in S. Piperidis et al. (eds.) Proceedings of the Fifteenth Language Resources and Evaluation Conference. LREC 2026, Palma de Mallorca, Spain: ELRA Language Resource Association, pp. 7568–7580. https://doi.org/10.63317/4sxku4i2piqq.

Shie, S.-S. et al. (Aug. 2026) “Direct Reconstruction of High-Fidelity Electrocardiogram Signals From Vector-Based PDF Files With Integrated Deep Learning for Multiparameter Estimation: Retrospective Methodological Study,” JMIR Formative Research, 10(1), p. e80597. https://doi.org/10.2196/80597.

Yajnik, A. and Pushkala, S.A. (Aug. 2026) “A Controlled Comparative Evaluation of Workflow Configurations for Source-Grounded Academic PDF Analysis,” IEEE Access. https://doi.org/10.1109/ACCESS.2026.3719291.

Zhang, X. et al. (2026) “Not as Sweet by Another Name: An Empirical Study of Format Robustness in LLM Document Workflows,” arXiv.org. 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026), p. 13. https://doi.org/10.1145/3832783.3837456.


WordPress Cookie Notice by Real Cookie Banner