PDF Association logo

Discover pdfa.org

Key resources

Get involved

How do you find the right PDF technology vendor?
Use the Solution Agent to ask the entire PDF communuity!
The PDF Association celebrates its members’ public statements
of support
for ISO-standardized PDF technology.

Member Area

156 Million Pages. How ITHAKA Built PDF/UA Compliance on AWS with PDFix

PDFix replaced Adobe in a production pipeline serving millions of researchers. Here’s what that looks like at scale.

Case studySeptember 25, 2026
aws, ITHAKA, JSTOR, and PDFix logo
156 Million Pages. How ITHAKA Built PDF/UA Compliance on AWS with PDFix
aws, ITHAKA, JSTOR, and PDFix logo

PDFix replaced Adobe in a production pipeline serving millions of researchers. Here’s what that looks like at scale.

Case studySeptember 25, 2026

Diana Kosovac

About Diana Kosovac, PDFix

DISCLAIMER
The views expressed in this article are those of the author(s) and do not reflect the policies or positions of the PDF Association.

JSTOR runs one of the largest digital libraries in the world – 20 million PDFs, 156 million pages, scholarly content going back to 1550. ADA Title II set the deadline. The scale set the challenge.

The pipeline started from an open source AWS solution originally built with the Adobe PDF Accessibility Auto-Tag API. ITHAKA’s team evaluated it and made a deliberate switch – replacing Adobe with PDFix and veraPDF for structural tagging and PDF/UA validation. The swap happened without rebuilding the pipeline. That modularity was the point.

Manual remediation at $1-4 per page meant a potential bill of up to $624 million for the full corpus. Bulk upfront processing wasn’t realistic for a nonprofit. So ITHAKA and AWS built on-demand: the pipeline remediates a PDF when a user actually requests it and caches it for every user after that.

PDFix handles structural tagging- headings, paragraphs, lists, tables – with veraPDF validating PDF/UA compliance and Amazon Bedrock generating alt text for images. The pipeline runs on AWS Lambda, Step Functions, and Fargate. A PDF comes in, gets processed page by page in parallel, reassembled, validated, and cached.

Architecture diagram of ITHAKA's PDF remediation pipeline on AWS: a request event triggers Lambda to split the PDF, AWS Step Functions processes each chunk through PDFix on Fargate for structural tagging and Amazon Bedrock for alt text generation, followed by PDF merging, veraPDF validation, and storage in an S3 pipeline bucket.

The results after moving to production:

  • 98% accessibility check pass rate
  • $0.026 per page – down from $1-4 manually
  • 97%+ cost reduction
  • Production-ready in approximately 2 months from a 2-day build workshop

The underlying solution is open source. Universities, libraries, government agencies, and nonprofits facing similar compliance deadlines can adapt it for their own environments.

Read the full case study on the AWS Public Sector Blog →


PDFix is a pioneering provider of AI-driven PDF accessibility solutions, empowering organizations to achieve PDF/UA and WCAG compliance through intelligent automation. With over 25 years of expertise in PDF technology, PDFix transforms complex document workflows by reducing manual remediation time by while ensuring accuracy, scalability, and full regulatory compliance environments…

Read more

WordPress Cookie Notice by Real Cookie Banner