Skip to content
Aback Tools Logo

PDF Metadata Extractor

Upload any PDF to extract all embedded metadata. The PDF Metadata Extractor reads both the Document Info Dictionary (title, author, subject, keywords, creator, producer, creation/modification dates) and XMP metadata (Dublin Core elements, Adobe XMP properties, rights information, document identifiers, language, and custom properties) using pdf-lib and pdfjs-dist. Fields are organized by source group (File Info, Document Info, XMP Metadata) and categorized by type (Identity, Content, Dates, Statistics, Software). With tab-based filtering, export as text, per-field descriptions, and a health indicator - all processing runs locally in your browser with zero uploads. Free, no signup required.

PDF Metadata Extractor

Upload any PDF to extract all embedded metadata, including the Document Info Dictionary (title, author, subject, keywords, creator, producer, creation/ modification dates) and XMP metadata when available. The tool uses pdf-lib and pdfjs-dist to read metadata locally - all processing runs entirely in your browser with zero uploads. No signup required.

Drop a PDF here or click to browse

Upload any PDF to extract its metadata - no file upload, all processing is local

Why Use Our PDF Metadata Extractor?

Document Info & XMP Extraction

Extracts both the PDF Document Info Dictionary (title, author, subject, keywords, creator, producer, creation and modification dates) and XMP metadata (Dublin Core, Adobe XMP) when present. Two metadata sources for maximum completeness.

Detailed Field Display with Categories

Metadata fields are organized into File Info, Document Info, and XMP Metadata groups. Each field is categorized as Identity, Content, Dates, Statistics, or Software. Tab-based filtering lets you focus on specific metadata sources.

Exportable Results with Field Descriptions

View all metadata in a clear, organized layout with per-field descriptions explaining what each field means. Copy all metadata as plain text with one click for documentation, reporting, or record-keeping.

100% Browser-Based & Private

All PDF processing happens entirely in your browser using pdf-lib and pdfjs-dist. Your documents never leave your device - no uploads, no servers, no third-party services. Complete privacy for sensitive documents.

Common Use Cases for PDF Metadata Extractor

Privacy & Data Leak Prevention

Before sharing PDFs externally, extract metadata to check for hidden author names, company information, software versions, and document properties that could expose sensitive organizational information to unintended recipients.

Document Forensics & Attribution

Identify the creator, author, and production toolchain of any PDF. The creator, producer, and XMP metadata fields help trace the origin of anonymous documents and verify authorship claims.

Intellectual Property Verification

Verify copyright ownership and document provenance through creation and modification timestamps. The XMP metadata also provides rights and source information for legal document verification.

Document Management & Archiving

When cataloging large PDF collections, extract metadata to populate document management systems, digital libraries, and archives with accurate title, author, subject, and date information.

Software & Workflow Compatibility

Check which application and PDF version created a document to ensure compatibility with your workflow. PDFs created in newer versions may have features not available in older PDF readers.

Regulatory Compliance & Auditing

Extract PDF metadata for regulatory compliance audits that require tracking document creation and modification history. The document info provides a verifiable record suitable for compliance documentation.

Understanding PDF Metadata Extraction

What Is PDF Metadata?

PDF metadata is descriptive information embedded inside PDF files that documents the file's properties - who created it, when it was created, what software was used, and keywords for categorization. This metadata is stored in two primary locations: the Document Info Dictionary (a simple key-value structure specified by the PDF standard) and XMP (Extensible Metadata Platform) metadata, an Adobe XML-based standard that provides richer, more extensible metadata including Dublin Core elements, rights information, and custom properties.

How PDF Metadata Extraction Works

The PDF Metadata Extractor uses two complementary PDF libraries for maximum coverage. pdf-lib reads the Document Info Dictionary directly to extract author, title, subject, keywords, creator, producer, and creation/modification dates. pdfjs-dist (Mozilla's PDF rendering library) then extracts XMP metadata by parsing the embedded XML metadata stream, providing access to Dublin Core elements (dc:title, dc:creator, dc:description), Adobe XMP properties (xmp:CreateDate, xmp:ModifyDate, xmp:CreatorTool), and rights management information (dc:rights). Results from both sources are merged and organized by source group for clarity.

Document Info vs XMP Metadata

The Document Info Dictionary is the older, simpler metadata system defined in the PDF specification. It supports a fixed set of fields: Title, Author, Subject, Keywords, Creator, Producer, CreationDate, and ModDate. XMP metadata, introduced with PDF 1.4 (Adobe Acrobat 5.0), provides a more comprehensive and extensible structure using XML. XMP can include Dublin Core metadata (title, creator, description, subject, rights, language), Adobe PDF properties (producer, PDFVersion), and custom namespaces. Modern PDFs often contain both, while older PDFs may only have the Document Info Dictionary.

Browser-Based & Private

All PDF processing runs entirely in your browser - your documents are never uploaded to any server. The tool uses pdf-lib to read the Document Info Dictionary and pdfjs-dist to extract XMP metadata, both running locally through WebAssembly and JavaScript. This makes it suitable for analyzing confidential PDFs, legal documents, contracts, and sensitive business files without any privacy concerns. No signup, no data collection, no usage tracking.

Related Tools

Steganography Detector

Upload an image to detect potential hidden data using LSB (Least Significant Bit) steganography analysis. Analyzes pixel-level modifications, shows a heatmap of altered pixels, detects statistical anomalies, and extracts hidden text messages if found. Compatible with standard LSB encoding and the STEG magic header format. All processing happens locally in your browser - free online steganography detector, no signup required.

Image File Size Analyzer

Upload any image to see a detailed binary-level breakdown of what contributes to its file size - pixel data, compression overhead, metadata (EXIF/IPTC/XMP), color profiles (ICC), and structural headers. Get prioritized optimization tips with estimated savings for web performance, mobile app optimization, and storage reduction. Supports JPEG, PNG, GIF, WebP, BMP, TIFF, AVIF, and HEIC - all processing runs locally in your browser, no server upload required. Free online image file size analyzer.

File Magic Byte Detector

Upload any file to instantly detect its true file type by reading the magic bytes (file signature / header). The tool bypasses incorrect file extensions and reveals the real format. Shows hex dump with matching signature bytes highlighted, ASCII interpretation, MIME type, file category, and extension match analysis. Supports over 120 file signatures across 16 categories: images, audio, video, documents, archives, executables, fonts, certificates, disk images, and more. All processing runs locally in your browser - free online File Magic Byte Detector, no signup required.

Image Clone Detector

Detect copy-move forgeries in images using pixel-block matching. Upload any image to find cloned/copied regions, view heatmap overlays showing the location and intensity of detected clones, and get confidence scores with region details. Three sensitivity presets for different detection needs. 100% private browser-based processing - free online Image Clone Detector.

Frequently Asked Questions About PDF Metadata Extractor
What is a PDF Metadata Extractor?
A PDF Metadata Extractor reads the hidden metadata embedded in PDF files and displays it in a readable format. This includes the Document Info Dictionary (title, author, subject, keywords, creator, producer, creation/modification dates) and XMP metadata (Dublin Core elements, Adobe XMP properties, rights information). This metadata is not visible when viewing the PDF content and can reveal valuable information about the document's origin and history.
What is the difference between Document Info and XMP metadata?
The Document Info Dictionary is a simple key-value metadata store defined in the PDF specification, supporting a fixed set of fields like Title, Author, Subject, and Keywords. XMP (Extensible Metadata Platform) is an XML-based standard that provides richer, extensible metadata including Dublin Core elements, custom properties, rights management, and provenance information. Modern PDFs often contain both, with XMP providing more comprehensive metadata that can include custom namespaces and structured data.
What metadata fields can be extracted from a PDF?
From the Document Info Dictionary: title, author, subject, keywords, creator application, producer application, creation date, and modification date. From XMP metadata: title, author, description, subject, creator tool, create/modify/metadata dates, document identifier, language, rights/copyright information, source URL, and any custom namespace properties. The tool also extracts basic file info: file name, file size, PDF version, page count, and encryption status.
Can this tool extract metadata from encrypted or password-protected PDFs?
The tool attempts to read metadata from encrypted PDFs using pdf-lib's ignoreEncryption option, which can sometimes access the Document Info Dictionary without the password. However, strongly encrypted PDFs (AES-256) with restricted metadata access may not be readable. XMP metadata extraction from encrypted PDFs typically requires the correct password. The tool will clearly report if a PDF is encrypted and whether metadata extraction was successful.
Is my PDF file uploaded to any server?
No. All processing happens entirely in your browser. The PDF file is read locally using the File API and processed in memory by pdf-lib and pdfjs-dist. Your documents never leave your device - no uploads, no data collection, no third-party server requests. This is the core privacy guarantee of all Aback Tools.
Why are some metadata fields missing or empty?
Metadata fields may be missing because the PDF creator did not populate them, the metadata was stripped for privacy reasons, or the PDF was created with a tool that doesn't fill in certain fields. The tool clearly shows which fields have values and which are empty. The health indicator (Complete/Partial/Minimal) provides a quick assessment of how much metadata was found.
What PDF versions are supported?
The tool supports PDF files of all versions from PDF 1.0 through the latest PDF 2.0 specification. The PDF version is detected from the file header. XMP metadata is supported from PDF 1.4 onward. Older PDFs may only have Document Info Dictionary metadata, which is fully supported. The tool works with linearized (fast web view) PDFs, tagged PDFs, and PDF/A archival documents.
Is this PDF Metadata Extractor free to use?
Yes - 100% free, forever. No signup, no account, no premium tier, no usage limits, and no ads. Analyze as many PDFs as you need, from small one-page documents to large multi-hundred-page files. All processing is done locally in your browser with zero server costs.