PDF Metadata Extractor
Upload any PDF to extract all embedded metadata. The PDF Metadata Extractor reads both the Document Info Dictionary (title, author, subject, keywords, creator, producer, creation/modification dates) and XMP metadata (Dublin Core elements, Adobe XMP properties, rights information, document identifiers, language, and custom properties) using pdf-lib and pdfjs-dist. Fields are organized by source group (File Info, Document Info, XMP Metadata) and categorized by type (Identity, Content, Dates, Statistics, Software). With tab-based filtering, export as text, per-field descriptions, and a health indicator - all processing runs locally in your browser with zero uploads. Free, no signup required.
Upload any PDF to extract all embedded metadata, including the Document Info Dictionary (title, author, subject, keywords, creator, producer, creation/ modification dates) and XMP metadata when available. The tool uses pdf-lib and pdfjs-dist to read metadata locally - all processing runs entirely in your browser with zero uploads. No signup required.
Drop a PDF here or click to browse
Upload any PDF to extract its metadata - no file upload, all processing is local
Document Info & XMP Extraction
Extracts both the PDF Document Info Dictionary (title, author, subject, keywords, creator, producer, creation and modification dates) and XMP metadata (Dublin Core, Adobe XMP) when present. Two metadata sources for maximum completeness.
Detailed Field Display with Categories
Metadata fields are organized into File Info, Document Info, and XMP Metadata groups. Each field is categorized as Identity, Content, Dates, Statistics, or Software. Tab-based filtering lets you focus on specific metadata sources.
Exportable Results with Field Descriptions
View all metadata in a clear, organized layout with per-field descriptions explaining what each field means. Copy all metadata as plain text with one click for documentation, reporting, or record-keeping.
100% Browser-Based & Private
All PDF processing happens entirely in your browser using pdf-lib and pdfjs-dist. Your documents never leave your device - no uploads, no servers, no third-party services. Complete privacy for sensitive documents.
Privacy & Data Leak Prevention
Before sharing PDFs externally, extract metadata to check for hidden author names, company information, software versions, and document properties that could expose sensitive organizational information to unintended recipients.
Document Forensics & Attribution
Identify the creator, author, and production toolchain of any PDF. The creator, producer, and XMP metadata fields help trace the origin of anonymous documents and verify authorship claims.
Intellectual Property Verification
Verify copyright ownership and document provenance through creation and modification timestamps. The XMP metadata also provides rights and source information for legal document verification.
Document Management & Archiving
When cataloging large PDF collections, extract metadata to populate document management systems, digital libraries, and archives with accurate title, author, subject, and date information.
Software & Workflow Compatibility
Check which application and PDF version created a document to ensure compatibility with your workflow. PDFs created in newer versions may have features not available in older PDF readers.
Regulatory Compliance & Auditing
Extract PDF metadata for regulatory compliance audits that require tracking document creation and modification history. The document info provides a verifiable record suitable for compliance documentation.
What Is PDF Metadata?
PDF metadata is descriptive information embedded inside PDF files that documents the file's properties - who created it, when it was created, what software was used, and keywords for categorization. This metadata is stored in two primary locations: the Document Info Dictionary (a simple key-value structure specified by the PDF standard) and XMP (Extensible Metadata Platform) metadata, an Adobe XML-based standard that provides richer, more extensible metadata including Dublin Core elements, rights information, and custom properties.
How PDF Metadata Extraction Works
The PDF Metadata Extractor uses two complementary PDF libraries for maximum coverage. pdf-lib reads the Document Info Dictionary directly to extract author, title, subject, keywords, creator, producer, and creation/modification dates. pdfjs-dist (Mozilla's PDF rendering library) then extracts XMP metadata by parsing the embedded XML metadata stream, providing access to Dublin Core elements (dc:title, dc:creator, dc:description), Adobe XMP properties (xmp:CreateDate, xmp:ModifyDate, xmp:CreatorTool), and rights management information (dc:rights). Results from both sources are merged and organized by source group for clarity.
Document Info vs XMP Metadata
The Document Info Dictionary is the older, simpler metadata system defined in the PDF specification. It supports a fixed set of fields: Title, Author, Subject, Keywords, Creator, Producer, CreationDate, and ModDate. XMP metadata, introduced with PDF 1.4 (Adobe Acrobat 5.0), provides a more comprehensive and extensible structure using XML. XMP can include Dublin Core metadata (title, creator, description, subject, rights, language), Adobe PDF properties (producer, PDFVersion), and custom namespaces. Modern PDFs often contain both, while older PDFs may only have the Document Info Dictionary.
Browser-Based & Private
All PDF processing runs entirely in your browser - your documents are never uploaded to any server. The tool uses pdf-lib to read the Document Info Dictionary and pdfjs-dist to extract XMP metadata, both running locally through WebAssembly and JavaScript. This makes it suitable for analyzing confidential PDFs, legal documents, contracts, and sensitive business files without any privacy concerns. No signup, no data collection, no usage tracking.
Related Tools
Steganography Detector
Upload an image to detect potential hidden data using LSB (Least Significant Bit) steganography analysis. Analyzes pixel-level modifications, shows a heatmap of altered pixels, detects statistical anomalies, and extracts hidden text messages if found. Compatible with standard LSB encoding and the STEG magic header format. All processing happens locally in your browser - free online steganography detector, no signup required.
Image File Size Analyzer
Upload any image to see a detailed binary-level breakdown of what contributes to its file size - pixel data, compression overhead, metadata (EXIF/IPTC/XMP), color profiles (ICC), and structural headers. Get prioritized optimization tips with estimated savings for web performance, mobile app optimization, and storage reduction. Supports JPEG, PNG, GIF, WebP, BMP, TIFF, AVIF, and HEIC - all processing runs locally in your browser, no server upload required. Free online image file size analyzer.
File Magic Byte Detector
Upload any file to instantly detect its true file type by reading the magic bytes (file signature / header). The tool bypasses incorrect file extensions and reveals the real format. Shows hex dump with matching signature bytes highlighted, ASCII interpretation, MIME type, file category, and extension match analysis. Supports over 120 file signatures across 16 categories: images, audio, video, documents, archives, executables, fonts, certificates, disk images, and more. All processing runs locally in your browser - free online File Magic Byte Detector, no signup required.
Image Clone Detector
Detect copy-move forgeries in images using pixel-block matching. Upload any image to find cloned/copied regions, view heatmap overlays showing the location and intensity of detected clones, and get confidence scores with region details. Three sensitivity presets for different detection needs. 100% private browser-based processing - free online Image Clone Detector.