AI Writing Diagnosticsby marcoderspace™ studio
Document forensics

PDF provenance is not the same as AI authorship.

A PDF can carry useful technical clues about how the file was created. Those clues describe the artifact, not necessarily the origin of every sentence inside it.

Creator and Producer

PDF metadata commonly records the application or library that created the document. Values such as Word, Acrobat, wkhtmltopdf, Chromium or a PDF library can help reconstruct the export path. They do not prove who wrote the underlying content.

Invisible Unicode

Zero-width characters, bidirectional controls, non-breaking spaces and soft hyphens can appear for legitimate formatting reasons. AI Doc Scan reports them because they are relevant to document inspection, but does not label them as an AI watermark.

Content Credentials and C2PA

C2PA Content Credentials are structured provenance metadata designed to describe origin and editing history for supported media workflows. They are richer than ordinary PDF metadata, but provenance metadata can be absent, stripped or unsupported. Absence is not proof that content was human-created.

Statistical text watermarking is different

OpenAI textGrain, Anthropic’s described approach and Google SynthID for text operate through token-selection statistics. Copying such text into another file does not create a special hidden Unicode tag that a generic PDF inspector can simply remove or reveal.

Think in layers.

File metadata answers questions about the artifact. Statistical watermarking answers questions about model-generation signals. Stylometry describes writing behavior. Authorship and responsibility are broader human and legal questions.