A PDF can carry useful technical clues about how the file was created. Those clues describe the artifact, not necessarily the origin of every sentence inside it.
PDF metadata commonly records the application or library that created the document. Values such as Word, Acrobat, wkhtmltopdf, Chromium or a PDF library can help reconstruct the export path. They do not prove who wrote the underlying content.
Zero-width characters, bidirectional controls, non-breaking spaces and soft hyphens can appear for legitimate formatting reasons. AI Doc Scan reports them because they are relevant to document inspection, but does not label them as an AI watermark.
C2PA Content Credentials are structured provenance metadata designed to describe origin and editing history for supported media workflows. They are richer than ordinary PDF metadata, but provenance metadata can be absent, stripped or unsupported. Absence is not proof that content was human-created.
OpenAI textGrain, Anthropic’s described approach and Google SynthID for text operate through token-selection statistics. Copying such text into another file does not create a special hidden Unicode tag that a generic PDF inspector can simply remove or reveal.
File metadata answers questions about the artifact. Statistical watermarking answers questions about model-generation signals. Stylometry describes writing behavior. Authorship and responsibility are broader human and legal questions.