AI Writing Diagnosticsby marcoderspace™ studio
Method overview

How AI Doc Scan works.

The scanner deliberately separates four different evidence layers: document extraction, model classification, stylometric description and file provenance. They answer different questions and should not be collapsed into a single claim about authorship.

Your PDF stays in the browser.

The document is opened with PDF.js on the device. Extracted text is analyzed locally. The page does not upload the PDF to a marcoderspace server. A first deep scan may download a quantized model into the browser.

1. Full-document PDF extraction

Every page is inspected for extractable text. Text objects are reconstructed into lines and approximate paragraphs using their PDF coordinates. This matters because page-level analysis is more useful than sampling only the opening pages of a long report.

2. Chunk-level model evidence

For eligible English text, the browser can run the open TMR RoBERTa AI-text detector through Transformers.js. Long documents are divided into sentence-preserving chunks, and the document score is computed as a word-weighted average rather than allowing one short section to dominate the result.

3. Explainable writing diagnostics

The classifier score is kept separate from descriptive features. The scanner measures sentence and paragraph variation, repeated n-grams, sentence-openers, lexical diversity, local lexical overlap, punctuation distribution, formulaic discourse, citations and other context signals. These features help explain the shape of the writing; none is treated as proof on its own.

4. Provenance and file signals

The PDF layer checks ordinary metadata such as Creator, Producer, Author and creation date, plus unusual invisible Unicode controls. Those observations are reported separately because a file artifact and a statistical text watermark are fundamentally different things.

5. Page heatmap and confidence

Where enough text exists, pages receive local scores. Confidence also considers document length and whether individual chunks agree. A detector that says “82%” on a tiny or highly unusual sample should not be interpreted in the same way as a consistent result across a long document.

What the score means

Evidence that the classifier associates with AI-generated text in its training domain.

What it does not mean

That a model authored a stated percentage of the work, or that a particular person did or did not write it.