The scanner deliberately separates four different evidence layers: document extraction, model classification, stylometric description and file provenance. They answer different questions and should not be collapsed into a single claim about authorship.
The document is opened with PDF.js on the device. Extracted text is analyzed locally. The page does not upload the PDF to a marcoderspace server. A first deep scan may download a quantized model into the browser.
Every page is inspected for extractable text. Text objects are reconstructed into lines and approximate paragraphs using their PDF coordinates. This matters because page-level analysis is more useful than sampling only the opening pages of a long report.
For eligible English text, the browser can run the open TMR RoBERTa AI-text detector through Transformers.js. Long documents are divided into sentence-preserving chunks, and the document score is computed as a word-weighted average rather than allowing one short section to dominate the result.
The classifier score is kept separate from descriptive features. The scanner measures sentence and paragraph variation, repeated n-grams, sentence-openers, lexical diversity, local lexical overlap, punctuation distribution, formulaic discourse, citations and other context signals. These features help explain the shape of the writing; none is treated as proof on its own.
The PDF layer checks ordinary metadata such as Creator, Producer, Author and creation date, plus unusual invisible Unicode controls. Those observations are reported separately because a file artifact and a statistical text watermark are fundamentally different things.
Where enough text exists, pages receive local scores. Confidence also considers document length and whether individual chunks agree. A detector that says “82%” on a tiny or highly unusual sample should not be interpreted in the same way as a consistent result across a long document.
Evidence that the classifier associates with AI-generated text in its training domain.
That a model authored a stated percentage of the work, or that a particular person did or did not write it.