AI Doc Scan is intentionally transparent about what it measures. The scanner combines a local text classifier with independent descriptive diagnostics and PDF-level observations.
The browser uses the ONNX Community conversion of Oxidane/tmr-ai-text-detector, a RoBERTa-base classifier trained on 50,000 stratified RAID samples. The model card identifies label 0 as human and label 1 as AI and lists English as its primary language. AI Doc Scan runs the quantized model locally through Transformers.js.
TMR ONNX model card · Model label configuration
Current diagnostics include sentence-length coefficient of variation, paragraph-length variation, repeated n-grams, repeated sentence openers, formulaic-discourse density, normalized lexical entropy, type-token ratio, hapax share, MTLD, Yule’s K, adjacent-sentence lexical overlap, function-word share, punctuation entropy, approximate citation density and quoted-text share.
PDF.js reads text objects locally. The scanner reconstructs approximate lines and paragraphs using coordinates and analyzes every page with extractable text. Image-only scans require OCR and are reported as unreadable rather than silently scored.
OpenAI textGrain announcement · OpenAI provenance signals · Anthropic watermark explainer · Google DeepMind SynthID
Detector models, thresholds and research evolve. This page is intended to make changes visible rather than presenting the scanner as an opaque oracle.