Information gain
Information gain is the measurable new information a page adds beyond what already exists in the corpus a model was trained on. It is the value a page contributes that could not be reconstructed from existing sources — original data, a named position, or first-hand results — and it is what makes a page worth retrieving rather than paraphrasing.
Why information gain matters for an existing site
If an LLM can write your page without retrieving it, you don’t have a page. You have training data with a URL. That’s the whole test, and most content fails it — it’s a competent rephrase of things the model already knows, which is why it never gets cited. The bar to clear is low and brutal: one original number per page that exists nowhere else on the internet. Not a rounded-up stat you borrowed. Yours. Ours is the 8% median citation share from our audit of 50 established sites — a journalist can quote it, a model can retrieve it, and nobody else has it. That’s the difference between being a source and being a paraphrase.
Three places gain actually comes from:
- Original data — your study, your benchmark, your numbers. One dataset earns citations for a year.
- A defended position — a stance stated plainly, with the reasoning shown. “It depends” contributes nothing.
- First-person results — what happened when you did the thing. “I deleted 400 pages and traffic went up” can’t be paraphrased from anyone else.
Related: Topical authority · E-E-A-T Read up: Information Gain: Pages LLMs Can’t Paraphrase · Content Audit for Live Sites · How to Market a Website in 2026