Content · Auditing

Half the Site
Was a Copy of Itself

Every page passed a word-count check. Every page had a distinct title, a distinct URL, and numbers nobody else had. And slightly under half of all the text on the site existed on more than one page. Here is the fifteen-line measurement that found it, and the three structurally different ways a page ends up duplicating its neighbour without anyone deciding it should.

Duplication inside your own site is hard to see by reading. You open two pages, they look different (different headings, different numbers, different titles in the tab), and you move on. The eye compares the parts that change. It does not add up the parts that don't.

So measure it instead. The question worth asking is not "are these two pages similar" but what fraction of this page's text exists only on this page. That number is cheap to compute and it ranks the whole site at once.

The Measurement

Break every page into overlapping 8-word windows (shingles) and count how many pages each window appears on. A window that appears on exactly one page is that page's own writing. Everything else is shared with a neighbour.

import re, html, glob
from collections import Counter

K = 8  # shingle length, in words

def text(path):
    s = open(path, encoding="utf-8").read()
    s = re.sub(r"<script.*?</script>", "", s, flags=re.S | re.I)
    s = re.sub(r"<style.*?</style>", "", s, flags=re.S | re.I)
    s = html.unescape(re.sub(r"<[^>]+>", " ", s))
    return re.sub(r"\s+", " ", s).strip()

def shingles(t):
    w = t.split()
    return {" ".join(w[i:i + K]) for i in range(max(0, len(w) - K + 1))}

pages = {f: shingles(text(f)) for f in glob.glob("*.html")}

df = Counter()
for s in pages.values():
    for x in s:
        df[x] += 1

for f, s in sorted(pages.items(), key=lambda kv: len(kv[1]) and
                   sum(df[x] == 1 for x in kv[1]) / len(kv[1])):
    if not s:
        continue
    unique = sum(1 for x in s if df[x] == 1)
    print(f"{unique / len(s):6.1%}  {unique:5d}/{len(s):<5d}  {f}")

Eight words is long enough that ordinary phrases ("of the following") don't collide by accident, and short enough to catch a sentence that was copied and had two numbers changed. Sorting ascending puts the worst offenders first, which is where the interesting failures are.

Strip the scripts first

If your pages carry inline JavaScript that builds markup from template literals, a naive tag-stripper will read those literals as page text. That inflates the shared count on every page that ships the same bundle and buries the signal. On the site I ran this against, skipping that step once produced a table-width measurement for a table that did not exist.

Shape One: The Template With One Variable

Eighty pages, one per birth year, each around 750 characters. Each had a distinct title, a distinct URL, and a distinct heading. Four body sections carried the prose.

All four were keyed on the same single field, a five-value classification derived from the year, and read their text straight out of a five-entry dictionary. The pages for two consecutive years were identical below the heading, character for character, because both years mapped to the same class.

Eighty pages. Five distinct bodies. Nothing in the code looked wrong; each function did what its name said. The defect only exists at the level of the whole set, and no test that looks at one page can see it.

Put generally, a generated page is as distinct as its narrowest input, not its widest one. If the URL varies over 80 values and the prose varies over 5, you have 5 pages wearing 80 URLs. Count the distinct outputs, not the distinct inputs.

Shape Two: The Hub That Renders Its Own Child

A comparison page with four product tabs. The tab contents are prerendered into four separate landing pages, one per product, so each is indexable on its own terms. Sensible design.

But the hub itself has to render something before you touch a tab, and it renders the first tab. Which is the first landing page. The two URLs came back with 131 lines of visible text each, differing in exactly one line: the title.

This one is invisible in code review because the duplication is not in the source. It is in what the source produces. One file, two URLs, one page. It also survives every "does each page have unique metadata" check, because the metadata is unique. Only the rendered body gives it away.

Shape Three: The Component That Outgrew Its Page

A 25-row comparison table, written once, injected by a build step into every page that might want it. It ended up on 51 pages.

On the page it was written for, it is the content. On the 18 budget-bracket pages that also received it, it was roughly half of all the text, and it did not answer those pages' question at all. Strip the shared shell and the table away, and what those 18 pages had of their own was three sentences.

Word count never flagged them. Every one was 3,000–3,100 characters, comfortably above any thin-content heuristic you would write. The volume was real; it just wasn't theirs.

What the Fix Is

The obvious move, appending more text to the thin pages, is the wrong one, and it fails in a specific way: if every page gets the same five extra sentences, a four-sentence template becomes a nine-sentence template. The shingle count barely moves, because you added shared text to solve a shared-text problem.

What works is the opposite: give each page permission to say less. Compute a set of candidate observations from that page's own underlying data, put a threshold on each one, and emit only the ones that cross. A page with nothing distinctive in its data then says nothing distinctive, which is honest, and the pages that do have something say different things from each other, which is the entire point.

Set the thresholds before you look at the results

The temptation is to tune a threshold until a particular page qualifies. That is fitting the rule to the answer, and it produces observations that are technically true and practically meaningless. Pick thresholds that already mean something outside your dataset: a regulatory line, a standard size class, a bootstrap interval computed from resampling your own population — and then accept whatever they select. If a page crosses nothing, saying "nothing here stands out" is a real finding and reads as one.

The Number to Watch

Site-wide, that first run came back at 51.7% unique, which means slightly under half of all shingles appeared on more than one page. The per-page ranking mattered more than the total: three pages came in under 3% unique, and those three turned out to be shapes one and two, which nobody would have found by reading.

Run it on your own site before you assume the answer. It takes about a minute, it needs nothing but the built HTML, and the pages at the top of that sorted list are almost never the ones you would have guessed.