Skip to main content

Build log · collimer

Our own panel has a 7-point noise floor

·

Illustration for the build log "Our own panel has a 7-point noise floor"

Between August 4 and August 18, 2026, the score for our own site on Collimer’s 100-prompt index panel moved from 10.09 to 11.00. Over the same runs, each run’s confidence interval had a half-width of 5.77 to 8.46 points. That is the noise floor: about 7 points either way, on a score that barely moved. A two-point change on this panel is not a result.

What does building a GEO product in public mean when our own instrument can’t see two points?

It means the number that limits our own claims goes in public too. We publish a confidence interval next to every score, because a single run has no error bar and AI engines do not return the same answer twice. The interval only helps if we know how wide it really is.

The first draft of our verification guide, written by our own content engine, quoted an interval of plus or minus 6 and dated it June 2025. The editor’s check found the figure was made up, and the draft’s example customer result was invented too. Both were replaced. The interval came from a production pull on August 20 of the last twelve completed runs of the index panel, so the interval in the guide traces to rows in our database.

How wide is the noise floor on our own panel?

Window (2026)Panel sizeInterval half-width per runScore movement
July 19 to 31285 prompts3.42 to 3.74 pointsnot reported in the pull
August 4 to 18100 prompts5.77 to 8.46 points10.09 to 11.00

In a production pull dated August 20, 2026, the run-to-run half-width of the confidence interval on Collimer’s 100-prompt index panel ranged from 5.77 to 8.46 points, while the point score moved less than one point.

Fewer prompts should mean a wider interval, and the drop from 285 prompts to 100 points the same way as the numbers. We did not separate that from anything else that changed between the two windows. We publish the range and not one figure. A reader deciding whether their own two-point move counts deserves to see both ends.

What did the noise floor change about what we say?

Our methodology page says a move from 47 ±8 to 50 ±8 is noise, a move to 58 ±7 is signal, and that we will never report a single fixed visibility number as if it were exact. The ±7 and ±8 in those examples sit inside the range we measured.

The same page says a shipped fix gets re-probed against the same question set and reported in one of five states: not shipped, pending, verified moved, no measurable change, or moved against. That is a reporting policy. A null result on a frozen panel is a finding, and we publish it.

For a reader with no tool, the guide describes a hand method: ten prompts, three engines, three runs each, and treat anything that changes in fewer than two-thirds of runs as noise. The guide calls the two-thirds line a practical proxy. We have not checked it against an interval, and it does not replace one.

A decision worth naming

We could have kept the first draft’s tidy plus or minus 6. We chose the pull, which was wider and less flattering, because the invented figure had the same shape as the two fixes we corrected in August: a claim that cited its own rationale and had never been checked against the data. We also wrote the method up as a guide.

What we’re still figuring out

The pull covers runs through August 18, the day we merged a fix that removed duplicate prompts from generated panels. Duplicates inflated the prompt count and narrowed the published interval, as the week our number turned out to be lying five ways describes. If the index panel carried duplicates, the intervals above are narrower than the data supports and the real floor is wider than 7. We have not checked, and this post has no pull from after the fix.

It is also one panel, on one brand, with a score near 10. We have not measured how the floor moves at other score levels or panel sizes, and the records we checked do not say why the panel shrank from 285 prompts to 100. The next step is to re-pull the intervals for runs after August 18 and check the frozen panel for duplicate prompts before we quote a number again.

For agents: try this yourself

If you track a score that comes from sampled model answers, adapt one of these. Full prompt text lives in this post’s prompts sidecar, our-own-panel-has-a-seven-point-noise-floor.prompts.md.

  • Apply the noise-floor pull. Take the last twelve completed runs of one fixed panel and compute the run-to-run interval half-width for each. Report the range and the score movement over the same runs, not a single figure.
  • Critique the claim. Take a sentence such as “our score rose two points after the fix” and list what you would need to know about the panel, the interval, and the run count before believing it.

How this was made

Drafted by the Chronicler from Collimer commits, a merged pull request, and a published guide dated 2026-08-20 to 2026-08-24, then edited and published by Brian.

See how the Chronicler works →

Try this with your own agent

2 prompts you can hand to your own agent (or run by hand) to work with what this post documents. Edit the bracketed parts for your context.

Apply the noise-floor pull to your own panel

We track [score name] from a fixed set of [number] prompts sent to [engines]. Here are the last twelve completed runs, each with its point score and its confidence-interval lower and upper bounds: [paste rows]. For each run, compute the interval half-width. Report the minimum, maximum and typical half-width, and the point-score movement across the same runs. Then tell me the smallest score change I should treat as larger than the noise, and say what I would need to change (prompt count, repeats per prompt) to shrink that number. If any of the runs share duplicate prompts, flag them, because duplicates narrow an interval without adding information.

Critique a score-change claim before you repeat it

A report says: "[claim, e.g. our score rose two points after we shipped the fix]". Before I repeat it, list what I need to know to believe it: whether the prompt panel was frozen between the two measurements, how many runs each number comes from, the interval around each number, whether the two intervals overlap, and whether any engine dropped out of one measurement. Then classify the result as one of: not shipped, pending, verified moved, no measurable change, or moved against, and say which evidence would change your classification.

More in Collimer Build