Skip to main content

Build log · collimer

The Week the AI-Visibility Number Turned Out to Be Lying in Five Different Ways

·

Illustration for the build log "The Week the AI-Visibility Number Turned Out to Be Lying in Five Different Ways"

Collimer’s whole pitch is a trustworthy AI-visibility number. In one four-day window, the engineer who owns it found five separate, unrelated ways that number had been quietly wrong, and fixed all five before a customer ever saw one of them move.

Building a GEO product in public means admitting when your own score is wrong

None of the five bugs crashed anything. Nothing paged anyone. Every one of them shipped inside code that compiled, passed review, and kept running exactly as designed, just designed around an assumption that wasn’t true. A passing test suite proves the code runs. It does not prove the metric means what the product’s own marketing says it means, and this week was proof that the gap between those two claims is wider than “measure honestly” sounds like it should allow, the same gap a database-pool fix and an attribution default fell into two weeks earlier.

Five different ways the number was wrong

The most visible one lived in the confidence interval Collimer publishes next to every score: duplicate rows in the underlying prompt panel had been inflating the published Wilson interval, making the number look more certain than the sample it came from actually supported. A related but separate bug hit Share of Voice, which was supposed to group a brand’s mentions together but grouped by the wrong key instead, splitting one real brand’s mentions across its own aliases and reporting them as smaller, separate competitors. Citation-host normalization had the same shape one level down: several near-identical copies of the same “is this the same site” logic had drifted apart from each other, so which citations counted as duplicates depended on which copy of the code happened to run that day.

The other two lived outside the score itself. A repeat free-scan submission ran the probe, charged the customer, and never mailed the report, silently, for eight days, because a dedup gate meant to stop double-billing a true duplicate had been written to apply to every completed scan in a rolling thirty-day window instead of just the actual repeats. And brand accuracy had been folding two different failures into one number: a brand the model had genuinely never heard of scored identically to a brand the model described incorrectly. The fix excludes “unknown” from the denominator, and it moved the score most for the small brands the old version had been quietly hurting worst.

A decision worth naming

Every one of these got treated as a correctness bug in the product’s core claim, not a cosmetic scoring tweak, and all five got fixed in the same four-day window that also shipped seven releases, v1.11.0 through v1.17.0, on the same release discipline built two weeks earlier. The same window added a release-boot smoke gate to CI, so a production release boots itself before Fly does, and started posting each release’s own summary to the team’s Slack channel, so a red release or a silently failing scheduled workflow can no longer just sit there unnoticed. Speed and honesty about being wrong shipped together this week, not one after the other, and that pairing was a choice, not an accident: nobody deferred the scoring fixes to “after the release push.”

What’s next

  • Go back through other scoring paths beyond the five found this window and ask the same question: does a passing test prove the metric is right, or only that the code runs.
  • Decide whether the Wilson-interval and Share-of-Voice fixes need a customer-facing note, since both could have changed a number a design partner was already looking at.
  • Watch whether the new release-boot gate and Slack TLDR catch the next regression before a customer does. That is the real test of whether this week’s fix generalizes.

What we’re still figuring out

We don’t know whether five is the whole list or just the five that were easiest to notice. None of them was found by a customer complaint or an alert firing; each one surfaced because someone went looking, asked why a number would ever move that much, and kept pulling the thread. We don’t yet have a standing process for going looking on a schedule, only the instinct that surfaced it once, and the honest next step is turning that instinct into a habit before the next quietly-wrong number sits there for eight days.

For agents: try this yourself

If you own a metric or score your product publishes, adapt one of these. The full prompt text lives in this post’s prompts sidecar, the-number-was-lying-five-ways.prompts.md.

  • Apply the duplicate-row audit. Find one statistic your product publishes that’s computed from raw event or panel rows. Check whether a dedup step actually runs before the calculation, the way Collimer’s Wilson interval was inflated by duplicate prompt-panel rows nobody had deduped.
  • Reproduce the unknown-versus-wrong split. Find a metric in your own product that averages a “not applicable” case together with a genuinely wrong case, the way brand accuracy folded “the model never heard of you” into the same bucket as “the model described you wrong.” Calculate how much the score moves once the two are separated.

How this was made

Drafted by the Chronicler from commits, releases, and CHANGELOG entries on 2026-08-17 to 2026-08-20. Edited and published by Brian.

See how the Chronicler works →

Try this with your own agent

2 prompts you can hand to your own agent (or run by hand) to work with what this post documents. Edit the bracketed parts for your context.

Apply the duplicate-row audit to a metric you publish

Find one statistic our product publishes to customers that's computed from raw event, log, or panel rows rather than a pre-aggregated table. Trace the calculation back to its source query and check whether a deduplication step actually runs before the statistic is computed, or whether it's silently assuming every row is unique. If duplicates are possible, calculate roughly how much they could be inflating or deflating the published number.

Reproduce the unknown-versus-wrong split on our own scoring

Find a metric or score in our own product that currently averages two conceptually different outcomes into one number -- for example, "we don't have data" scored the same as "we have data and it's bad." Identify the specific case, then recalculate the metric with the "no data" case excluded from the denominator instead of counted against it, and report how much the score moves and for which segment it moves the most.

More in Collimer Build