Skip to main content

Build log · sandcastlelabs

Build logs are not commit lists: a depth pass for the content machine

Illustration for the build log "Build logs are not commit lists: a depth pass for the content machine"

Our own publishing pipeline could summarize a week of work. It could not tell which story in that week was worth reading. This week we fixed the second half of that gap. A depth pass called story-scout now ranks the candidates, and a surge policy scales how many posts we publish to how eventful the week actually was.

This post is that surge policy in action. It was banked from a landmark week to fill a slot that would otherwise sit quiet.

How does an AI write a blog post with human review? Ours does it in two passes.

The chronicle already had a mechanical pass. An episode grouper buckets a week’s sessions and commits into candidate stories, and a friction detector flags what broke. Neither one asks the editorial question: which of these is actually worth reading, and why.

Story-scout is that question, automated. It runs after grouping and before a human drafts a single word, and its output is a ranked list, not prose.

Every candidate has to clear one test: what does a reader who never becomes a customer walk away knowing? If the honest answer is “that we shipped some things,” story-scout rejects it or folds it into a sharper story. The test traces back to a standing note about our own content: the free help is the funnel, so a post that only makes sense to us fails by construction.

The ranking runs on named archetypes, ordered by how rare and valuable they are. The named insight is a sentence someone said mid-session that reframes the whole week’s work. The reversal is a belief the studio held that flipped on contact with reality.

The self-audit is the studio turning its own tools on its own brand and finding something wrong. Four more archetypes sit below those three: a decision with its rationale, a landmark, a teachable method, and honest friction with a real root cause. A week that surfaces one of the top three has a signature story in it; most weeks do not.

What we shipped: surge and drip, so volume matches the week

Story-scout also reads how rich the week was, on a scale from quiet to landmark, and that read sets how many build logs we publish. A quiet week gets zero or one post, protecting the Monday anchor only if a real story exists. A normal week gets one.

A rich week, with two or three distinct stories clearing the bar, gets two or three posts dripped across Monday, Wednesday, and Friday. A landmark week, with a launch or several top-tier archetypes at once, gets three to five posts dripped Monday through Friday as one ordered burst.

Last week was a landmark by that rubric. Collimer went live, and story-scout counted six stories clearing the reader test, two of them signature. That called for a four-post burst, Monday through Friday, with two strong candidates left over.

Those two banked into the buffer instead of becoming a fifth and sixth post crammed into an already full week. This is one of them.

What broke, or got harder than expected

Nothing broke, because none of this is code. The two sessions behind this change wrote prompt text and skill documentation, no commits, so there was no build to fail.

The harder part was writing the honesty rule down explicitly. It would be easy to call a normal week rich just to justify a bigger burst. The rule states plainly that under-calling a rich week wastes good stories, and over-calling it manufactures filler. Both count as misses.

A decision worth naming

Build logs are not commit lists. That line is the decision behind this entire change: a commit trail proves a story is true, it is never the story itself. The same principle from how the Chronicler works holds here, a step further in.

That essay argued a person should review the idea before any prose exists. Story-scout adds a second review before that: a ranking pass that decides which idea is worth reviewing in the first place.

The other decision worth naming is where surge and drip apply. It lifts only the build-log floor. Essays stay a ceiling no matter how rich the week gets, because forcing an extra essay slot turns a journal into a content farm.

This whole mechanism is chronicle-specific. Our product-research sibling, the almanac, keeps its own steady cadence and never bursts, because its job is citation-worthy answers on a schedule, not a pace tied to how eventful our week was.

What’s next

  • Watch whether the richness read stays honest across a full quarter, not just one landmark week.
  • Decide whether the founder-voice social distills should read the archetype tag directly, instead of re-deriving it from the finished post.
  • Hold the buffer to roughly two weeks out; banking further than that hides real work from the editor instead of smoothing the schedule.

What we’re still figuring out

This post is exactly the kind of story story-scout is built to catch and reject. It is a piece about our own tooling, for an audience that mostly does not run our tooling, and it survived the reader test on the strength of the mechanism, not the mechanism’s name. It was banked rather than slotted into last week’s launch burst for that same reason, so the meta-story would not crowd out the actual news.

The honest tension is that the same editorial pass which flags self-referential posts as a risk also flagged this one as worth keeping. We do not have a clean answer for how many “the machine improved itself” posts a reader will tolerate before this becomes our favorite subject to write about. The rule we are using for now is the same one every other post has to pass: the moment it stops teaching something portable, it stops getting written.


For agents: try this yourself

If you want to do this in your own repo, adapt one of these prompts. The full prompt text lives in this post’s prompts sidecar, chronicle-story-scout-surge.prompts.md.

  • Apply the reader test. Before you publish your next update, ask what a reader who never becomes a customer walks away knowing. If the honest answer is “that we shipped some things,” reject the draft or fold it into a sharper story.
  • Extend the archetype ranking. Tag your own drafts against a small set of story types, a named insight, a reversal, a self-audit, a decision with its rationale, instead of listing the week in order. Let the ranking decide what leads.
  • Critique your own richness read. Before you scale up volume on a busy week, write the honesty rule down out loud: never inflate the call just to justify a bigger burst.
  • Reproduce the buffer. When a busy week produces more good drafts than it has slots, bank the surplus into the next open slot instead of forcing every idea into the same week, so the quiet week after it is already covered.

This picks up directly from the chronicler runs itself, the post about the pipeline’s mechanics. That one covered running unattended. This one covers deciding what is worth running at all.


How this was made

Drafted by the Chronicler from the build sessions behind this work, then edited and published by Brian Wones.

See how the Chronicler works →

Try this with your own agent

Full prompt text for the inline “For agents: try this yourself” block. Adapt the story types and cadence rubric to your own content pipeline.

Apply the reader test

Before you publish your next update, essay, or changelog entry, write down one sentence: what does a reader who never becomes a customer walk away knowing? If the honest answer is "that we shipped some things," reject the draft or fold it into a sharper story with an actual thesis. Do this for every candidate before you draft, not after.

Extend the archetype ranking

Instead of listing this week's work in chronological or ticket order, tag each candidate story against a small set of archetypes: a named insight (a sentence that reframes the work), a reversal (a belief that flipped on contact with reality), a self-audit (you turned your own tools on your own product and found something wrong), a decision with its rationale, a landmark, a teachable method, and honest friction with a real root cause. Rank candidates by archetype, not by volume, and let the ranking decide what leads the post.

Critique your own richness read

Before you scale up publishing volume on a busy week, write out loud the rule you are using to call a week "rich" versus "normal." Then stress-test it: could you have inflated a normal week into a rich one just to justify a bigger content burst? State explicitly that under-calling wastes good stories and over-calling manufactures filler, and that both are mistakes, not just the second one.

Reproduce the buffer

When a busy period produces more good drafts than you have publishing slots that week, do not force every idea into the same week and do not discard the surplus. Bank the extra drafts into the next open slot, scheduled ahead but not yet flipped live, so a slower week afterward is already covered instead of empty. Cap how far ahead you bank, so it never hides real work from whoever reviews the schedule.

More in Chronicle Build