Skip to main content

Build log · sandcastlelabs

The Studio Built a Second Content Machine Instead of Stretching the First One

·

Illustration for the build log "The Studio Built a Second Content Machine Instead of Stretching the First One"

What actually running an agent-native studio looked like this week

Sandcastle now runs on two separate writing machines, not one. This week we shipped the second one: a founder-content-plan that publishes buyer-facing guides under real names, kept deliberately apart from the Chronicle you’re reading right now. The two want opposite things from the same sentence. The Chronicle wants “this week we shipped X.” The founder-content-plan wants a timeless answer to a question someone actually typed into a search box or an assistant. A sentence that satisfies one almost always fails the other, so instead of teaching the Chronicle to do both jobs, we built a second machine and gave it its own rule: anything that opens “this week we…” belongs back here.

What broke (or got harder than expected)

The two founder guides that opened the new pipeline, one under Nick’s byline and one under Brian’s, both came back from the writing agent’s first draft with numbers a human hadn’t checked yet. Nick’s guide claimed a measurement confidence interval of “roughly ±6, dated June 2025.” Neither number was real. Both drafts also carried an invented “Last updated: 2025-07-14” line under the headline that nothing in the repo could account for.

Neither guide shipped with those numbers. Both PRs itemize the fix: the fabricated interval and dates were replaced with the real, sourced ones, and the invented last-updated line was deleted before either founder signed his name to the page. That’s what the checklist we’re now calling the editor’s hour is for, and this week is the first time it ran on content meant for a stranger typing a question into ChatGPT rather than a reader following our own build log.

A decision worth naming

The same anonymized devtools receipt anchors both founder guides: citation rate rose from 1.25% to 5.83% in six weeks after a README and docs wave landed. That number is real, it is the same evidence we have cited before, and it is the one piece of proof both guides share, because the guides are answering different questions with the same underlying fact.

The harder decision was the split itself. We could have added a “guides” section to the Chronicle and pointed the same machine at buyer questions. We didn’t, because the two jobs actively fight each other in a single sentence: point-in-time versus timeless, narrating the work versus answering a query, “we” versus “the answer is.” The founder-content-plan gets its own rule instead: one lane per founder, so two guides never compete for the same search query, and no fact in a guide without a source, which is the same rule the Chronicle already runs on, just enforced for a colder, less forgiving reader. It’s the same question this arc opened with, asked of a second machine instead of the first one: can it run on rules alone, without someone watching every sentence.

The handoff wasn’t clean everywhere. Brian’s PR for the second guide notes that the plan’s own tracking table, the one that would mark those two guides as shipped, exists only as an uncommitted edit on a different branch, not on the main line. A small gap, and an honest one: the guides are live, the record of them being live is not yet where it belongs.

What’s next

  • Flip the founder-content-plan’s F1 and F2 tracking rows once the table that holds them actually lands on the main branch.
  • Watch whether the “this week we” test still catches every borderline sentence once a third byline joins Brian and Nick.
  • Decide whether Chase’s author page ships this week now that his bio and headshot are in hand, per the plan.

What we’re still figuring out

We don’t know yet whether one lane per founder holds once the founder-content-plan is publishing three or four pieces a week instead of two. The rule is easy to state and easy to follow when there are two guides and two names. What we haven’t tested is whether it survives volume, or whether “this week we” turns out to be a fuzzier line than it sounds once more people are writing against it under real deadlines.

For agents: try this yourself

If you want to do this in your own repo, adapt one of these prompts. The full prompt text lives in this post’s prompts sidecar, second-content-machine.prompts.md.

  • Apply the “this week we” test to your own content backlog. Read through your last ten published pieces and flag every sentence that only makes sense as of the day it was written; that’s the tell for which pipeline a piece actually belongs to.
  • Critique whether your own generated drafts get an editor’s hour before a real name goes on them. Name the specific failure modes (invented dates, fabricated stats, made-up “last updated” lines) your last AI-drafted piece would have shipped with unchecked.

How this was made

Drafted by the Chronicler from commits and PR bodies on 2026-08-17 to 2026-08-23. Edited and published by Brian.

See how the Chronicler works →

Try this with your own agent

2 prompts you can hand to your own agent (or run by hand) to work with what this post documents. Edit the bracketed parts for your context.

Apply the “this week we” test to your own content backlog

Read through our last ten published pieces of content (blog posts, guides, docs, whatever we publish regularly) and flag every sentence that only makes sense as of the day it was written -- anything that implicitly says "this week" or "recently" or "we just." Group the flagged pieces by whether they're journey content (narrating our own work) or evaluation content (answering a reader's question). If any single piece mixes both, that's the piece to split first.

Critique whether our own generated drafts get an editor’s hour before a real name goes on them

Take the last AI-generated draft we published or almost published under a real person's byline. Go back through it line by line and name every specific failure mode it would have shipped with if nobody had checked it: invented dates, fabricated statistics or examples, made-up "last updated" lines, unverified claims about competitors or vendors. Then argue whether our current review process would actually have caught each one, or just some of them.

More in Chronicle Build