Skip to main content

Build log · collimer

The TLS setting that stopped being safe when the database moved

·

Illustration for the build log "The TLS setting that stopped being safe when the database moved"

A TLS setting can be correct for years and then become a live hole the moment the infrastructure underneath it changes, a database migration, a new transport, anything that alters what was actually providing the trust. Nothing about the code will tell you when that happens. It still compiles. It still connects. It still works, in every sense the deploy pipeline checks for.

Why did the TLS setting stop being safe after the database migration?

Because its safety was never coming from the setting itself. It was borrowed from the transport underneath it, and that transport got replaced without anyone going back to see what the setting had been quietly depending on. The database connection here used verify_none, which accepts a certificate from anything that answers on the far end. On its own, that is a real gap. But it wasn’t operating on its own: the connection ran over an authenticated WireGuard tunnel to Fly Postgres, so the tunnel, not the certificate, was doing the actual trust work. TLS was belt-and-braces on top of a connection that was already authenticated a different way. verify_none was a defensible second layer for that one specific architecture.

What changed under the setting, without changing the setting

The database moved to Neon, reached over the public internet instead of an authenticated tunnel. No diff touched verify_none itself; nothing about the setting’s own text changed. But verify_none accepts a certificate from anything that answers on the far end, and once the transport underneath it was ordinary public internet, that meant the encryption authenticated nobody. Anyone positioned on the path between the app and the database could present any certificate and be accepted, and from there could take the database credentials and read every row in transit: users, public scans, Stripe webhook payloads. The setting’s text never moved. The thing that had been making it safe was gone.

The fix, and the line that’s load-bearing rather than boilerplate

The connection now sets verify_peer with :public_key.cacerts_get(), which reads the OS trust store the runtime image already installs via ca-certificates in the Dockerfile, so there’s no certificate bundle to vendor and nothing to rotate by hand. The part that looks like standard hardening boilerplate isn’t: customize_hostname_check has to be set explicitly, because OTP’s default hostname match rejects the wildcard certificate Neon actually serves. Without it, every connection would fail the moment verify_peer turned on. :https is the match function that implements the same wildcard rules browsers already use. Server name indication stays derived from the URL host rather than hardcoded, which is also what lets Neon route to the right compute, so an earlier switch to Neon’s direct endpoint needed no change here.

The failure mode is now the opposite of before: if the certificate chain the database serves isn’t anchored in the image’s trust store, the app does not boot. It fails closed instead of connecting anyway.

The checklist item most migration runbooks don’t have

The standard migration checklist asks whether the new thing works: does the app connect, does the query run, does the dashboard turn green. It rarely asks the question that mattered here, which is which of the settings you’re keeping unchanged were only correct because of the infrastructure you just replaced. A setting that predates the migration and never throws an error is invisible to the checklist, because “still connects” and “actually authenticates the other end” produce the exact same green check. This is the same shape of failure as three mechanisms reporting success while verifying nothing, and the same reason checking the work matters more as more of it ships unverified by default: a system that crashes gets noticed immediately, but one that keeps compiling while quietly trusting the wrong thing can run that way for a long time. The actionable version of the checklist item: after any infrastructure migration, grep for every trust or auth assumption that predates it, not just the ones that throw an error if they’re wrong, and ask what was actually making each one safe.

What we’re still figuring out

This closes the one connection we found stale. We haven’t gone back through the rest of the pre-Neon configuration asking the same question, whether some other setting’s safety was similarly borrowed from the WireGuard-era architecture and never re-verified after it changed. That audit is the honest next step, and it hasn’t happened yet.


For agents: try this yourself

If you’ve migrated a database, message queue, or any service that changes its own trust boundary, adapt one of these. The full prompt text lives in this post’s prompts sidecar, the-tls-setting-that-stopped-being-safe.prompts.md.

  • Reproduce the inheritance audit. For every trust or auth setting in your infrastructure, write down what actually provides its safety: the setting itself, or something underneath it, like a tunnel, a network boundary, or a firewall rule. Then check whether that underlying thing has changed since the setting was written.
  • Critique your own post-migration configs. List every infrastructure migration your team completed in the past year. For each one, check the security-relevant settings that didn’t change in the diff, TLS verification, IP allowlists, auth modes, and ask whether their justification depended on the infrastructure you just replaced.

How this was made

Drafted by the Chronicler from the build sessions behind this work, then edited and published by Brian Wones.

See how the Chronicler works →

Try this with your own agent

2 prompts you can hand to your own agent (or run by hand) to work with what this post documents. Edit the bracketed parts for your context.

Reproduce the inheritance audit

List every trust or authentication setting in this codebase's infrastructure config (TLS verification modes, IP allowlists, auth handshakes, credential scopes). For each one, write down what actually provides its safety: the setting's own logic, or something underneath it, like a tunnel, a private network boundary, or a firewall rule it assumes is in place. Then check whether that underlying thing is still true today, and flag any setting whose safety depends on infrastructure that has since changed or been removed.

Critique your own post-migration configs

List every infrastructure migration this team has completed in the past year (database, hosting provider, VPN, message queue, load balancer). For each one, diff the old and new config for security-relevant settings that did NOT change (TLS verification mode, IP allowlists, auth modes, credential handling). For each unchanged setting, check whether its original justification depended on the infrastructure that was just replaced, and report any that need re-verification.

More in Collimer Build