- DOC
- writing/loud-failures-are-a-gift
- VER
- 1.0
- STATUS
- MAINTAINED
- LAST REVIEWED
- 2026-08
Loud Failures Are a Gift
Something broke in my first Marketo instance at 6am on a Tuesday. A smart campaign hit a bad token and the send stopped. I found out because my phone made a noise I still have a physical reaction to, years later, in unrelated contexts.
I want to argue that this was the good outcome.
A loud failure hands you everything. A timestamp, a component, usually an actual sentence about what went wrong. It has done most of the diagnostic work already, and it did it inside ten minutes, while the blast radius was still one campaign instead of one quarter.
The failures I've spent the last few years cleaning up don't do any of that.
Everyone told the truth
Here's the shape of it. A field stops arriving in the CDP because a permission changed somewhere upstream. The connector doesn't throw an error, because from where the connector sits nothing went wrong — it synced every field it could see, and it just couldn't see that one anymore. Mapping doesn't complain either, because it mapped everything that arrived. The segment builds. The flow runs. The campaign goes out to sixty percent of the people it was meant to reach.
Every layer reported success. Every layer was being honest.
That's what makes these miserable to unpick. There's no bug to find. Each component answered correctly about its own job, and the question of whether the chain still does end to end what it did last month belongs to nobody.
So it doesn't reach you as an error. It reaches you as an opinion. Someone says email feels soft this quarter. Someone else thinks that segment looks smaller than it used to. You get a Slack message with a question mark in it. All of this lands six weeks after the permission changed, having passed through three people's hunches on the way, and now you're doing archaeology instead of debugging.
Monitoring is the wrong word
The usual response here is that we need better monitoring. I've watched that sentence eat a lot of budget, and I don't think it's right.
Monitoring catches departures from normal. When a field quietly stops arriving, the new state becomes normal on day one. Nothing spikes. No connection drops. The row count is lower, but lower than what? Lower than a number nobody ever wrote down.
You can't catch a silent failure without a prior claim. Something has to have said, ahead of time, what was supposed to be there.
Write it down while it's still true
What actually works is dull. Write down what correct looks like while it's still correct. I keep a short list: how many rows the account object had last time I looked, what share of records are missing the two fields our segmentation actually leans on, roughly how many distinct identity keys we should be seeing.
It doesn't need to be a dashboard. A number in a document with a date next to it will do, provided a person opens the document on purpose.
Then check it on a schedule you don't get to skip. Not because you expect a problem, but because the failure we're talking about is specifically the kind that never tells you to look.
Which is roughly why an earlier entry here is called Measure Twice, Cut Once. Measuring isn't the tedious bit before the real work starts. When every component reports success independently, the measurement is the only thing you have that runs across the whole chain.
The job changed underneath us
There's a version of this work where you're excellent at fixing things that announce themselves. I was good at that job and I liked it. It's less and less the job.
The stack got quieter as it got more distributed. More boundaries between systems means more places where a handoff can succeed at being a handoff while failing at being useful, and every one of those boundaries has a person on each side who is right about their own half.
So I've mostly stopped asking teams whether anything is broken. The answer is always no, and it's always accurate, because nothing is erroring. I ask when they last checked that a number was what they expected. The answer to that one takes a while.