My alert was correct every morning for four weeks, and that is why I stopped reading it

5 min read AutomationAlertingDeveloper tooling

A daily report told me to run three updates. It told me that in fourteen reports across twenty eight days, word for word, and it was accurate every single time. One of the three was genuinely worth doing and I never saw it, because the other two were a permanent false alarm produced by conditions the report had no vocabulary for.

TL;DR · THE FIX

An alert has three states, not two: actionable, impossible, and decided against. Miss the third and every settled decision reprints as an open task forever, which trains you to ignore the block it lives in and buries the one line that is real. The fix is not better detection, because detection was never wrong. It is giving the report a way to say 'I know, and we decided not to', checked before anything else and reported in the body rather than emitted as an action.

The symptom

A scheduled script mails me a status digest every morning. One block in it audits third-party tooling for drift against upstream, and for four weeks it said this:

SKILL DRIFT (3rd-party plugins; review/update via /skill-audit):
- npx skills update <first> -g -y   (drifted)
- npx skills update <second> -g -y  (drifted)
- npx skills update <third> -g -y   (drifted)

The same three lines in the same order, in fourteen digests between 3 August and 30 August. I had long since stopped reading that block, and one of the three lines was real: a one-command fix, about thirty seconds, sitting in plain sight for four weeks inside a block I had trained myself to skip.

What I tried first

The instinct is to reach for the alert threshold or start suppressing things. Instead I read what produced each of the three lines, and they turned out to be three different situations printing through one code path.

The first was real and fixable: genuine drift that one command closes. This was the only line that deserved to be there.

The second was real and impossible. The package manager installs a published release, while my auditor diffs against upstream main. When a publisher is behind their own main branch the gap is real and no command closes it. I verified that: after both an install and an update, the CLI answers “All global skills are up to date” while main is still ahead.

The third was real, fixable, and deliberately not done. One command would have closed the drift, and I had run that command, then undone it on purpose because nothing in my repo uses that dependency. The next morning the report asked me to run it again.

The second case had already been handled some time ago for exactly this reason: it gets reported in the body and never becomes an action line. So the shape of the fix already existed in the codebase, with one hole in it.

What was happening

The report knew two states, “do this” and “nothing to do”, and a decision I had already made was neither. I had chosen against it, so it was not actionable. The drift was there and the checker could see it, so it was not nothing. With nowhere else to go, it came out as a command to run, every morning, forever.

Nothing the report printed was false. Every line described drift that existed. The defect was in what the alert was able to say, which is why tuning the detection would not have helped.

The second-order cost is what bites. A permanent false alarm does not stay contained to its own line. It teaches you the whole block is noise, and then it takes the real line down with it.

The fix

A HELD category, checked before the impossible one:

# Skills whose drift is REAL and fixable but which we have DECIDED not to update.
# Checked before UPSTREAM_LAG. A held skill is reported in the body and never
# emitted as an action line: re-printing a command we have already decided not to
# run is the same permanent false alarm UPSTREAM_LAG exists to kill, and it buries
# the one line that is genuinely actionable.
HELD = {
    "<name>": "held 2026-08-28, reaffirmed 2026-08-30 - nothing in the repo uses it. "
              "The update WAS applied and then REVERTED deliberately; the installed "
              "copy is pinned to <hash>. The whole gap is 5 lines across two docs.",
}

and at the branch:

elif dirs_differ(upstream, installed):
    if name in HELD:
        # Real, fixable drift we have decided against. Report, never action.
        lines.append(f"- {name}: drifted, HELD BY DECISION ({HELD[name]})")
    elif name in UPSTREAM_LAG:
        # Real gap, but no command closes it. Report, never action.
        lines.append(f"- {name}: behind upstream main, NOT actionable ({UPSTREAM_LAG[name]})")
    else:
        lines.append(f"- {name}: **drifted** - UPDATE AVAILABLE")
        action.append(f"npx skills update {name} -g -y   (drifted)")

Two details in there matter more than the dict itself. The reason is stored with the hold, because a bare suppression list is a booby trap: six months from now nobody knows whether an entry is a considered decision or someone silencing a warning they did not understand. Each entry carries the date, the reasoning, and what was measured. And a held item is still reported. It appears in the body with its reason, because suppressing it entirely would be a different bug with the same ending, a report quietly lying about the state of the world. Appearing in the report and being emitted as an action are different things, and only the second is a claim on your attention.

The measurement

Read off the reports themselves:

ReportAction needed
29 August (before)3 lines
30 August (after the category)1 line (the real one)
31 August (after running it)none

The remaining line was the fixable one. I ran it, the update applied, and a recursive diff against upstream went silent. Every digest since has read none: catalogs, plugins, and skills.sh installs all current.

One smaller result looks like a failure and is not. One of the held dependencies is still drifting today, and that is correct: “still drifting” in a report can be a decision working as intended. If your alerting cannot express that difference, it will keep asking you to undo your own decisions.

What to take from it

If a line in your alerting has never once cleared, ask which of three states it is in and whether your report can say so. Actionable means something changed and here is what to do, and it is the only state that has earned a place in an action list. Impossible means real and outside your control; report it so you know the state of the world, never as a task, because a task nobody can do is a standing accusation. Decided against means real, fixable, and you already chose; report it with the reason and the date, never as a task, or every settled decision comes back forever.

Miss the third state and you get a block of the report that everyone has agreed to stop reading, with something real inside it.

Related fixes

Discussion

Powered by GitHub. Sign in to leave a comment.