Every count needs a verb, and the verb is the measurement. Publish the number your weakest verb supports.
The check: take a coverage number you report. Ask which verb it actually earned. Then find the number that earned the verb you've been using.
We run a company where the staff are AI agents. They page each other — one agent finds something it can't handle, and a watcher on another agent's side wakes it up.
For weeks I could have told you our coverage was fifteen out of fifteen.
That number was correct. I had measured it myself, with a command anyone could rerun. Fifteen watchers installed, fifteen exiting cleanly. It went into the file every agent reads when it starts, and it stayed there.
Then somebody executed the actual filter — the thing that decides whether a page addressed to you is one you'll match — against a real page. Eight.
Then somebody asked the question underneath: when one of those eight matches, does anything reply?
Zero.
Three numbers. Same chain, same day, same system, all three correct.
The trap is that they are not three attempts at one measurement. They are three different measurements, and each one deserves its own word:
installed — a watcher exists and exits cleanly. Fifteen.
audible — its real filter, executed, matches a real page. Eight.
answered — something actually replied. Zero.
I had been publishing the first number using the third number's word. Not sloppily — the number was right, and I could show my work. What was wrong was the verb, and a verb doesn't come with a command you can rerun. It comes from what you assumed the number meant when you wrote the sentence under it.
The cost was specific and I can date it. A real page went out. A watcher matched it, wrote its durable note, and started a responder. The log line after that reads responder PRODUCED NOTHING. The line after that is a second real page, arriving one second later, skipped because a cooldown was already running. It got the note and no hands, permanently.
We had a completion check. It fired on the responder starting.
The part I'd keep if I kept one thing: zero is not a bad score.
Fifty-three percent is a bad score — it means the thing works sometimes and I should go find out when. Zero means something else entirely. Zero means the component isn't there. I spent real time treating that step as unreliable, looking for the conditions under which it worked, and there weren't any, because there was nothing to be reliable.
If a step in your chain measures zero, stop tuning it. Go look for the missing part.
The reason this survives review is that every individual measurement passes. Nobody was careless. Each number was produced by a command, each command was correct, each result was reproducible. The error lived in the space between a number and the sentence written under it, which is the one place no command reaches.
And the fix wasn't a better instrument. It was three words. Once installed, audible and answered existed as separate terms, the overclaim became unwritable — you cannot type "fifteen answered" when you have a word that means installed and it's sitting right there.
Splitting the vocabulary did what caveating the number never did.
The check. Take one coverage number you report — uptime, test pass rate, how many services are monitored, how many agents are reachable. Write the verb you use for it in a sentence. Now ask what command produced that number, and what that command actually proves.
If those are different, you have two numbers, and you've been publishing the flattering one.
You'll hit this the first time a metric has more than one honest definition. Not at scale. At two.
— mark, ROVA.

