Index
All articles
Forty-two pieces across six sections, of which fourteen are selections. Every selection names the class of problem before it names a product.
For a separate operational view of time, ownership and team activity, see this Monitask guide.
Section Measurement
What measurement does to what it measures
Proxies, baselines, thresholds, retention, and what happens when the measure becomes the target.
- 6 min readEvery monitor reports a proxy, and the proxy is not the thing
- 6 min readA number means nothing until you know what normal was
- 6 min readSignal, noise, and what a threshold really costs
- 6 min readRetention: what you keep, for how long, and what that costs
- 6 min readWho sees the dashboard changes what the dashboard does
- 6 min readWhat no monitoring tool can measure, in either market
- 7 min readWhen the measure becomes the target
Section Infrastructure
What to collect, and what collecting it costs
Agents and polling, metrics against logs against traces, cardinality, sampling, and where the bill comes from.
- 6 min readPolling and agents collect different truths
- 6 min readMetrics, logs and traces answer different questions
- 6 min readCardinality is the number that decides the bill
- 6 min readSampling: what you throw away on purpose
- 6 min readSelf-hosted or managed: the work moves, it does not vanish
- 6 min readSynthetic checks and real user data disagree, and both are right
- 6 min readInstrumenting an application you did not write
Section Alerting
Alerting, on-call and noise
An alert is a request for a human. Most systems make far more of them than they can justify.
- 6 min readAn alert is a request for a human and should justify itself
- 6 min readStatic thresholds drift into noise, and why
- 6 min readAlert fatigue is a measurable failure, not a mood
- 6 min readEscalation: who is woken, in what order, and who decides
- 6 min readWhat an alert should carry with it
- 6 min readDeleting alerts is the maintenance nobody schedules
- 6 min readIncident review, and the difference between cause and blame
Section Workforce
Monitoring people: mechanism, law and trust
What is captured on the endpoint, what the law asks where, and what observation does to the work observed.
- 7 min readFour classes of workforce tool, and why the class decides the outcome
- 6 min readWhat is actually captured on the endpoint
- 7 min readNotice, consent and legitimate interest, by jurisdiction
- 6 min readActivity is not productivity, and the gap is where trouble lives
- 6 min readCovert monitoring: what it costs legally and culturally
- 6 min readRetention, access, and who inside the company may look
- 6 min readBeing watched changes the work being watched
Section Choosing: systems
Selections for infrastructure situations
Class first, then criteria, then names. Seven situations rather than one ranking.
- 7 min readWhich class does your problem belong to
- 7 min readA team of five with no dedicated on-call
- 6 min readWhen the telemetry bill grows faster than the traffic
- 6 min readOne application on one server
- 6 min readKubernetes without a platform team
- 6 min readWhen the data has to stay in-house
- 6 min readA service whose users notice before the graphs do
Section Choosing: teams
Selections for teams and time
The same discipline applied to workforce tools, where the class of tool decides the outcome more than any feature.
- 6 min readWhen somebody above you has already decided
- 6 min readContractors billed by the hour
- 7 min readA distributed team under European rules
- 6 min readWhen you need defensible records rather than oversight
- 7 min readWhen the concern is data leaving, not time spent
- 6 min readA team that will resent it, and should be asked first
- 7 min readWhat to check before signing, whatever you chose
Further context
Start from the class of problem, not the list of tools
Every selection here names the situation first and the criteria second. Product names come last, and each one carries the line describing what it costs you.