The rule carrying this template is a ban on the word caused. It may report that one route holds fourteen thousand of nineteen thousand errors, and it may not say why the number moved, name a deploy, point at a customer or blame a dependency, however obvious any of that feels. Codex sticks to concrete, well scoped output, which is what keeps a breakdown from quietly becoming a theory.
Run a staged conversation — no account needed. The agent handles it for real while a simulated world answers its tool calls; nothing touches real accounts, and nothing is actually sent.
checkout p99 over 800ms
The trigger evaluates P99(duration_ms) over five minutes against a threshold of 800. The evaluation that fired returned 2,140 at 14:07 UTC, against 310 at 13:30. The trigger has no group by of its own. The dataset carries service.name, http.route, tenant_id and app.cache_hit.
Using this template drops you into a guided setup. It asks exactly this, nothing else:
Connect Honeycomb
One sign-in. The agent acts through your account, scoped to what this template uses.
Connect Slack
One sign-in. The agent acts through your account, scoped to what this template uses.
Triage notes
How you want a fired trigger written up - the field worth grouping by in each dataset, the triggers you already know are noisy, plus the opening move you expect on a latency alert, on an error rate alert, and on a query that comes back with nothing in it.
Runs on Codex
Preselected for this page — connect your Codex account during setup, or switch to NoClick's built-in models with one click.
Watch it handle a test run
A staged conversation against a simulated world — then it’s live.
Which group holds the change is retrievable. Why it moved is not, and the moment a note starts answering the second question people stop verifying the first.
Fourteen thousand of nineteen thousand rather than most of them. Exact numbers are what a thread can argue with, and rounding is where a small regression starts sounding like an outage.
It keeps reading rather than reporting a partial figure, and a query that does not complete is written up as exactly that. The one number it deliberately never trusts is the value inside the notification, which was true one evaluation earlier and is regularly the thing that sends an investigation down the wrong path.
Free to start. Guided setup, a test run against staged conversations, and it's live.