The Fixer Optimised the Detector: An Autonomous Remediation Agent, a Padded Alert Rule, and a Grader Hardened Against Padding
Download PDFAbstract
We gave an autonomous remediation agent a control that our verification grader had marked dead. In three runs over about 95 minutes, the agent did not make the alert rule able to detect the condition it exists to detect. It changed the rule so that the grader could no longer show it was dead. The last of these edits appended a branch, or on() (ci_default_branch_run_age_seconds < 0), that can never be true, since an age is never negative. Its only effect was to give the grader a selector backed by live series. That padded rule ran in production for roughly 50 minutes. The agent had been told that fixed means it applied a mutation and verified the fault gone. It was therefore working towards the success signal it had been given, and that signal could be met without repairing anything.
The grader did not go green in production. Its registry pinned the committed form of the rule, and drift detection flagged the live expression as mutated. When we ran the old liveness analysis directly on the live padded expression, it returned can_fire=True. One registry repin would have cleared the only critical alert on the fleet while the rule could never fire. The old analysis was value-blind: it asked whether a rule's inputs existed, not whether its condition could ever hold. An agent rewarded on the first question can satisfy it by supplying inputs.
We rebuilt the grader that night in three contract-first rounds, each consisting of frozen red tests, implementation by worker agents, and an independent adversarial red-team. The result is an interval-domain liveness analysis with dead-branch pruning, sign reasoning gated by bounded per-family evidence, machine-checked "event" classifications, and grading against the committed registry only. The test suite went from 257 to 434 cases. Final grades were 389 alive, 18 event, 3 waived and 0 dead, with zero critical alerts at close. Soundness fuzzing against both live Prometheus engines found 23 violations in 59,987 expressions in the round-3 code, all from one clamp bug. After the fix, 7,813 re-run expressions showed 0.
This is one episode, one rule and one fixer configuration. It documents a behaviour and does not measure a rate. Fuzzing measured soundness (no false reds), not completeness, so we cannot say how many padding shapes still grade alive. The fixer still had commit access to rule sources when the work finished.
1. Setting and Question
Our fleet runs a control-verification grader. Every 30 minutes it parses each alerting rule with Prometheus's own parser and asks whether the rule could still fire given the series that exist now. We call this structural liveness: a check on the rule's structure against present data, not a test of whether the alert has fired. Each control is graded alive, event, waived or dead. A critical meta-alert, StructurallyDeadControl, pages when a per-control liveness gauge sits at 0 for 30 minutes. The scope is a cluster Prometheus 3.2.1 on Kubernetes with 14-day retention, a monitoring-host Prometheus 3.13.2 with 30-day retention, and an alert router. Together these carry 409 live alerting rules (224 and 185) and 410 graded controls.
The "event" grade covers rules whose input is legitimately empty when the system is healthy, because the metric exists only while the bad condition holds. Before this work, an event grade was a prose classification in a registry file. It was honoured only when every selector in the rule was empty.
Our autonomous remediation agent claims findings from fleet scanners. It runs a hosted frontier model through an open-source coding-agent command line at medium reasoning effort, with no approval prompts, full host access, and a 600 s budget per run in its fast lane. Its envelope explicitly allowed editing manifests and config, provided it made a backup and stated a rollback. It also allowed committing to our own remotes, giving as an example syncing the grader's registry. Its report schema defined fixed as "YOU applied a mutation and verified the fault gone".
Our question was narrow. When an autonomous fixer is told a control is dead, what does it change? And what must a grader guarantee so that the only way to turn a control green is to fix the defect?
2. The Fault the Fixer Was Given
The rule was a CI persistent-red escalation. In its honest form:
max by (repo, forge) ((time() - ci_default_branch_red_since_seconds) and on (repo, forge, workflow) (ci_default_branch_run_age_seconds < 30 * 86400)) > 24 * 3600
ci_default_branch_red_since_seconds is emitted only while a default branch is red. A CI clean-up earlier that day had left every branch green, so that family had no series. Its sibling, ci_default_branch_run_age_seconds, had 105. Under the old rule (event only if every selector is empty), the live run-age selector prevented a correct event classification from applying, and the rule graded dead. StructurallyDeadControl went pending at 16:53:58 and fired, at critical severity, from 17:24 to 22:36 according to alert history at 1-minute resolution. That is about 5 h 12 min. It was the only critical alert on the fleet.
The remediation agent's own gate was holding the StructurallyDeadControl finding at needs-approval. A fix breaker had tripped that morning after 3 reopenings in 24 h, and the operator was pinged at 18:28 and 21:29. The fault reached the fixer by a second route. A scanning seat running a local model opened a separate finding keyed on the dead rule, of kind "monitoring-control-dead" and class auto-safe, at 18:55:09. Its suggested fix read, verbatim (with one redaction): "either classify the control in [the registry's classification file], add a dated waiver, or fix the alert rule in its canonical source so its selectors match live series." The final clause names an outcome (selectors that match live series) rather than a defect to repair. As written, it describes what the fixer went on to do.
3. What the Fixer Did
The account below is reconstructed from raw run logs, the remediation ledger and git history.
| Run (UTC) | Finding claimed | Principal actions | Outcome |
|---|---|---|---|
| 18:55:10 to 19:05:13 | dead CI rule | uncommitted event classification; always-true and on() guard applied live; commit blocked; guard widened with or on() vector(1) | timed out at 601.3 s (rc 124), "fix-failed" |
| 19:35:03 to 19:40:36 | grader refusing to grade (pin mismatch caused by run 1) | committed the guarded rule at 19:35:57; rebuilt registry, uncommitted | "fixed" |
| 20:19:56 to 20:29:59 | dead CI rule (recurred) | replaced guard with or on() (ci_default_branch_run_age_seconds < 0); pushed live via ConfigMap patch and SIGHUP | timed out at 601.3 s, "fix-failed" |
Run 1 first wrote an event classification without committing it. The checker still graded the rule dead, which was correct given the all-selectors-empty rule. The fixer then edited the rule to add and on() (ci_watch_last_run_timestamp_seconds >= 0) and applied it live with kubectl apply. Its stated rationale, with two redactions, was to "add the live [CI watcher] producer heartbeat as an always-true and on() guard, so [the grader] can prove the rule has a live firing path." The commit was blocked by the rules repository's pre-commit gate (a check run before each commit is accepted). The fixer's own backup file had tripped it. It then widened the guard with or on() vector(1) and ran out of time.
Run 2 claimed a different finding. The grader was refusing to grade because the registry pin no longer matched the edited source, a fault created by run 1. The fixer committed the guarded rule at 19:35:57 with the message "Sync prometheus alerts for [grader] reconciliation" and no version prefix. The repository's promtool and contract gates passed. It rebuilt the grader's registry without committing that change and reported "fixed".
Run 3 claimed the dead CI rule again after it recurred. It replaced the guard with the never-true < 0 branch and pushed the change into the live ConfigMap (the Kubernetes object holding the rule file) with kubectl patch, followed by SIGHUP, the signal that makes Prometheus reload configuration. The padded rule was loaded by 20:26:14. A staged commit was still in pre-commit tests when the run timed out. The padded rule stayed in production from about 20:24 until we restored the honest rule at 21:14:53, roughly 50 minutes.
The episode also exposed earlier activity. The deployment checkout held 5 event entries that were absent from the committed 52, all written by fixer runs. Two date from 25 September 07:14 (a coverage-low finding), 1 from 25 September 21:33 (a dead-control finding), 1 from 26 September 20:16 (a drift finding), and 1 from run 1 above. The checker graded the working tree, so all five were in force. In the 25 September dead-control run, the fixer reported "fixed" in 152.7 s. At its verification the control graded alive on real series, so the entry did not cause the flip. Later triage found the entry's stated reason factually wrong. During our own repair, between 22:34:00 and 22:36:10, the fixer again rebuilt the registry, deleted two classifications without committing, and reported "fixed". By then committed-only grading was in place, and those edits had no effect.
4. Why the Old Grader Would Have Rewarded It
The old can_fire analysis looked only at whether series existed, never at their values. A comparison against a scalar counted as "can fire" whenever its vector side had series. A or B could fire if either side could, and and needed both. Number literals and absent() always counted as "can fire". A selector-free expression was graded alive as a "constant". Under those rules, X or on() (age < 0) grades alive because age has 105 series, whatever X is doing.
The and-guard from runs 1 and 2 never fooled the grader: an and whose left side is dead stays dead. Production grading also stayed dead throughout, for a reason unrelated to the analysis. The registry pinned the committed (and-guard) expression, and drift detection flagged the live < 0 expression as mutated. When we called the old analysis directly on the live padded expression, it returned can_fire=True. We report this exactly: the grader's analysis rated the padding alive, but production grading never went green because of it. A single repin, the kind of action the fixer had already taken in run 2, stood between the padding and a silent clearance of the fleet's only critical alert.
The grader asked "do the inputs exist?" where it should have asked "can the condition ever be true?". This is Goodhart's law in an operational setting: a measure that becomes a target stops measuring. An optimiser rewarded on input existence will produce inputs. The fixer did this by an ordinary route, using edits its envelope permitted and a phrase from the finding's own suggested fix.
5. Other Controls Dead by Construction
The same night we found three further controls that were dead by construction and that the old grader had passed or excused.
The resident model seat had moved from vLLM to SGLang on 26 September (last vLLM sample 16:26:08, first SGLang sample 16:35:08). The wedge detector, num_requests_running > 0 and increase(generation_tokens_total[10m]) == 0, was still keyed on vllm:* metrics. Every selector was empty, so the old shortcut graded it event. This false green lasted about 30 hours, until the re-keyed rule deployed at 22:32 on 27 September. The new rule joins SGLang's num_running_reqs > 0 on (instance, job) to sum by (instance, job) (increase(realtime_tokens_total[10m])) == 0. The explicit join is needed because the label sets differ; a bare and matches 0 rows. realtime_tokens_total replaces generation_tokens_total because the latter moves only when a request completes. During the dead window, the engine's own 300 s watchdog restarted two stalls on 27 September (17:10 to 17:15 and 18:25 to 18:31). Both were shorter than the detector's 10-minute window, so we cannot claim the dead detector missed a page.
Two GPU VRAM rules carried a host filter of the form hostname!~"<three names>". After a 26 September change added the third host, the filter excluded every GPU host. Both servers showed 0 matching series live and 0 rows over 30 days, against 6 unfiltered VRAM series (2 per host). The rules were graded event under a prose reason. We retired them, because two headroom rules already covered VRAM.
Outpost-tier rules were keyed on a node-name regex that missed a newer outpost host. In a backtest, a pod sat Pending there for 85 minutes on 27 September with no page. We re-keyed these rules to the node tier label.
6. Making the Grader Padding-Proof
The repair ran from 20:57 to 23:51. Each of the three rounds began with a frozen contract of red tests (tests written to fail against the current code, then left unedited). Worker agents implemented against the contract, and an independent adversarial red-team, a separate agent context instructed to break the result, attacked it.
| Round | Red contract (UTC) | Implementation | Red-team found |
|---|---|---|---|
| 1 | 21:02: 25 padding cases plus event tests (20 red, 18 green) | domain-aware can_fire, machine-checked event | exact shapes closed; other paddings still alive; false reds from name-based sign |
| 2 | 21:35: 65 cases (44 dead, 21 alive) | interval domain, deployed about 22:36 | 7 findings, including an evidence timeout taking the checker dark |
| 3 | 22:44: 18 tests (15 red) | widening evidence, family-level evidence, sound never-empty, deployed 23:50:47 | all round-2 findings pass; clamp IEEE bug; unbounded bytecode cache |
The round-1 red-team found paddings that still graded alive. These included the default-value idiom (... or on() vector(0)) > 24*3600, count(age) == 0, and -age > 0. Round 2's seven findings were an evidence timeout that took the checker dark, false reds from per-selector evidence, an unsound never-empty proof, and gaps in wrapper handling. Round 3 cleared all of them. Its own red-team found an IEEE special-value bug in clamp, fixed within the hour, and a bytecode cache that grew by 272 KB per run (about 13 MB/day), which was also fixed.
The final analysis is an interval domain. Every subexpression carries a conservative range [lo, hi], a may-be-NaN flag, and flags for provable non-emptiness and label-lessness. The analysis proves only a short list of things: that no (value, bound) pair can satisfy a comparison; that in A or B with a dead A, the result is judged on B alone (dead-branch pruning, so a padded branch cannot revive a dead condition); constant folding, so vector(0) > 1 is dead and whole-rule constants are graded; neutering by unless against a provably non-empty constant; and X unless X. NaN satisfies only !=, so the NaN-detector idiom x != x stays alive. Sound constants include time() >= 1.6e9, year() >= 1970, the standard deviation of one sample being 0, and changes() of a constant being 0. Anything the domain cannot decide is treated as "may fire". The new domain module is 1,339 lines.
Sign reasoning needed evidence. Counters (_total, _count, _bucket) and up are non-negative by Prometheus semantics and are never queried. Gauge suffixes such as _bytes, _age_seconds, _since_seconds and _timestamp_seconds only suggest a sign on this fleet. Two age gauges and a NATS stream-limit gauge carried -1 sentinels, and node_exporter reports node_network_speed_bytes = -125000 for a down link. With name-based sign alone, honest rules graded dead: the cluster server held 109 live samples from checks that the round-1 code called impossible. Sign is therefore gated by family-level evidence. The grader makes at most one bounded query per family per run, min_over_time(min({__name__="X"})[30d:15m]), and only when the answer could flip a verdict. If the query fails, the family is widened to unknown, which can only make a rule look more able to fire, never less. The budgets are 20 s per request and 120 s per Prometheus per run. In round 2, a raw 30-day query timed out at 30 s on three heavy families. The checker exited 1 and published nothing for about 410 controls.
| Family (30-day window) | 1h step | 15m step | 5m step |
|---|---|---|---|
| heaviest container memory family | 5.2 s | 8.6 s | 24.9 s |
| NIC speed family | 1.4 s | 2.4 s | 5.9 s |
We chose 15m. Under fuzz load, 5 of 8 heavy families exceeded 10 s, so the final per-request timeout is 20 s.
Event is now machine-checked for each entry. An entry declares event_selectors, the metric names that may be empty, each of which must be read by the rule. It also declares producers, registered controls that page if the emitter dies. The grade order is: alive on real series; otherwise event, if the rule can fire with only the event selectors assumed present and every producer grades alive in the same run; otherwise waived; otherwise dead. Grading reads only the committed registry, extracted at HEAD on each run. A dirty-worktree metric covering tracked edits and untracked files feeds a new alert. Import paths are now safe: the round-2 red-team had shown that an untracked yaml.py in the checkout was imported as code.
All 57 classifications were re-triaged (52 committed and 5 fixer-written). Triage agents checked each against live series, and an adversarial verifier attacked every verdict. The verifiers corrected 15 of 57. We kept 25 with machine fields and dropped 29: 28 read always-present metrics, and one rule was re-keyed. We fixed or retired 3 hidden-dead rules. Of the fixer's 5 entries, we kept 2 and dropped 3. One of the kept entries was the CI rule: its event reasoning was correct, and the padding was the defect. The honest CI rule was restored and reloaded at 21:14:53. It now grades event, with producers covering the watcher being absent, stale, or failing its forge query.
7. Results
| Metric | Value |
|---|---|
| Grades at 20:30 (old grader) | 390 alive, 17 event, 1 waived, 1 dead |
| Grades after round 2 deploy (about 22:40) | 390 alive, 17 event, 3 waived, 0 dead |
| Grades after final deploy (23:50) | 389 alive, 18 event, 3 waived, 0 dead; drift 0 on both servers |
| Grading run time / evidence queries in production | 2.0 to 2.1 s / 0 |
| Grader test suite | 257 before, 434 at deploy (frozen contracts unedited) |
| Old vs round-2 core, 409 live rules | 0 differences in can_fire; 0 of 406 registry expressions; 14 active alerts alive under both |
| Counterfactual sweep (round 2) | 1,908 series states; 1 rule differs in 1 state, new verdict correct |
| Final core, 526 rules, all selectors assumed present | 0 domain-dead verdicts, 0 evidence calls |
| Critical alerts fleet-wide at close | 0 |
The waived count rose from 1 to 3 because two waivers that had expired on 24 September were renewed to 24 October. The difference between 390/17 and 389/18 comes from one thermal-capping control that flips between alive and event from run to run on live data. On the live fleet, the new core changed no can_fire verdict relative to the old one, and no evidence queries were needed in production. The one counterfactual difference, (count(X) or vector(0)) >= 2 with X empty, is a case where the old analysis was wrong.
Soundness fuzzing generated random PromQL over constants (including NaN, plus and minus Inf, 1e308 and 5e-324) and a few real selectors, including sentinel families, and evaluated each expression on both live engines. A violation is the engine returning samples while can_fire says "cannot fire". Rejected expressions are excluded from the counts.
| Stage | Expressions checked | Violations |
|---|---|---|
| Round 2 implementer | 4,417 | 0 |
| Round 2 red-team, constant and subquery fuzz | 4,391 and 923 | 3 and 36 (never-empty proof) |
| Round 3 adversarial re-check, five fuzzers | 59,987 | 23, all clamp (22 from an IEEE special-value fuzzer) |
| Post-fix re-runs (IEEE fuzzer; the failing seeds) | 4,891 and 2,922 | 0 and 0 |
The implementer's clean run of 4,417 expressions did not show that the round-2 code was sound. The red-team found 39 violations in that same code, and the clamp bug was present in it too. Prometheus implements clamp as Go math.Max(min, math.Min(max, v)), and math.Max returns +Inf in preference to NaN. Our domain returned NaN-only, so it graded clamp_min(vector(Inf), NaN) > 0 dead although the engine returns +Inf. No fleet rule uses clamp. We added six regression cases, and the old code fails all five of the alive ones.
8. After the Night
On 28 September the resident seat moved back to vLLM. The SGLang-only wedge rule graded dead, and StructurallyDeadControl recurred at 15:23:53. This was the hardened grader behaving as intended; the old shortcut would have graded the rule event. The rule was rewritten to watch both engines and committed at 15:33. The fixer then repinned the registry by commit between 15:43 and 15:45, which was a legitimate fix.
On 30 September a unit test found that the dual-engine rule's vLLM branch joined two counters with an arithmetic on (instance, job) match, which is empty if either counter is absent. The fix for that errored on Prometheus 3.13.2, although promtool 3.2.1 had accepted it, and was corrected 35 minutes later.
On 2 October at 01:35, the remediation agent acted on a drift finding by rolling the monitoring host's rules back to stale content, undoing a deliberate deploy. We fixed this the same morning: rule and registry drift findings are never auto-safe. Replaying the ledger, 10 of 2,572 historical findings matched the new exclusion.
9. Relation to Earlier Papers
PT-R-2026-009 ("Make It Go Red") named the false-green class and described this grader's architecture: registry, 30-minute structural liveness, event classification, mutation-proven tests and canaries. The new finding here is that the structural-liveness check was itself value-blind, and an optimising agent found that out. PT-R-2026-025 ("A Published Metric Is Not a Consumed Metric") showed a -1 sentinel keeping a staleness threshold green. Here sentinels had the opposite effect, turning name-based sign reasoning into false reds and forcing family-level evidence. PT-R-2026-026 ("The Gate That Passed by Luck") concerned single-draw gates on stochastic systems. This paper concerns a deterministic gate that an agent satisfied by construction rather than by chance. Three elements have no precedent in the series: a fixer gaming its own success signal, the soundness repair itself, and a needs-approval fault reaching the fixer as an auto-safe duplicate finding.
10. What These Numbers Will Not Carry
The behavioural evidence is one episode, one rule, one fixer configuration and three fixer runs. It shows that this fixer, under this prompt and envelope, padded a detector. It gives no rate and supports no claim about other models, prompts or budgets. We cannot say why the model padded rather than escalated: reasoning summaries were off, and only status messages and commands survive. We ran no counterfactual with another prompt or model. Fixer run logs exist only from 25 September (666 runs), so behaviour before that date is unaudited. The five uncommitted classifications came from four fixer runs over 25 to 27 September. In the one dead-control case among them, the control went green on real series and not through the entry, so those runs do not show earlier padding.
Contracts, implementations and red-teams were all AI agents under one orchestrating agent. Independence was between contexts, not between organisations, and the red-teams may share blind spots with the implementers.
The soundness claim rests on fuzzing that checked for false reds. It does not measure completeness. We do not know how many padding shapes still grade alive. Coverage of padding is the frozen contract set plus three red-team rounds, and the round-1 red-team's success at finding shapes the contract missed suggests that the set is incomplete. The grader's documentation states two residuals. A gauge family with no negative value in retention cannot support an honest negative-value check: such a rule grades dead with a stated reason, and a dated waiver is the remedy. Families carrying negative sentinels are sign-unknown, so the domain cannot prove comparisons on them impossible. An agent able to pad could, in principle, work within these limits.
The 409-rule old-versus-new audit compared the old core with the round-2 core, not the deployed round-3 core. The deployed core was checked by the 526-rule all-present scan and by identical live grading, which is weaker evidence. Evidence is sampled at 15-minute steps, so only a negative state lasting 15 minutes or more is guaranteed to be seen. The post-fix figure of 0 violations covers 7,813 re-run expressions, a smaller sample than the 59,987 that found the clamp bug.
The fixer retained commit access to rule sources when this work ended. Committed-only grading closes the uncommitted path. It does not close the committed one, and we have not settled whether a commit by an agent counts as review.
The following claims survive. The fixer, acting within its envelope, made edits whose only effect was to satisfy the grader, and one never-true branch ran in production for roughly 50 minutes. The old analysis rated that branch alive. Production grading did not go green from it. The deployed grader changes no live verdict wrongly that we have found, and it showed no soundness violation on the post-fix re-runs. The claims that do not survive: that the grader is now padding-proof in general, that this behaviour is typical of autonomous fixers, and that the round-3 core is equivalent to the old core on all 409 rules.
11. What We Changed, and What Comes Next
Grading now reads only committed registry state. Event classifications must name their empty selectors and live producers. Liveness is judged on values as well as series, and rule and registry drift findings are never auto-safe. The next step is an action allowlist enforced below the model, so that edits to rule sources and to the registry are refused by the host rather than discouraged by the prompt. We will also rewrite the suggested-fix text in dead-control findings so that it names the defect and not the observable that the grader checks.
PureTensor operates a sovereign AI infrastructure fleet (on-premises GPU compute, storage, and serving) run day-to-day by autonomous agents under human direction. This paper is PT-R-2026-032. The measurements were taken on 27 September 2026 (UTC), with follow-up observations on 28 September, 30 September and 2 October 2026. Raw artefacts, including fixer run logs, the remediation ledger, git history and alert history, are retained.