SMAD PickleBot · Post-mortems

SMAD PickleBot Post Mortems

Every production outage that reached players or money, newest first. 3 so far.

← All post-mortems
SMAD PickleBot · Post-mortem

Post Mortem: 9/29/26 New members got no welcome DM and no sheet row for 7 days

From 9/22 to 9/29 the daily member sync could not add any new member to the Player sheet, so three new players got no row and no welcome DM. A 9/9 change widened the row the sync writes without widening the range it writes to; Google Sheets refused every insert, and the sync reported success anyway.

Outage
7 days
Failed
9/22/26 12:20:40 PM
Detected
9/29/26, time not recorded
Fix pushed
9/29/26 11:57:47 AM
Restored
9/29/26 12:13:06 PM

Summary

Once a day, the Sync Group Members workflow compares the SMAD WhatsApp group with the Player sheet. Each new member gets a row on the sheet, in name order, and a welcome DM with the address, parking, court fees, how the weekly poll works and the survey link. Without the row the bot does not know the player: no vote tracking, no charges, no reminders.

On 9/9 at 9:24 PM PT, Join Date, Referrer and Join Source on the Player sheet, backfilled from history and stamped by sync-members added three columns to the sheet and changed how the sync builds a new member's row: the row now ran through every fixed column, including the three new ones. The range the row is written to was left as it was, ending at the win/loss columns. Google Sheets refuses a row that is wider than its range, so from then on every new member's row was refused, and without a row no welcome DM is sent.

Nobody joined for the next 13 days, so nothing showed. The first new member, ND, arrived on 9/22 and was refused at 12:20:40 PM PT. John Stowell (9/27) and cunningham dan (9/28) followed. The sync retried each of them twice a day and failed every time. It logged each failure and then exited as a success, so the workflow watchdog, which alerts on failed runs, saw 14 green runs in a row. Gene found the outage on 9/29 because two new players told him they had heard nothing.

Outage: 7 days, from 9/22 12:20:40 PM PT, the first refused member, to 9/29 12:13:06 PM PT, the last of the three welcome DMs, sent by a sync Snow White ran by hand on the fixed code at Gene's request. The defect itself was live from 9/9 9:24 PM PT, 13 days before any player met it. ND is Andy Tien, "John S" is John Stowell and "cunningham dan" is Dan Cunningham: the sync now takes names from Google Contacts first.

End-user impact: three new members got no welcome and no row. ND waited 7 days, John Stowell 2, cunningham dan 1. They had no location, parking, fee or poll instructions from the bot, and the bot did not know them: poll votes from them were not tracked and they got no vote reminders. Financial impact: none. From the Pickle Poll Log, only Andy Tien voted in the window: "Can't play this week" on the 9/27 poll, on 9/28 at 9:43 AM, before he had a row. His row now carries it. John Stowell and Dan Cunningham did not vote, and none of the three played, so nothing went uncharged. Side effect: the refused inserts left rows with no name in the Player sheet. The logs show 20 refused attempts (ND 14, John Stowell 4, cunningham dan 2), each of which inserted a row before its write was refused. Snow White found 18 such rows, under Chris Lee (2), Jesse David (4) and Nardo Manaloto (12). Each had copied the formulas of the row above it, that player's Mobile included, so the sync saw Nardo's number on 13 rows. Why 18 and not 20 is not established. All 18 were deleted on 9/29 between 12:14 and 12:37 PM PT (time not recorded) with Gene's approval.

Who did what

Timeline (PT)

WhenWhat happened
9/8 12:05:14 PMScheduled sync inserts Jerry Farrell and sends his welcome DM (last scheduled success)
9/9 4:33 PMGene asks for Join Date and Referrer columns; Mr Sandman hands it to Snow White
9/9 9:01:00 PMA manual sync inserts Sean Liu and sends his welcome DM: the last successful insert, on the old code
9/9 9:24:50 PM9428645 ships: row widened, range not. The backfill adds the three columns to the live sheet the same evening (time not recorded)
9/10 to 9/21Every run green; no new member joins, so nothing is inserted
9/22 12:20:40 PMFirst refused member: ND. Requested writing within range ['2026 Pickleball'!A58:V58], but tried writing to column [W]. The run reports success
9/22 to 9/26ND refused in every run, twice a day, each attempt leaving a nameless row; every run green
9/27 12:23:21 PMJohn Stowell joins and is refused too
9/28 2:27:22 PMcunningham dan joins and is refused too; 3 refused per run from here
9/29, time not recordedGene: "Why don't you do daily player sync anymore? I got 2 new players who haven't received welcome dm yet". Detected
9/29Mr Sandman reads the 9/28 run log, finds the refused rows and the cause, and fixes it (commits below)
9/29 11:57:47 AMThe fix (a9873cf) is pushed with the incident record (0a68198); its deploys and tests run green. The next scheduled sync would have been the next morning
9/29 12:13:04 to 12:13:06 PMRestored: Snow White runs sync-members --execute on the fixed code at Gene's request. Andy Tien, John Stowell and Dan Cunningham get their rows and welcome DMs (delivered or read, from GREEN-API's outgoing log)
9/29 12:14 to 12:37 PM, time not recordedThe 18 nameless rows are deleted with Gene's approval, after delete-blank-player-rows was changed to "no first or last name" (66f2c31)

Why it happened

  1. The row and its range were computed separately. The insert built its row from one number (every fixed column) and its range from another (the last column it filled). Every other write to the Player sheet goes cell by cell through the column map, so this was the only place where the two could disagree. 9428645 changed one of the numbers and not the other.
  2. A refused insert did not fail the run. sync-members logged the error and exited 0. The watchdog alerts on failed runs, so 14 runs that failed at their one job looked healthy.
  3. The insert added the row before writing it, with nothing to undo it. Each refused attempt left a nameless row behind: 20 attempts in the logs, 18 rows found on the sheet.

Why it shipped

Action items

#ActionStatus
1A player row and its range come from one column map: ColumnMapper.static_row() and static_range(), used by the insertDone a9873cf
2A member who could not be added fails the run, so the watchdog files an alertDone a9873cf
3tests/test-sync-members.py, the first suite for the insert. Its fake Sheets service refuses a row wider than its range, as Sheets does. It fails on 9428645's codeDone a9873cf
4A refused write deletes the row it inserted, so a retry cannot leave a nameless rowDone 0a68198
5The survey sync and jokes steps of sync-members.yml run even when the member sync fails, now that item 2 can fail itDone 4324176 (Snow White)
6Remove the nameless rows the refused inserts left in the Player sheetDone 18 rows deleted on 9/29 between 12:14 and 12:37 PM PT with Gene's approval, by the rule "no first or last name" (66f2c31, Snow White). The first command, 08cf29b, looked for rows with no value and no formula and found none
7Check whether ND, John Stowell or cunningham dan voted or played between 9/22 and 9/29, and charge any game they playedDone (Snow White): only Andy Tien voted, "Can't play"; his row now carries it. Nobody played; nothing owed
8Get the three their rows and welcome DMs, and record the restore timeDone Snow White's run at 12:13 PM PT on 9/29; restored 12:13:06 PT
9New members are named from Google Contacts first, so a WhatsApp handle like "cunningham dan" is welcomed by nameDone 8f7dfae (Snow White)
10A phone number on two player rows fails the run (the copied Mobile formulas put Nardo's number on 13 rows)Done 66f2c31 (Snow White)
11After every real run the sync checks one active row per group member, no shared numbers and no nameless rows, and fails if any check failsDone 8b92ced (Snow White); the first live check passed, 48 members on 48 rows
12A member who leaves the group is told their stats are archived and how to come back (DM + email, Gene's wording); the sync used to archive them silentlyDone 66f2c31 (Snow White). Jerry Farrell, archived by the 9/29 sync before this existed, was sent the notice by DM at 12:36 PM PT at Gene's request (no email on file)

Written by Mr Sandman on 2026-09-29 PT at Gene's request ("This is a player facing outage, members sync broken since 9/8 due to an untested regression. Do full post mortem"). Sources: the job logs of all 49 Sync Group Members runs from 9/5 to 9/28, the git history and the session notes. Every timestamp is observed and given in PT; where a time was not recorded this says so. Commits are named by their subject line and linked.

Incident record: ops/incidents.json → 2026-09-29-new-member-sync

Shareable page: https://smadpicklebot.com/postmortems#pm-2026-09-29-new-member-sync

Alert that started it: none. No alert fired: every run of the sync reported success. Gene noticed.

← All post-mortems
SMAD PickleBot · Post-mortem

Post Mortem: 9/27/26 WhatsApp token rotation took WhatsApp down for 4 minutes

On 9/27 a planned rotation of the GREEN-API token took WhatsApp down for 4 minutes. GREEN-API switches to a regenerated token a few minutes after showing it; the rotation script checked too early, stopped, and the old token then expired. Nothing was sent in those minutes, so no player was affected.

Outage
4 minutes
Failed
9/27/26 10:40:04 AM
Detected
9/27/26, before 10:40:04 AM (not recorded)
Fix pushed
9/27/26 10:42:36 AM
Restored
9/27/26 10:43:41 AM

Summary

The GREEN-API token is the one key the bot uses for every WhatsApp call. It was being rotated on purpose. The day before, the session had put the token's first half into a test file, which GitGuardian flagged, and had printed the whole token into its own transcript. The rotation script, scripts/rotate-greenapi-token.ps1 (9/26), checks a new token with GREEN-API before writing it anywhere. If GREEN-API refuses it, the script stops and changes nothing.

On 9/27 that safety check caused the outage. GREEN-API switches to a token regenerated in its console a few minutes after the console shows it. The script checked the new token straight away, got a 401, treated it as a bad copy and stopped. The old token was still live at that point, so nothing had broken yet. The session then misread the 401 as "some other GREEN-API key" and shipped a second way to regenerate the token, which also got a 401. A few minutes later GREEN-API retired the old token, and every call from the three functions started failing. Service came back on the fourth run, using a paste-at-the-prompt mode written during the outage.

Outage: 4 minutes, from the first logged failure at 10:40:04 AM PT to the first confirmed send at 10:43:41 AM PT (the incident record's start and end). The session saw 401s shortly before 10:40:04 but did not record the time, so the record starts at the first logged one.

End-user impact: none observed. Nothing tried to send in the window: no poll, reminder or /pb command. Had one arrived, its WhatsApp send would have failed. Financial impact: none. No payment, booking or charge touches GREEN-API.

Who did what

Timeline (9/27/2026, PT)

TimeWhat happened
10:30:04Webhook instance-state poll: authorized on the old token
not recordedRun 1. The clipboard held the token already in use (the console had not regenerated). GREEN-API answered 429 (rate limit). The script stopped with nothing written. Correct outcome, confusing message
not recordedRun 2. The console regenerated; the new token starts 259b4d. The script's check got 401 and it stopped with nothing written. The session checked the old token: still authorized
not recordedThe session concluded the pasted value was "some other GREEN-API key". Wrong: it was the new token, not yet active
10:37:42Pushed rotate-greenapi-token.ps1 regenerates the token itself through GREEN-API's updateApiToken
not recordedRun 3. updateApiToken, called with the old token: 401
not recorded, before 10:40:04The session checked the old token: 401 on getStateInstance, getSettings and getWaSettings. Outage identified
not recordedThe session gave a one-line .env update that read the clipboard. Copying that command from the session replaced the token on the clipboard, so .env was not updated
10:40:04Webhook poll: getStateInstance returned 401 (WARNING, not paged)
10:41:19 – 10:41:41Run 4 with -Paste: the new token was accepted (authorized), both secrets were updated and the new revisions went live (smad-picklebot-00523-nkt, whatsapp-message-sender-00397-wbh, smad-whatsapp-webhook-00488-nm8)
10:42:36Pushed rotate-greenapi-token.ps1 -Paste: start it, then copy the console's new token
10:43:41Restored: the first send on the new token (the reply to Gene's /pb help) was delivered

Why it happened

  1. The script assumed GREEN-API switches tokens instantly. On 9/26 a rotation went through on the first check, and the delay was never looked for. A refusal was treated as final, so the script stopped at exactly the moment the switchover was already under way.
  2. The session diagnosed from one sample. After run 2, the old token still worked, and that was read as proof nothing had been regenerated. "Not active yet" was never considered, and a second, untested mechanism was shipped instead of waiting a few minutes and checking again.
  3. An instruction that fought the clipboard. The fix-up command read the token from the clipboard, but running it meant copying the command, which replaced the token.

Why it shipped

The script can only be tested by a real rotation, which is itself the risky act. The first real rotation (9/26) happened not to show the delay. Its parse was checked; its behaviour against a slow switchover never was.

Action items

#ActionStatus
1Paste at a prompt: start the script, then copy the token, so nothing overwrites the clipboardDone df583ba
2The check waits out a 401 on a new token, and a 429, every 15 seconds for up to 3 minutes, with the token already saved in .envDone in the commit that adds this post-mortem
3The console-paste flow is the default; updateApiToken moves behind -Api because its one try failedDone in the commit that adds this post-mortem
4Never grep across .env; never build a test fixture from a real secret (why this rotation happened at all)Done session memory rule, and the redact test's fake token is visibly fake

Written by Snow White on 2026-09-27 PT at Gene's request. Sources: Cloud Logging for the three functions, the Cloud Run revision times, the git history and this session's own checks. Every timestamp is observed. Where a time was not recorded, this says so rather than estimating it. Commits are named by their subject line and linked.

Incident record: ops/incidents.json → 2026-09-27-greenapi-token-rotation

Shareable page: https://smadpicklebot.com/postmortems#pm-2026-09-27-greenapi-token-rotation

Alert that started it: none. No alert issue was filed. The webhook's instance-state poll logged the 401 as a WARNING, which is correct: one failed read is not paged. The session found it first.

← All post-mortems
SMAD PickleBot · Post-mortem · issue #69

Post Mortem: 9/20/26 Sunday 8am New Games Poll Broke because of refactor regression

On Sunday 2026-09-20 the 8 AM automation that creates the week's games poll and sends every reminder did not start. The poll reached the group at 5:54 PM, ten hours late. A cost-cutting change to the GitHub Actions workflows the night before caused it.

Outage
9 h 55 min
Failed
9/20/26 08:00:04
Detected
9/20/26 09:28
Fix pushed
9/20/26 17:43
Restored
9/20/26 17:55

Summary

PickleBot's daily 8 AM job is a GitHub Actions workflow, Daily Reminder Runner. One of its jobs calls a second workflow, Release Notes, to send the daily digest. GitHub's rule for that arrangement: the calling job must grant every permission the called workflow asks for, or GitHub refuses to start the run at all. The night before, a refactor to cut Actions minutes, Actions minutes: retire GREEN-API Watch, gate close-on-green on an OPEN_ALERTS variable (2026-09-19, 10:03 PM PT, Snow White), gave Release Notes actions: write and did not give it to the only job that calls it, daily-reminder-runner.yml's release-digest. At 8:00:04 AM GitHub refused the run (startup_failure, zero jobs). Nothing that rides that job happened: no games poll, no Games This Week report, no Last Call scheduling, no game day, vote, payment or survey reminders, no Venmo sync, no court cache refresh, no digest. The watchdog filed the alert 88 minutes in; the fix waited until 5:43 PM for a session that was awake and allowed to make it; a hand dispatch restored service at 5:55 PM PT. Outage: 9 hours 55 minutes.

The same file had failed the same way on 2026-09-04, and the comment documenting that failure sat six lines above the block that was not updated. The same push carried a second defect: it gated the alert closer on a repository variable that the Actions token cannot write (HTTP 403); the refusal was logged as a WARNING under a step that reported success, so alert auto-closing was silently off for a day.

End-user impact. No player could vote on the new week's games for ten hours: the Sunday poll that normally opens at 8 AM did not exist until 5:54 PM. Every reminder due that morning was also missed. One knock-on found the next day: Shyam's vote-to-court handover did not fire for Tuesday 7 PM. He voted at 7:54 PM Sunday, 47.1 hours before the game, and the handover requires 48 (a member can only book his own free court 48+ hours ahead: Shyam court: skipping 09/22/26 - ... 47.1 hours away in the vote webhook's log). With the poll out at 8 AM he would have voted 11 hours earlier and it would have run. The automation itself is intact; the late poll pushed his vote inside the window.

Financial impact. No direct money loss. Potential loss from players who normally vote on Sunday for the week's games, could not, and may not get around to it later in the week, which means fewer paid slots for those games.

Who did what

Three Claude Code sessions and two pieces of automation acted in this incident. Each is named with what it did right and what it did wrong.

actorwhat it ispart in this incident
Snow WhiteClaude Code session on Gene's desktop; the only session with credentials, run logs, a dispatch token and, by project rule, the right to edit workflowsCaused the regression. Wrote and pushed the change that gave Release Notes a permission its caller did not grant, and the variable gate the workflow token could never open; verified only the top-level path; wrote tests that encoded the same mistake. Was off Sunday morning, so could not repair it for ten hours. Then: dispatched the runner that restored service, removed the variable gate and rebuilt the closer, corrected the incident record, updated the investigator's instructions, wrote this post-mortem.
Alert InvestigatorHourly cloud routine: reads open alert issues, triages from the runs API and the code, writes its finding on the issue and in the sessions notes, wakes Mr Sandman; cannot read run logs, dispatch or change codeTriaged correctly and wrote one ambiguous line. Named the cause, the one-line fix and the need for a hand dispatch 46 minutes after detection, and separately spotted that the alert closer's gate had not opened. But reported the Gmail watch as "renewal (day 20, even)": accurate shorthand for "the runner renews on even calendar days and today is the 20th", and open to any reading by someone without the workflow file in front of them. Its live instructions now forbid that shape.
Mr SandmanClaude Code session in a 24/7 cloud sandbox; has the repo and git push; no credentials, no run logs, no dispatch tokenFixed the outage, and added noise to the investigation. Once Gene handed it the workflow edit, it pushed the caller fix and the lint check that catches the whole class within the hour. But it read the investigator's line as "day 20 of a 7-day watch", repeated that in two notes and in the incident record, and told Gene the Gmail watch had a hard deadline that day. It did not verify: the runner's own workflow file states the even-day rule, and the runs API showed the 9/18 renewal step green (the step's log text is out of the sandbox's reach; the file and the step conclusion were not). Its incident entry also said "no monitor caught it", when the watchdog had filed the issue 88 minutes in. Both are corrected.
Workflow WatchdogGitHub Actions workflow, hourly sweep over every workflow's latest runDetected it. Read the failed conclusion and filed #69 at 9:28 AM. The 8:41 sweep ran 47 minutes late: GitHub's scheduling, not the watchdog's.
Cloud SchedulerGCP job that dispatches the runner at 8:00 AMDispatched on time. Its retries cover a failed dispatch call, not a run GitHub accepted and then refused to start.
GeneOwner; the only human in the loop, and the only one who can turn the desktop session onAway all day, and the decision-maker once back. Spent Sunday driving his son to UCSD and packing, and did not notice the 8 AM poll had not gone out. Around 5 PM, on the drive back from San Diego, opened the Mr Sandman session, saw the outage, and told Mr Sandman what to do. Snow White's desktop had been off since Saturday night. He did not want Mr Sandman making production fixes: it has no live logs and no keys, and it had not caused the regression, so he allowed it only the one-line workflow edit that every weekday run depended on. Home around 6 PM, he turned the desktop on and told Mr Sandman to hand everything else to Snow White. Chose the full dispatch over poll-only, and asked for this post-mortem.

Timeline (PT)

whenwhat
09/19 22:03Actions minutes: retire GREEN-API Watch, gate close-on-green on an OPEN_ALERTS variable pushed. Its own runs (four deploys, Tests, Lint) green: they exercise release-notes.yml only as a top-level workflow, where its own block sets the token. The called form runs once a day.
09/20 08:00:04Cloud Scheduler 8am-daily-runner dispatches; run 35518182853 → startup_failure at 08:00:06. Scheduler retries (2) do not apply: the dispatch itself returned 200.
09:28Workflow Watchdog hourly sweep (cron :41; GitHub ran it 47 minutes late) reads the conclusion and files #69. Detection: 88 minutes.
10:14Alert Investigator triages on #69 and in the sessions repo: cause, the one-line fix, "dispatch by hand". Correctly notes the Gmail watch renewal was due that day (even calendar day).
10:14 → 17:00Nobody who could act was available or allowed to. Snow White's desktop had been off since Saturday night. Mr Sandman cannot dispatch and, by CLAUDE.md, does not write workflows. Gene was on the road to UCSD with his son and did not see the alert. Nearly 7 hours of the outage is this gap.
~17:00Gene, driving back from San Diego, opens the Mr Sandman session, sees the outage and directs it. He does not want Mr Sandman making production fixes (no live logs, no keys, not the author of the regression) and allows it only the one-line workflow edit every weekday run depends on.
17:43The 8am runner starts again: a caller must grant what the workflow it calls asks for pushed by Mr Sandman: caller fixed; check-workflow-reporting.py check 3; tests/test-workflow-permissions.py; incident entry with restored: null.
~17:50Gene, home, turns the desktop on and tells Mr Sandman to hand everything else to Snow White.
17:53Snow White dispatches run 35549088614 (reminder_type=all, Gene's choice).
17:54:13[OK] Availability poll created in SMAD Pickleball group.
17:55:10Run complete, every step green, release digest ran. Restored.
18:01Alerts close from the watchdog's hourly sweep; the OPEN_ALERTS gate is gone pushed: defect 2 removed, closer rebuilt, scopes reverted, both incidents recorded.
18:02A hand dispatch of the watchdog on the new closer closes #69 with the sweep's own comment.

What the refactor was

Actions minutes: retire GREEN-API Watch, gate close-on-green on an OPEN_ALERTS variable, one push of 23 files after GitHub's "90% of Actions minutes" alert, did three things:

  1. Deleted greenapi-watch.yml, a workflow README already recorded as retired (126 billed minutes that month for a check the keepalive poll makes). Correct, and not involved in the outage.
  2. Gated close-on-green.yml on a repository variable OPEN_ALERTS, to stop a six-second job from billing a minute 193 times a month. The reporter was to set the variable, the closer to clear it, the watchdog to re-derive it hourly. Defect: GITHUB_TOKEN cannot write repository variables at all; actions: write does not cover it. The gate never opened; the refusal was a printed warning; the step reported success.
  3. To let workflows write that variable, added actions: write to the permissions block of every workflow containing the string report-failure: fifteen files, including release-notes.yml. Defect: daily-reminder-runner.yml calls release-notes.yml as a reusable workflow (uses: ./.github/workflows/release-notes.yml), does not contain the string, and was not patched. Its release-digest job granted contents: read, issues: write; the callee now asked for actions: write; GitHub refused to start the run.

Defect 3 caused the outage. Defect 2 is why #69 then sat open behind green CI.

Why it shipped

Three failures of method, in order of weight.

  1. The verification proved the wrong path. I read a green push as proof of the change. The push's runs exercised release notes only as a top-level workflow. The called form runs at 8 AM, once a day, and nothing I did before pushing exercised it. daily-reminder-runner.yml:333-345, written after the 09/04 outage, says exactly this: "Direct dispatches of release-notes.yml worked all along... The failure only exists in the called form... so the verification runs the night before proved the wrong path." I was editing permissions across the repo and did not read the one file where permissions have a second consumer.
  2. The blast-radius search matched a string, not a relationship. "Who uses report-failure?" is answerable by grep and I answered it. "Who inherits this job's permissions?" is a graph question with a second edge for every callee, and I never asked it. The static check that now exists (check 3) is the search I should have done by hand.
  3. The tests I added encoded my own predicate. "Every file containing report-failure grants actions: write" passed on the callee and could not see the caller. docs/LESSONS.md already has a section titled The Author's Tests Inherit the Author's Blind Spot, written on 2026-08-30 for three instances of this shape. I repeated it.

Contributing: the change was framed as a low-risk cost cut and shipped late on a Friday in a single 23-file push; and I declared the variable gate "proven on the first alert", which is to say unproven, with its failure mode set to report success. Both halves of the rule a permissions block is a denylist by omission; a called workflow is bounded by its caller's job were in CLAUDE.md; I applied the first.

The Gmail watch question

Gene asked why the investigator said the Gmail watch renewal "hasn't been done in 20 days". It had not said that, but what it did say was no better for a reader: "Gmail watch renewal (day 20, even)". What that shorthand meant: the runner renews the watch on even-numbered calendar days, 9/20 is one, so the skipped run also skipped a routine renewal. Mr Sandman rendered it as "day 20 of a 7-day watch" in its 5:55 PM note, again in the 6:25 PM relay to Snow White, and in the incident entry's impact line, without checking the runner's workflow file or the 9/18 run, and the number changed meaning on the way from "the 20th" to "20 days old".

The facts from the run logs: the watch was renewed on 9/18 at 8:01 AM PT (Expires: 2026-09-25 15:01:28). The 9/20 attempt was skipped with the rest of the run; the next even day, 9/22, renews it three days before expiry. Nothing was at risk. The incident entry is corrected.

Two lessons, both now written down. The investigator's live instructions (the routine's prompt, and Appendix A of Investigator.md) now require every figure to carry its meaning in the same sentence, and every commit, run and issue to be named by title and link, never a bare hash: "day 20, even" was accurate and useless to anyone without the workflow file open. And a number that crosses a relay must carry its source (docs/LESSONS.md, A Relayed Number Carries Its Source).

A side finding: gmail-watch-renewal.yml ("every 6 days") has no schedule, last ran in March, and failed then. The runner's even-day step is the only live renewal. See action item 9.

Detection and repair

Detection worked: the watchdog's hourly sweep filed #69 88 minutes after the failure, and the investigator had the cause and the fix written within 46 minutes of that. The incident entry as first written said "no monitor caught it"; that was wrong and is corrected.

Repair did not: nearly 8 of the 10 hours were spent waiting for a session that was both awake and allowed to edit a workflow and dispatch a run, and for the one human who could turn that session on, who was on the road all day and saw the alert at 5 PM. The alert reached the right places; it reached nobody who could act. That wait is a policy, not an accident.

Earlier instance: 2026-09-04

The same file failed the same way sixteen days earlier and was never recorded until now. Send the whole day's list, to three chats, at 8am from Cloud Scheduler (09/03) added the call to release-notes.yml with no permissions block on the calling job; the 09/04 8 AM run failed to start (run 33886955002); Let the 8am runner start again: grant the digest job the scopes it calls for fixed it at 9:56 AM and a hand dispatch (run 33898063769) restored service at 9:58 AM. Outage 1 h 58 min. The fix's comment described the trap precisely; a comment is not a check.

Action items

#itemstatuswhere
1Caller must grant what callee asks: scripts/check-workflow-reporting.py check 3, run by Lint Workflows on every push and by tests/test-workflow-permissions.py in the local suite, so the local suite fails before a pushdoneThe 8am runner starts again: a caller must grant what the workflow it calls asks for (Mr Sandman)
2Remove the variable gate: the closer runs from the watchdog's hourly sweep with no state to write; actions: write removed from all sixteen callers; OPEN_ALERTS deleted; proven on #69 under the workflow tokendoneAlerts close from the watchdog's hourly sweep; the OPEN_ALERTS gate is gone (Snow White)
3Incident record corrected: 09/20 restored time, Gmail wording, detection credited to the watchdog; 09/04 outage entered with the commits that introduced and repaired itdonePost-mortem: the Sunday poll never went out (2026-09-20) and the revision that added this section's links
4The investigator writes for a reader who has not seen the code: every figure carries its meaning; commits, runs and issues by title and linkdoneLive routine trig_019cYNV8RgZ5k9gvmPhvnPCR updated 2026-09-20 6:40 PM PT; Appendix A of Investigator.md in the same revision
5Before any permissions: change: grep uses: ./.github/workflows/, read every caller's job, and name which trigger paths the push's own runs will not exercise before calling it provendoneSnow White's standing rule; docs/LESSONS.md
6One push per task, and a permissions change is its own task: never bundled with a deletion and a new gate in one late-evening pushdoneOne push per task (Gene, 2026-09-19 PT): the project reading of the shared rule
7Give the cloud session a repair path when the desktop is off: a committed ops/rerun.txt that a push-triggered workflow reads and dispatches the named workflow with the given inputs. Turns an 8-hour wait into a push. One billed minute per usedoneThe rerun lever: a push to ops/rerun.txt dispatches the workflow it names, so the cloud session can repair production when the desktop is off; gmail-watch-renewal.yml is retired (Snow White, approved by Gene 2026-09-22 PT) — rerun.yml, .github/scripts/rerun_lever.py, tests/test-rerun-lever.py
8Or remove the edge entirely: have the runner's release-digest job dispatch release-notes.yml through the API instead of calling it as a reusable workflow, so no caller→callee permission inheritance is left to get wrongdeclinedGene, 2026-09-22 PT: check 3 in scripts/check-workflow-reporting.py fails Lint Workflows on the whole class before it can reach main, and tests/test-workflow-permissions.py pins it locally; an API dispatch needs the same actions: write scope, bills the same one job, and turns one run into two to watch
9Retire gmail-watch-renewal.yml (no schedule, dead since March, failing) and its README rows; the runner's even-day step is the renewaldonesame commit as item 7: The rerun lever … gmail-watch-renewal.yml is retired — verified from the runs API 2026-09-22 PT: workflow_dispatch only, last run 2026-03-04, failed
10A relayed number carries its source: session notes and incident entries quote the log line, not a paraphrase, for any figure with a deadline in itdonedocs/LESSONS.md, A Relayed Number Carries Its Source

Written by Snow White on 2026-09-20 PT at Gene's request, from the runs API, the run logs, the git history and the session notes. Every timestamp is observed, none estimated. Commits are named by their subject line and linked; a bare hash is never the reference.

Incident record: ops/incidents.json → 2026-09-20-daily-runner-startup-failure

Shareable page: https://smadpicklebot.com/postmortems#pm-2026-09-20-sunday-poll-never-sent