00System health
Full end-to-end sweep of every site, worker, feed and pipeline — run against live systems on 2026-07-27, ~02:30–03:00 UTC. Each row names how it was checked, so any of it can be re-run.
Verified working
| Area | Result | How it was checked |
|---|---|---|
| All 12 sites | ✅ up | HTTP 200 on all; DryRun 401 is its password gate — expected, not a fault |
| Ops feed (CMD_CTR) | ✅ fresh | flight-data.js fetchedAt ~2 min old; CI green on the last 12 consecutive runs |
| Progress feed (DB001) | ✅ fresh | cache.json _updated ~5 min old; 4 batches + 3 curricula present |
| CMDV2 snapshots | ✅ fresh | Refresh job green, chained off CMD_CTR — ~5 min cadence, not the documented hourly |
| Dispatcher | ✅ running | Cron */5; DB001 + CMD_CTR dispatch runs landing on schedule |
| Watchdog worker | ✅ healthy | /status: healthy:true, lastError:null, runCount 5263, anomalyStreak 0 |
| Watchdog tests | ✅ 119/119 | npm test in AP127_V2/watchdog — 3 files, all pass |
| 07-27 stabilization fix | ✅ deployed | wrangler deployments list — live version created 02:29:34Z |
| Portal structure drift | ✅ none | portal_fingerprint.json current — 48 RPC fns, 3 Timeline modes, Daily Schedule present |
| CMDV2 in-browser | ✅ clean | 17 views load, zero console errors, now serving p114 |
Findings
✅ Fixed — duplicate flight rows were inflating every view and every hours KPI CMDV2 p114
attachCancelDetails() pushed one synthetic row per cancellations[] record — but upstream emits one record per cancel event (its id embeds a submission timestamp), so re-cancelled bookings duplicated: 5 bookingIds → 6 phantom rows, 3 of them AP-127, one rendering 3× with a colliding React key.2. The same flight ingested twice as two
ACTUAL_ONLY rows under different _ACT_<n> ids survived the existing dedup pass — which only ever removes planned rows — and double-counted block hours: 67 rows.
MEETING bookings were deliberately excluded from the dedupe (they legitimately repeat at the same time); all 26 verified preserved.flight_schedule.pre_migration_archive.json, re-applied as an override every run, so it can never self-heal. Walking one commit per day back to 2026-06-15 shows the count climbing daily to 437 by 07-10, then pinned at 425 from 07-11 onward — exactly when that archive was frozen. The post-migration scraper adds none.⚠ Open — the dead-man's-switch monitor still cannot send alerts
wrangler secret list on ap127-watchdog-monitor returns [].monitor:state reads {alertedDown:false, downStreak:0, reason:"ok"}, and its seemingly-stale lastCheck is by design (it persists on a 6-hour heartbeat to save KV writes). But runMonitor() gates sending on if (alert && env.TELEGRAM_BOT_TOKEN), so with no token it silently no-ops. It detected the 19-hour outage on 2026-07-21 correctly and still could not page anyone. Every future silent watchdog death repeats that blind spot. One command fixes it:cd /Users/nugui/AP127_V2/watchdog-monitor && npx wrangler secret put TELEGRAM_BOT_TOKEN
✓ Closed 2026-09-02 — CMDV3 retired
dispatcher/worker.js targeted only DB001 and CMD_CTR, so CMDV3 lived permanently on its own delayed schedule: cron — measured gaps of 1–2.5 h versus CMDV2's ~5 min, meaning V2 and V3 could show different numbers for the same moment. The proposed fix was one more dispatcher target; it was never applied. The finding is now moot: CMDV3 has been retired entirely (repo archived, Pages project and local dir deleted) because CMDV2 remained the site actually in use. The divergence risk flagged here was the real cost of running two dashboards off one feed — removing the second resolved it more cleanly than syncing it harder would have.⚠ Open (upstream, low impact) — ID collisions in the frozen pre-migration data
ACTUAL_ONLY_ with an empty suffix, where 21 distinct flights share one id. Data is correct; identity is not. All 206 such rows fall on 2026-05-05→2026-06-10, outside the watchdog's rolling snapshot window — which matters, because buildSnapshot() keys by id, so a collision inside the window would hide real changes. Notifications are therefore unaffected; the only cost is duplicate React keys when those historical dates render. Fixing it means rewriting ids inside the archive file that is explicitly marked do-not-hand-edit, so it is recorded rather than done.01Bird's-eye overview
Two independent data domains — Operations (what is flying, when, with whom) and Progress (how far each student has advanced) — are merged only at the presentation layer. Every site is a static front-end; all "backend" work is GitHub Actions cron jobs that commit data files into repos, plus four tiny Cloudflare Workers.
per-batch CSV")]:::src PORTAL[("Flight Ops Portal
Google Apps Script")]:::src RELAY[/"Apps Script RELAY
CORS bridge"/]:::src GS --> RELAY --> UC["update-cache.js
DB001 · Action"]:::data UC --> CACHE["cache.json
4 batches"]:::data CACHE --> KV[("CF KV
ap127_slice")]:::data KV --> API["ap127-data-api
worker"]:::wkr PI[("Orange Pi Zero 2W
every 5 min, concurrency 2")]:::src PORTAL -->|"getStudentSchedule RPC"| PI PI -->|"git push, CI-Skip"| FS["fetch_schedule.py
+ CI fallback"]:::data FS --> GH[("GitHub main
flight-data*.js")]:::data GH --> RAW2[("raw.githubusercontent")]:::data RAW2 --> DATA["ap127-data Worker
proxy · 60s cache"]:::wkr DATA --> CC["CMD_CTR
ops dashboard"]:::site CACHE --> DB001["DB001
admin site"]:::site DATA --> DBSHARE["DB_Share
student site (/mirror proxy)"]:::site API --> DBSHARE DATA --> V2["CMDV2 · PRIMARY
unified SPA"]:::site API --> V2 CACHE --> V2 RAW2 --> WD["ap127-watchdog
cron 2m + POST notify"]:::wkr WD -->|notify| TG(["Telegram SP"]):::src WD --> V2 WD -.->|"shared KV"| MON["ap127-watchdog-monitor
dead-man switch, cron 10m"]:::wkr MON -.->|"no bot token — cannot alert"| TG DISP["ap127-dispatcher
cron 5m"]:::wkr -.->|"only if feed ≥35m stale"| FS DISP -.->|"only :00/:15/:30/:45"| UC CC -.->|"trigger"| V2 PORTALL["Portal
launcher"]:::site -.links.-> CC PORTALL -.links.-> DB001 PORTALL -.links.-> DBSHARE PORTALL -.links.-> V2
TELEGRAM_BOT_TOKEN. As of 2026-09-02 a second detector (flight-data staleness) routes through that same missing secret, so nothing can page a human at all. Tracked in System health.02The twelve websites
All reachable as of 2026-07-27 (DryRun answers 401 by design — it is password-gated). Plus one experimental optimizer (flight-scheduler) that is local-only and not part of the deployed set.
ap127-cmdv2* duplicate Pages projects. Corrected 2026-07-27 against gh repo list AP127CMD: RPT is private (previously documented public), the repo AP127CMD/FlightTraining does not exist (that site is a direct-upload with no remote), and AP127_Docs and IFR_Flight_Deck were missing from this list entirely.flight-scheduler — experimental optimizer (not deployed)
~/flight-scheduler, no git remote. docker compose up --build seeds 24 students / 8 instructors / 10 aircraft / 101 lessons.03Cloudflare Workers — 4 services
All free-tier, all in account ae38e04e56d0ae52d3ec47ad29977587.
⏱ ap127-dispatcher
workflow_dispatch to DB001 update-cache.yml (only at :00/:15/:30/:45 via shouldDispatchDb001()) and CMD_CTR fetch_schedule.yml (only if the feed is ≥35 min stale). Does NOT dispatch CMDV2 — CMD_CTR chains that itself.GITHUB_PAT secretcron '0 * * * *' is only a fallback.DB001/dispatcher/🔌 ap127-data-api live
ap127_slice, returns JSON, CORS-locked to the student site.KV → AP127_STUDENT_DATA (c5c88c81…)ALLOWED_ORIGIN = ap127-dashboardr1.pages.dev📡 ap127-watchdog live
ap127-watchdog-AP127_WD (b42f3202…ccd5)healthy:true, runCount 5263, no errors, 119/119 tests🐕 ap127-watchdog-monitor muted
watchdog:status straight from shared KV, not over HTTP — a same-account Worker→workers.dev fetch is blocked (CF error 1042), and a CPU-killed watchdog stops writing KV, so a frozen lastRun is the truest death signal.lastCheck is normal.wrangler secret list returns []. Open since 2026-07-17.GET / on ap127-dispatcher returns HTTP 500. That is expected — it is a cron-only Worker with no meaningful fetch handler. Judge it by whether its dispatch targets actually run, not by that response.ap127-data-api worker source
export default {
async fetch(request, env) {
const allowedOrigin = env.ALLOWED_ORIGIN || '*';
if (request.method === 'OPTIONS') return new Response(null, { headers: {
'Access-Control-Allow-Origin': allowedOrigin,
'Access-Control-Allow-Methods': 'GET', 'Access-Control-Max-Age': '86400' }});
if (request.method !== 'GET') return new Response('Method not allowed', { status: 405 });
const data = await env.KV.get('ap127_slice', 'json');
if (!data) return new Response(JSON.stringify({ error: 'No data' }), { status: 503,
headers: { 'Content-Type': 'application/json' }});
return new Response(JSON.stringify(data), { headers: {
'Content-Type': 'application/json', 'Cache-Control': 'no-store',
'Access-Control-Allow-Origin': allowedOrigin }});
},
};
04Data sources & feeds
Operations feed → flight-data.js
A Google Apps Script web-app renders the academy Flight Operations Portal. The schedule lives in a JS object flightCache inside a sandboxed cross-origin iframe — so a headless Playwright browser is needed, not a plain GET.
/exec (90s budget; GAS cold-starts), wait for the iframe src, enumerate page.frames, frame.evaluate() out flightCache. Up to 3 attempts, 20/40s backoff.-→null, derive isSimulator/isStandby/durationMin.data/flight_schedule.json — fresh dates overwrite, dates outside the rolling ~10-day window preserved. Rolling backup written first.flight-data.js (window.FLIGHT_DATA), strips actuals from non-Completed flights, appends ?v=<unix-ts> cache token. Recovery: rebuild_history.py replays all git commits.Progress feed → cache.json
Google Sheets published as per-batch CSV, read through a second Apps Script relay (CORS bridge = the RELAY_URL secret).
RELAY_URL + "?batch=" + batch for each batch./^AUPRT/i (never counted anywhere).planned[], next_lesson, finish date, remaining, pct (workday/holiday-aware).ap124 ap126 ap127 ap129 monthly cur124 cur126 cur127 cap _updated. push-to-kv.js PUTs only {ap127,cur127,_updated} to KV ap127_slice.?batch= to a published Google Sheet CSV (gid 416854743) and proxies it as text/plain. Deploy as Web app · Execute as Me · Anyone → that /exec URL is the RELAY_URL secret.const CSV_URLS = {
AP124: "docs.google.com/spreadsheets/d/e/2PACX-1vRQK_to…/pub?gid=416854743&output=csv",
AP126: "docs.google.com/spreadsheets/d/e/2PACX-1vQB8PiZ…/pub?gid=416854743&output=csv",
AP127: "docs.google.com/spreadsheets/d/e/2PACX-1vQNxzCi…/pub?gid=416854743&output=csv",
};
function doGet(e){
const url = CSV_URLS[(e.parameter.batch||"").toUpperCase()];
if(!url) return json({error:"Unknown batch"});
return ContentService.createTextOutput(UrlFetchApp.fetch(url).getContentText("UTF-8"))
.setMimeType(ContentService.MimeType.TEXT);
}
AKfycbzs…/exec Apps Script is an academy system the owner cannot access — the operations feed is an external black box scraped by Playwright. This is the system's single most fragile point: if the academy changes the portal, fetch_schedule.py breaks (→ fetch-failure Issue).runScheduler — greedy capacity simulator (not a fetch)
After parsing the CSVs, update-cache.js projects every student's future lessons forward to compute planned[] · finish · monthly.
CFG.n129:13 placeholder students (AP129-01 "Student 01", 0 done) are generated in-code from ap129start 2026-06-01, using the AP127 curriculum as a stand-in — a capacity projection of a batch that hasn't started. No missing CSV.update-cache.js: AP127_NICKS / AP127_FI / AP127_SE (28 index-matched entries) + AP127_FI_FULL map, assigned to AP127 students by position. CSV student rows kept only if catc_id starts with 681; /^AUPRT/i lessons dropped.05Auto-fetching mechanism
Pi-primary since 2026-09-02. The Orange Pi Zero 2W does the scraping; Cloudflare + GitHub Actions are the automatic fallback. Before that the roles were reversed.
ap127-cmd-ctr ~95/day + ap127-ngt2 ~93/day + ap127-db001
~296/day) — every data commit triggered a full rebuild. Fixes, no payment method added:
- New
ap127-dataWorker (flight-schedule-feed/data-worker/) — a stateless proxy that re-servesflight-data.js/flight-data-recent.js/cache.jsonfromraw.githubusercontent.comwith a browserContent-Type+ 60 s edge cache (Range sliced locally, ETag/304, stale-fallback). No bindings, no storage, no secrets. R2 was the original design — dropped because enabling R2 needs a card on file. - Browsers (
ap127-cmd-ctrindex.html;ap127-ngt2index/legacy/ops/crosscheck/overview; DB_Share's/mirrorproxy) load the data from the Worker. CMDV2'sflight-data.jsmirror file was deleted. [CI Skip]in every data-commit message → Cloudflare Pages skips the build (deployment statusidle, off the 500/mo cap). Verified live onap127-db001+ap127-cmd-ctr. Build watch paths were the plan but the Pages API silently dropspath_excludes(dashboard-only). A real code push must NOT contain[CI Skip].- Backend Workers (
ap127-watchdog,ap127-dispatcher) stay onraw.githubusercontent.com— a Worker cannot fetch a same-account*.workers.devURL (CF error 1042; verified — pointingFLIGHT_SRCat the Worker made every watchdog runUpstream HTTP 404). - Watchdog
POST /notify(X-API-Key: dedicatedNOTIFY_KEYsecret) runs the diff immediately; the Pi +fetch_schedule.ymlcall it after a publish, so Telegram fires within seconds. Watchdog cron tightened*/5→*/2as the backstop. - DB001
update-cache.yml→ every 15 min (was every tick):dispatcher/worker.jsshouldDispatchDb001(event.scheduledTime)gates it to :00/:15/:30/:45 (~288 → ~96 runs/day). The dispatcher's own 5-min cron and the CMD_CTR stale-check are unchanged. - Phase 2 — parallel scrape (kept), 3-min cadence (reverted 2026-09-07). The per-date
getStudentScheduleRPC loop runsFETCH_RPC_CONCURRENCYdates at once (=1is byte-identical to serial) — full window ~12 min → ~3 min; that part stays. The paired 3-min timer +STANDBY_MAX_AGE_MIN=3were reverted to 5 min / 6 after the board hard-hung ~17 h (powered, unresponsive) ~4 h in — Chromium memory bloat + no idle window between tight cycles on a 1 GB board also running CUPS.FETCH_RPC_CONCURRENCYeased 4→2; new nightlyap127-chromium-restart.timerrecycles Chromium at 03:00. The cloud fallback covered the whole hang (feed never >44 min stale). Effective refresh ~10–13 min.
flight-schedule-feed/docs/superpowers/specs/2026-09-06-r2-data-plane-decoupling-design.md.The thresholds
| Threshold | Where | Value | Meaning |
|---|---|---|---|
| STANDBY_MAX_AGE_MIN | pi-native/run_fetch.sh | 6 min | Pi fetches unless someone else just committed |
| FETCH_RPC_CONCURRENCY | scripts/fetch_schedule.py + pi .env | 2 (was 4; 1 = serial) | per-date getStudentSchedule RPCs in parallel |
| timer OnUnitActiveSec | pi-native/ap127-fetch.timer | 5 min (briefly 3; reverted after the hang) | how often a cycle wakes; a full fetch takes ~8 min |
| ap127-chromium-restart.timer | pi-native/recycle-chromium.sh | 03:00 local | nightly Chromium recycle — clears memory bloat |
| STALE_TAKEOVER_MIN | dispatcher/worker.js | 35 min | cloud takes over as fallback |
| DATA_STALE_LIMIT_MIN | ap127-watchdog-monitor | 60 min | Telegram pages you |
*/2 cron and its
POST /notify push still mean a detected change reaches Telegram in seconds.fetch_schedule.yml also runs cron 0 */12 * * *
unguarded — a twice-daily proof run, because a fallback that never executes is a fallback
nobody knows is broken.Reliability features
git pull --rebase before push concurrency groups (no overlap) Playwright Chromium cached (~2 min/run saved) --with-deps always (cold-cache safe) failure → auto GitHub Issue (fetch-failure / refresh-failure) commit only on real change KV writes only on real change watchdog tracks feed freshness, not just job successKnown gap — the Pi is not a full peer
The Pi runs the CMD_CTR scrape only. It has no equivalent for DB001's
update-cache.js / build-student.js / push-to-kv.js, so the
progress pipeline is still 100% cloud-dependent. A full Cloudflare/GitHub outage stops progress data
even though flight data keeps flowing.
Issues: Read and write on AP127CMD/CMD_CTR to that token.5.4 The Pi is shared hardware — it is also the house AirPrint server
Added 2026-09-03. The Zero 2W is not AP127-dedicated: it also runs CUPS, sharing a USB-attached Canon PIXMA E410 to every Apple device on the LAN as an AirPrint printer. This is deliberately the opposite of the 3D-printer rule (one board per printer) — a print queue is idle almost all the time, where Klipper is realtime.
| Queue | Canon_E410, shared, advertised as "Canon PIXMA E410 @ DietPi" |
|---|---|
| Driver | Gutenprint 5.3.4 ships a native Canon PIXMA E410 PPD — Canon's proprietary cnijfilter2 is not needed and should not be attempted on arm64 |
| Cost to AP127 | cupsd idles at ~15 MB of the board's 969 MB; a rasterizing job can briefly slow one scrape cycle, which the thresholds above already absorb |
dr_mode = "peripheral" with its companion
EHCI/OHCI disabled — a USB printer plugged in there never enumerates and emits no error explaining
why. Only H5 usbhost overlays ship with the kernel, so a custom H616 one was written:
/boot/overlay-user/usb-otg-host.dtbo, enabled via user_overlays=usb-otg-host
in /boot/dietpiEnv.txt. If a kernel upgrade ever loses that overlay, USB dies
silently — that file is the first place to look.apt-get update && apt-get install running alongside the persistent headless
Chromium crash-rebooted the Pi (watchdog or brownout on a 1 GB board) and left nothing
installed — /var/log is RAMlog, so the apt history was gone too. One package group per
invocation. This is a property of the board under load, not of printing.Two CUPS settings that must stay set: IdleExitTimeout 0 — Debian's
60 s default lets socket-activated cupsd exit, which silently drops the Bonjour advert
(the classic "AirPrint worked yesterday" failure); and systemctl enable cups for the
service, not just the socket. Sharing itself is
cupsctl --remote-admin --share-printers — note --remote-printers is
not a valid cupsctl option and passing it aborts the whole command.
usblp0: removed / re-added pairs in dmesg bracketing a print
job are normal — the CUPS usb backend claims the interface and releases it afterwards. Not a
flapping cable.
06Telegram notification system — Watchdog
The Telegram system is entirely the ap127-watchdog Cloudflare Worker (repo CMDV2/watchdog/). No Telegram code exists anywhere else in the repos. It fetches AP-127 flights, diffs vs the previous snapshot, and sends one Telegram message per change to the affected student. Since 2026-09-06: cron */2 (was */5) plus a POST /notify push endpoint the Pi + CI call after a data publish, so a change reaches Telegram in seconds rather than on the next cron tick. It reads flight-data-recent.js straight from raw.githubusercontent.com (NOT the ap127-data Worker — a Worker cannot fetch a same-account *.workers.dev URL, CF 1042).
CMD_CTR/flight-data-recent.js")]:::k NOTIFY[["POST /notify (Pi + CI)"]]:::w --> F CRON[["cron */2 (backstop)"]]:::w --> F RAW --> F["fetch + filter
batch=AP-127"]:::w PREV[("KV watchdog:snapshot")]:::k --> DIFF F --> DIFF["diffSnapshots()"]:::w DIFF --> EV{"event type"}:::w EV -->|ADDED| TG EV -->|REMOVED| TG EV -->|STATUS| TG EV -->|CHANGED| TG TG["formatMessage +
sendTelegram()"]:::w --> BOT(["Telegram Bot API"]):::k DIFF --> SNAP[("write snapshot
if changed")]:::k TG --> LOG[("KV watchdog:log:YYYY-MM")]:::k
Tracked fields
Event → message
@username via roster config; 1s pause between sends (rate-limit).HTTP API (Watchdog tab)
/status · /config · /log?month=/config · /test · /notify (need X-API-Key; /notify uses NOTIFY_KEY)KV keys (AP127_WD)
watchdog:snapshot — diff basewatchdog:config — roster + prefswatchdog:status — heartbeatwatchdog:log:YYYY-MM[-A/B] — sharded @20MBMessage format (src/telegram.js)
✈️ New flight scheduled SP: @username 📅 10 Jun 2026 08:00–09:30 📖 Lesson: CDGL 04 🛩 HS-NGT | FI: ITTIPOL P.
07Deployment
Sites — confirmed live (all 4 main = CF Pages Git build on push to main)
| Pages project | Repo | Notes |
|---|---|---|
| ap127-cmd-ctr | CMD_CTR | fetch_schedule.yml is a data job, not a deploy |
| ap127-db001 | DB001 | Git-integrated CF Pages. Data commits carry [CI Skip] so the build is skipped; a real code push builds normally. deploy-pages job in update-cache.yml is legacy/dead. |
| ap127-data | flight-schedule-feed/data-worker | npx wrangler deploy — stateless raw.github proxy, no bindings/secrets. Serves the data files to browsers so data commits need no rebuild. |
| ap127-dashboardr1 | DB_Share (private) | content written by sync-dashboardr1.js; redeploys ~hourly |
| ap127-ngt2 | CMDV2 | the live unified SPA + watchdog CORS origin |
| GitHub Pages | Portal | static.yml → ap127cmd.github.io/Portal/ |
Git Provider: Yes in Cloudflare — CF Pages is authoritative; the in-repo GitHub-Pages steps are dead code. Two extra projects ap127-cmdv2 / ap127-cmdv2-ngt-imp1 appear to be old/staging CMDV2 deploys.Workers
| Worker | How deployed |
|---|---|
| ap127-dispatcher | deploy-dispatcher.yml on push to dispatcher/**: wrangler deploy + wrangler secret put GITHUB_PAT. Token CF_WORKERS_TOKEN = Workers Scripts:Edit + Account:Read. |
| ap127-data-api | Manual wrangler deploy (no in-repo workflow), 2026-05-21. Bind KV AP127_STUDENT_DATA (c5c88c81…), set ALLOWED_ORIGIN. |
| ap127-watchdog | wrangler deploy from CMDV2/watchdog/; KV bound by id; 3 secrets via wrangler secret put. |
Confirmed Cloudflare ids & URLs
| Resource | Value |
|---|---|
| Account | ae38e04e56d0ae52d3ec47ad29977587 · anusorn.tanmetha@gmail.com |
| workers.dev subdomain | anusorn-tanmetha.workers.dev |
| KV AP127_STUDENT_DATA | c5c88c813d8d4f668f6081506ad98bcd |
| KV ap127-watchdog-AP127_WD | b42f3202c5364f91aef3837132d6ccd5 |
| Worker ap127-data-api | ap127-data-api.anusorn-tanmetha.workers.dev (= CF_WORKER_URL) |
| Worker ap127-watchdog | ap127-watchdog.anusorn-tanmetha.workers.dev |
| Worker ap127-dispatcher | ap127-dispatcher.anusorn-tanmetha.workers.dev (cron-only) |
08Secrets & environment inventory
Values are never stored — this is where each secret lives and what it does.
GitHub Actions — DB001
| Secret | Purpose |
|---|---|
| RELAY_URL | Apps Script relay (CSV CORS bridge); injected into index.html |
| ADMIN_PASSWORD_HASH | SHA-256 of admin password; injected into index.html |
| CF_ACCOUNT_ID | Cloudflare account id |
| CF_KV_NAMESPACE_ID | KV namespace id for AP127_STUDENT_DATA |
| CF_API_TOKEN | CF token w/ KV write (push-to-kv.js) |
| CF_WORKER_URL | ap127-data-api URL; injected into student.html + DB_Share |
| GH_PAT_DASHBOARDR1 | fine-grained PAT, Contents R/W on DB_Share only |
| CF_WORKERS_TOKEN | CF token, Workers Scripts:Edit + Account:Read (dispatcher deploy) |
| GH_PAT_DISPATCHER | PAT uploaded to dispatcher worker as GITHUB_PAT |
GitHub Actions — CMD_CTR
| Secret | Purpose |
|---|---|
| GH_PAT_WORKFLOW | PAT to workflow_dispatch CMDV2 refresh-data.yml after a fetch |
Cloudflare Worker secrets/vars
| Worker | Secrets / vars |
|---|---|
| ap127-dispatcher | GITHUB_PAT |
| ap127-data-api | ALLOWED_ORIGIN (var) · KV binding |
| ap127-watchdog | TELEGRAM_BOT_TOKEN · TELEGRAM_CHAT_ID · WATCHDOG_API_KEY · KV |
09Reproduce from scratch
Order matters — later steps depend on earlier IDs/URLs.
RELAY_URL). The Flight Ops Portal (→ /exec in fetch_schedule.py) is a third-party academy system — point the scraper at the real portal or your own equivalent.AP127CMD/{CMD_CTR,DB001,DB_Share,CMDV2,Portal}; push code.AP127_STUDENT_DATA and AP127_WD; note ids.ap127-data-api, ap127-watchdog, ap127-dispatcher; bind KV + set secrets.TELEGRAM_BOT_TOKEN; chat id → TELEGRAM_CHAT_ID; generate WATCHDOG_API_KEY; map SP names → @usernames in watchdog:config.main./test message.flightCache. The relay (progress feed) and everything else is fully reproducible from these docs.ap127-watchdog-monitor still has no TELEGRAM_BOT_TOKEN (re-verified 2026-09-02, open 5+ weeks). As of 2026-09-02 two detectors route through that gate — watchdog-down and the new flight-data-staleness alert — so nothing in this ecosystem can page a human. The value is not recoverable from this machine. ·
2. The Pi runs the CMD_CTR scrape only — DB001's cache rebuild + KV push have no Pi equivalent, so the progress pipeline is still fully cloud-dependent. Needs RELAY_URL + CF KV creds in pi-native/.env. ·
3. The Pi's PAT cannot comment on or close issues (verified 403s), so a Pi-side failure opens an unlabelled issue nothing can auto-close — matters more now that the Pi is primary. ·
4. 60 upstream ID collisions in the frozen pre-migration archive (data correct, identity broken; outside the watchdog window, so notifications are unaffected). ·
5. Security: pi-native/mac-monitor/config.json (tracked in a public repo) contains a plaintext VNC password — still unfixed, and it is in git history. (The AP127_Portal/.git/config plaintext-PAT note is resolved 2026-09-07 — dead token, remote switched to keychain auth.) ·
Closed 2026-09-02: the two ap127-cmdv2* duplicate Pages projects were deleted. Still unconfirmed: any WAF rate-limit rule on the data worker · DB_Share index.html manual drift.