AP127 Flight-Training Ecosystem

Interactive reference — every website, worker, data feed, deployment & the Telegram watchdog. Reconstructed from source.
GitHub org · AP127CMD CF acct ae38e04e… 12 sites · 4 workers · 2 data feeds anusorn-tanmetha.workers.dev ✅ swept live 2026-07-27 ⚠ 3 open items

00System health

Full end-to-end sweep of every site, worker, feed and pipeline — run against live systems on 2026-07-27, ~02:30–03:00 UTC. Each row names how it was checked, so any of it can be re-run.

12/12
sites reachable
119
watchdog tests pass
~5 min
ops feed freshness
3
open items

Verified working

AreaResultHow it was checked
All 12 sites✅ upHTTP 200 on all; DryRun 401 is its password gate — expected, not a fault
Ops feed (CMD_CTR)✅ freshflight-data.js fetchedAt ~2 min old; CI green on the last 12 consecutive runs
Progress feed (DB001)✅ freshcache.json _updated ~5 min old; 4 batches + 3 curricula present
CMDV2 snapshots✅ freshRefresh job green, chained off CMD_CTR — ~5 min cadence, not the documented hourly
Dispatcher✅ runningCron */5; DB001 + CMD_CTR dispatch runs landing on schedule
Watchdog worker✅ healthy/status: healthy:true, lastError:null, runCount 5263, anomalyStreak 0
Watchdog tests✅ 119/119npm test in AP127_V2/watchdog — 3 files, all pass
07-27 stabilization fix✅ deployedwrangler deployments list — live version created 02:29:34Z
Portal structure drift✅ noneportal_fingerprint.json current — 48 RPC fns, 3 Timeline modes, Daily Schedule present
CMDV2 in-browser✅ clean17 views load, zero console errors, now serving p114

Findings

✅ Fixed — duplicate flight rows were inflating every view and every hours KPI CMDV2 p114

Two independent causes, both live, both measured against the real feed rather than inferred.
1. attachCancelDetails() pushed one synthetic row per cancellations[] record — but upstream emits one record per cancel event (its id embeds a submission timestamp), so re-cancelled bookings duplicated: 5 bookingIds → 6 phantom rows, 3 of them AP-127, one rendering 3× with a colliding React key.

2. The same flight ingested twice as two ACTUAL_ONLY rows under different _ACT_<n> ids survived the existing dedup pass — which only ever removes planned rows — and double-counted block hours: 67 rows.
Verified by replaying both preludes over one live feed: rows 3806 → 3733, overall hours inflation 2.81% → 1.11%, AP-127 inflation 0.67% → 0% (1337.9 h → 1328.9 h). Today's numbers are byte-identical before and after — the correction is entirely historical. Studentless MEETING bookings were deliberately excluded from the dedupe (they legitimately repeat at the same time); all 26 verified preserved.
Root cause is upstream and frozen: 425 of the 427 redundant raw rows fall on dates ≤ 2026-07-09 — inside flight_schedule.pre_migration_archive.json, re-applied as an override every run, so it can never self-heal. Walking one commit per day back to 2026-06-15 shows the count climbing daily to 437 by 07-10, then pinned at 425 from 07-11 onward — exactly when that archive was frozen. The post-migration scraper adds none.

⚠ Open — the dead-man's-switch monitor still cannot send alerts

5+ weeks outstanding — re-verified 2026-09-02. wrangler secret list on ap127-watchdog-monitor returns [].
The worker is healthy and its verdict logic is correct — monitor:state reads {alertedDown:false, downStreak:0, reason:"ok"}, and its seemingly-stale lastCheck is by design (it persists on a 6-hour heartbeat to save KV writes). But runMonitor() gates sending on if (alert && env.TELEGRAM_BOT_TOKEN), so with no token it silently no-ops. It detected the 19-hour outage on 2026-07-21 correctly and still could not page anyone. Every future silent watchdog death repeats that blind spot. One command fixes it:
cd /Users/nugui/AP127_V2/watchdog-monitor && npx wrangler secret put TELEGRAM_BOT_TOKEN

✓ Closed 2026-09-02 — CMDV3 retired

The original finding: dispatcher/worker.js targeted only DB001 and CMD_CTR, so CMDV3 lived permanently on its own delayed schedule: cron — measured gaps of 1–2.5 h versus CMDV2's ~5 min, meaning V2 and V3 could show different numbers for the same moment. The proposed fix was one more dispatcher target; it was never applied. The finding is now moot: CMDV3 has been retired entirely (repo archived, Pages project and local dir deleted) because CMDV2 remained the site actually in use. The divergence risk flagged here was the real cost of running two dashboards off one feed — removing the second resolved it more cleanly than syncing it harder would have.

⚠ Open (upstream, low impact) — ID collisions in the frozen pre-migration data

60 ids each map to several genuinely different flights — worst case ACTUAL_ONLY_ with an empty suffix, where 21 distinct flights share one id. Data is correct; identity is not. All 206 such rows fall on 2026-05-05→2026-06-10, outside the watchdog's rolling snapshot window — which matters, because buildSnapshot() keys by id, so a collision inside the window would hide real changes. Notifications are therefore unaffected; the only cost is duplicate React keys when those historical dates render. Fixing it means rewriting ids inside the archive file that is explicitly marked do-not-hand-edit, so it is recorded rather than done.

01Bird's-eye overview

Two independent data domains — Operations (what is flying, when, with whom) and Progress (how far each student has advanced) — are merged only at the presentation layer. Every site is a static front-end; all "backend" work is GitHub Actions cron jobs that commit data files into repos, plus four tiny Cloudflare Workers.

AP124 AP126 AP127 (focus) AP128 AP129
flowchart TD classDef src fill:#141a24,stroke:#3a4658,color:#e6edf3; classDef site fill:#1a1320,stroke:#d850c8,color:#f4d9ef; classDef wkr fill:#15110a,stroke:#f4a050,color:#ffe3c2; classDef data fill:#0e1219,stroke:#2a3340,color:#cdd6e0; GS[("Google Sheets
per-batch CSV")]:::src PORTAL[("Flight Ops Portal
Google Apps Script")]:::src RELAY[/"Apps Script RELAY
CORS bridge"/]:::src GS --> RELAY --> UC["update-cache.js
DB001 · Action"]:::data UC --> CACHE["cache.json
4 batches"]:::data CACHE --> KV[("CF KV
ap127_slice")]:::data KV --> API["ap127-data-api
worker"]:::wkr PI[("Orange Pi Zero 2W
every 5 min, concurrency 2")]:::src PORTAL -->|"getStudentSchedule RPC"| PI PI -->|"git push, CI-Skip"| FS["fetch_schedule.py
+ CI fallback"]:::data FS --> GH[("GitHub main
flight-data*.js")]:::data GH --> RAW2[("raw.githubusercontent")]:::data RAW2 --> DATA["ap127-data Worker
proxy · 60s cache"]:::wkr DATA --> CC["CMD_CTR
ops dashboard"]:::site CACHE --> DB001["DB001
admin site"]:::site DATA --> DBSHARE["DB_Share
student site (/mirror proxy)"]:::site API --> DBSHARE DATA --> V2["CMDV2 · PRIMARY
unified SPA"]:::site API --> V2 CACHE --> V2 RAW2 --> WD["ap127-watchdog
cron 2m + POST notify"]:::wkr WD -->|notify| TG(["Telegram SP"]):::src WD --> V2 WD -.->|"shared KV"| MON["ap127-watchdog-monitor
dead-man switch, cron 10m"]:::wkr MON -.->|"no bot token — cannot alert"| TG DISP["ap127-dispatcher
cron 5m"]:::wkr -.->|"only if feed ≥35m stale"| FS DISP -.->|"only :00/:15/:30/:45"| UC CC -.->|"trigger"| V2 PORTALL["Portal
launcher"]:::site -.links.-> CC PORTALL -.links.-> DB001 PORTALL -.links.-> DBSHARE PORTALL -.links.-> V2
Data is joined by student name, never by array position — upstream reordering can never shift everyone's labels.
Note the dotted warning above. The monitor can observe the watchdog but cannot alert on it — no TELEGRAM_BOT_TOKEN. As of 2026-09-02 a second detector (flight-data staleness) routes through that same missing secret, so nothing can page a human at all. Tracked in System health.

02The twelve websites

All reachable as of 2026-07-27 (DryRun answers 401 by design — it is password-gated). Plus one experimental optimizer (flight-scheduler) that is local-only and not part of the deployed set.

Updated 2026-09-02: CMDV3 and Chatbot were retired — repos archived (recoverable), Pages projects, KV namespaces and local dirs deleted. Also deleted the two stale ap127-cmdv2* duplicate Pages projects. Corrected 2026-07-27 against gh repo list AP127CMD: RPT is private (previously documented public), the repo AP127CMD/FlightTraining does not exist (that site is a direct-upload with no remote), and AP127_Docs and IFR_Flight_Deck were missing from this list entirely.

flight-scheduler — experimental optimizer (not deployed)

Google OR-Tools CP-SAT · FastAPI + SQLAlchemy 2 · PostgreSQL 16 · React/TS/Vite · Docker Compose. Treats scheduling as constraint optimization (airport hours, fleet, maintenance, duty windows, prerequisites, shared-FI, leaves). Local ~/flight-scheduler, no git remote. docker compose up --build seeds 24 students / 8 instructors / 10 aircraft / 101 lessons.

03Cloudflare Workers — 4 services

All free-tier, all in account ae38e04e56d0ae52d3ec47ad29977587.

⏱ ap127-dispatcher

ap127-dispatcher.anusorn-tanmetha.workers.dev · cron */5
JobPOSTs workflow_dispatch to DB001 update-cache.yml (only at :00/:15/:30/:45 via shouldDispatchDb001()) and CMD_CTR fetch_schedule.yml (only if the feed is ≥35 min stale). Does NOT dispatch CMDV2 — CMD_CTR chains that itself.
AuthGITHUB_PAT secret
WhyGitHub cron is ~hourly best-effort; this gives true 5-min cadence. In-repo cron '0 * * * *' is only a fallback.
RepoDB001/dispatcher/

🔌 ap127-data-api live

ap127-data-api.anusorn-tanmetha.workers.dev
JobReads KV ap127_slice, returns JSON, CORS-locked to the student site.
BindKV KV → AP127_STUDENT_DATA (c5c88c81…)
VarALLOWED_ORIGIN = ap127-dashboardr1.pages.dev

📡 ap127-watchdog live

ap127-watchdog.anusorn-tanmetha.workers.dev · cron */5
JobDiffs AP-127 flights, DMs the affected SP via Telegram Bot API; HTTP API backs the CMDV2 Watchdog tab.
BindKV ap127-watchdog-AP127_WD (b42f3202…ccd5)
SecretsTELEGRAM_BOT_TOKEN · TELEGRAM_CHAT_ID · WATCHDOG_API_KEY
Status✅ 2026-07-27 — healthy:true, runCount 5263, no errors, 119/119 tests

🐕 ap127-watchdog-monitor muted

ap127-watchdog-monitor.anusorn-tanmetha.workers.dev · cron */10
JobDead-man's switch for the watchdog. Reads watchdog:status straight from shared KV, not over HTTP — a same-account Worker→workers.dev fetch is blocked (CF error 1042), and a CPU-killed watchdog stops writing KV, so a frozen lastRun is the truest death signal.
LogicTwo consecutive unhealthy checks (~20 min) → one alert; recovery → one all-clear. Transitions only. Persists on a 6 h heartbeat, so a stale lastCheck is normal.
Status⚠ Running and observing correctly, but cannot send — wrangler secret list returns []. Open since 2026-07-17.
Fixcd AP127_V2/watchdog-monitor && npx wrangler secret put TELEGRAM_BOT_TOKEN
Judging the dispatcher: a plain GET / on ap127-dispatcher returns HTTP 500. That is expected — it is a cron-only Worker with no meaningful fetch handler. Judge it by whether its dispatch targets actually run, not by that response.

ap127-data-api worker source

export default {
  async fetch(request, env) {
    const allowedOrigin = env.ALLOWED_ORIGIN || '*';
    if (request.method === 'OPTIONS') return new Response(null, { headers: {
      'Access-Control-Allow-Origin': allowedOrigin,
      'Access-Control-Allow-Methods': 'GET', 'Access-Control-Max-Age': '86400' }});
    if (request.method !== 'GET') return new Response('Method not allowed', { status: 405 });
    const data = await env.KV.get('ap127_slice', 'json');
    if (!data) return new Response(JSON.stringify({ error: 'No data' }), { status: 503,
      headers: { 'Content-Type': 'application/json' }});
    return new Response(JSON.stringify(data), { headers: {
      'Content-Type': 'application/json', 'Cache-Control': 'no-store',
      'Access-Control-Allow-Origin': allowedOrigin }});
  },
};

04Data sources & feeds

Operations feed → flight-data.js

A Google Apps Script web-app renders the academy Flight Operations Portal. The schedule lives in a JS object flightCache inside a sandboxed cross-origin iframe — so a headless Playwright browser is needed, not a plain GET.

URLhttps://script.google.com/macros/s/AKfycbzsOcPHLUpD5U8Qyq-x78edIOMUr28NJAp0KTvJvYCW6IQ_yG-HB97aRue8aFoxGQ5lJg/exec
1fetch_schedule.py — launch headless Chromium, goto /exec (90s budget; GAS cold-starts), wait for the iframe src, enumerate page.frames, frame.evaluate() out flightCache. Up to 3 attempts, 20/40s backoff.
2validate_raw_cache() — hard-fail on schema break (data not saved). Then normalize_entry() per row: -→null, derive isSimulator/isStandby/durationMin.
3Merge into data/flight_schedule.json — fresh dates overwrite, dates outside the rolling ~10-day window preserved. Rolling backup written first.
4generate_flight_data.js → flight-data.js (window.FLIGHT_DATA), strips actuals from non-Completed flights, appends ?v=<unix-ts> cache token. Recovery: rebuild_history.py replays all git commits.

Progress feed → cache.json

Google Sheets published as per-batch CSV, read through a second Apps Script relay (CORS bridge = the RELAY_URL secret).

1update-cache.js fetches RELAY_URL + "?batch=" + batch for each batch.
2parseCSV() — 3-rows-per-student format; silently drops any lesson/entry matching /^AUPRT/i (never counted anywhere).
3runScheduler() — projects each student's planned[], next_lesson, finish date, remaining, pct (workday/holiday-aware).
4Writes cache.json — keys ap124 ap126 ap127 ap129 monthly cur124 cur126 cur127 cap _updated. push-to-kv.js PUTs only {ap127,cur127,_updated} to KV ap127_slice.
✅ Relay Apps Script (supplied). A trivial CORS bridge — maps ?batch= to a published Google Sheet CSV (gid 416854743) and proxies it as text/plain. Deploy as Web app · Execute as Me · Anyone → that /exec URL is the RELAY_URL secret.
const CSV_URLS = {
  AP124: "docs.google.com/spreadsheets/d/e/2PACX-1vRQK_to…/pub?gid=416854743&output=csv",
  AP126: "docs.google.com/spreadsheets/d/e/2PACX-1vQB8PiZ…/pub?gid=416854743&output=csv",
  AP127: "docs.google.com/spreadsheets/d/e/2PACX-1vQNxzCi…/pub?gid=416854743&output=csv",
};
function doGet(e){
  const url = CSV_URLS[(e.parameter.batch||"").toUpperCase()];
  if(!url) return json({error:"Unknown batch"});
  return ContentService.createTextOutput(UrlFetchApp.fetch(url).getContentText("UTF-8"))
           .setMimeType(ContentService.MimeType.TEXT);
}
✅ Ops Portal = third-party. The AKfycbzs…/exec Apps Script is an academy system the owner cannot access — the operations feed is an external black box scraped by Playwright. This is the system's single most fragile point: if the academy changes the portal, fetch_schedule.py breaks (→ fetch-failure Issue).

runScheduler — greedy capacity simulator (not a fetch)

After parsing the CSVs, update-cache.js projects every student's future lessons forward to compute planned[] · finish · monthly.

BATCHES = AP124 · AP126 · AP127 (only real feeds) cap 25 slots/operating-day horizon 800 workdays priority AP124→126→127→129 rest-gap 2d if last ≥120min else 1d rank by remaining ÷ workdays-left
AP129 is synthetic — not a feed. CFG.n129:13 placeholder students (AP129-01 "Student 01", 0 done) are generated in-code from ap129start 2026-06-01, using the AP127 curriculum as a stand-in — a capacity projection of a batch that hasn't started. No missing CSV.
Identity arrays hard-coded in update-cache.js: AP127_NICKS / AP127_FI / AP127_SE (28 index-matched entries) + AP127_FI_FULL map, assigned to AP127 students by position. CSV student rows kept only if catc_id starts with 681; /^AUPRT/i lessons dropped.

05Auto-fetching mechanism

Pi-primary since 2026-09-02. The Orange Pi Zero 2W does the scraping; Cloudflare + GitHub Actions are the automatic fallback. Before that the roles were reversed.

2026-09-06 — data plane decoupled from Cloudflare Pages builds; latency cut. The ecosystem was running at ~14,500 CF Pages builds/month against the 500/mo free cap (ap127-cmd-ctr ~95/day + ap127-ngt2 ~93/day + ap127-db001 ~296/day) — every data commit triggered a full rebuild. Fixes, no payment method added:
  • New ap127-data Worker (flight-schedule-feed/data-worker/) — a stateless proxy that re-serves flight-data.js / flight-data-recent.js / cache.json from raw.githubusercontent.com with a browser Content-Type + 60 s edge cache (Range sliced locally, ETag/304, stale-fallback). No bindings, no storage, no secrets. R2 was the original design — dropped because enabling R2 needs a card on file.
  • Browsers (ap127-cmd-ctr index.html; ap127-ngt2 index/legacy/ops/crosscheck/overview; DB_Share's /mirror proxy) load the data from the Worker. CMDV2's flight-data.js mirror file was deleted.
  • [CI Skip] in every data-commit message → Cloudflare Pages skips the build (deployment status idle, off the 500/mo cap). Verified live on ap127-db001 + ap127-cmd-ctr. Build watch paths were the plan but the Pages API silently drops path_excludes (dashboard-only). A real code push must NOT contain [CI Skip].
  • Backend Workers (ap127-watchdog, ap127-dispatcher) stay on raw.githubusercontent.com — a Worker cannot fetch a same-account *.workers.dev URL (CF error 1042; verified — pointing FLIGHT_SRC at the Worker made every watchdog run Upstream HTTP 404).
  • Watchdog POST /notify (X-API-Key: dedicated NOTIFY_KEY secret) runs the diff immediately; the Pi + fetch_schedule.yml call it after a publish, so Telegram fires within seconds. Watchdog cron tightened */5 → */2 as the backstop.
  • DB001 update-cache.yml → every 15 min (was every tick): dispatcher/worker.js shouldDispatchDb001(event.scheduledTime) gates it to :00/:15/:30/:45 (~288 → ~96 runs/day). The dispatcher's own 5-min cron and the CMD_CTR stale-check are unchanged.
  • Phase 2 — parallel scrape (kept), 3-min cadence (reverted 2026-09-07). The per-date getStudentSchedule RPC loop runs FETCH_RPC_CONCURRENCY dates at once (=1 is byte-identical to serial) — full window ~12 min → ~3 min; that part stays. The paired 3-min timer + STANDBY_MAX_AGE_MIN=3 were reverted to 5 min / 6 after the board hard-hung ~17 h (powered, unresponsive) ~4 h in — Chromium memory bloat + no idle window between tight cycles on a 1 GB board also running CUPS. FETCH_RPC_CONCURRENCY eased 4→2; new nightly ap127-chromium-restart.timer recycles Chromium at 03:00. The cloud fallback covered the whole hang (feed never >44 min stale). Effective refresh ~10–13 min.
Combined Pages builds: ~14,500/mo → real code deploys only (~30–40/mo). Design: flight-schedule-feed/docs/superpowers/specs/2026-09-06-r2-data-plane-decoupling-design.md.
sequenceDiagram participant PI as Orange Pi Zero 2W (*/5) PRIMARY participant D as ap127-dispatcher (*/5) participant CC as CMD_CTR Action FALLBACK participant DB as DB001 Action participant V2 as CMDV2 Action participant WD as ap127-watchdog participant DATA as ap127-data Worker participant SH as DB_Share repo PI->>PI: gate: feed younger than 6 min ? stand by PI->>PI: fetch_schedule.py (CDP Chromium, concurrency 2) PI->>CC: git push flight-data.js commit ending CI-Skip Note over CC: CI-Skip token means Pages build skipped (status idle) PI->>V2: workflow_dispatch refresh-data.yml PI->>WD: POST /notify (X-API-Key) D->>DB: workflow_dispatch update-cache.yml (only :00/:15/:30/:45) DB->>DB: update-cache.js, commit also ending CI-Skip D->>CC: workflow_dispatch fetch_schedule.yml Note over D,CC: ONLY if the feed is >= 35 min stale CC->>V2: workflow_dispatch refresh-data.yml V2->>V2: refresh_snapshots.mjs to progress-data.js + ngt-data.js DATA-->>CC: browsers GET flight-data.js (proxies raw.github) DATA-->>V2: browsers GET flight-data.js DATA-->>SH: /mirror proxy GET flight-data.js / cache.json

The thresholds

ThresholdWhereValueMeaning
STANDBY_MAX_AGE_MINpi-native/run_fetch.sh6 minPi fetches unless someone else just committed
FETCH_RPC_CONCURRENCYscripts/fetch_schedule.py + pi .env2 (was 4; 1 = serial)per-date getStudentSchedule RPCs in parallel
timer OnUnitActiveSecpi-native/ap127-fetch.timer5 min (briefly 3; reverted after the hang)how often a cycle wakes; a full fetch takes ~8 min
ap127-chromium-restart.timerpi-native/recycle-chromium.sh03:00 localnightly Chromium recycle — clears memory bloat
STALE_TAKEOVER_MINdispatcher/worker.js35 mincloud takes over as fallback
DATA_STALE_LIMIT_MINap127-watchdog-monitor60 minTelegram pages you
Effective schedule refresh is a fetch every ~10–13 min (was ~12–18; briefly ~3–6 on the 3-min timer before the hang). The watchdog's */2 cron and its POST /notify push still mean a detected change reaches Telegram in seconds.
CMD_CTR's fetch_schedule.yml also runs cron 0 */12 * * * unguarded — a twice-daily proof run, because a fallback that never executes is a fallback nobody knows is broken.

Reliability features

git pull --rebase before push concurrency groups (no overlap) Playwright Chromium cached (~2 min/run saved) --with-deps always (cold-cache safe) failure → auto GitHub Issue (fetch-failure / refresh-failure) commit only on real change KV writes only on real change watchdog tracks feed freshness, not just job success

Known gap — the Pi is not a full peer

The Pi runs the CMD_CTR scrape only. It has no equivalent for DB001's update-cache.js / build-student.js / push-to-kv.js, so the progress pipeline is still 100% cloud-dependent. A full Cloudflare/GitHub outage stops progress data even though flight data keeps flowing.

Open: the Pi's fine-grained PAT can create and read issues but cannot comment or close them, so a Pi-side failure opens an unlabelled issue nothing can auto-close. Grant Issues: Read and write on AP127CMD/CMD_CTR to that token.

5.4  The Pi is shared hardware — it is also the house AirPrint server

Added 2026-09-03. The Zero 2W is not AP127-dedicated: it also runs CUPS, sharing a USB-attached Canon PIXMA E410 to every Apple device on the LAN as an AirPrint printer. This is deliberately the opposite of the 3D-printer rule (one board per printer) — a print queue is idle almost all the time, where Klipper is realtime.

QueueCanon_E410, shared, advertised as "Canon PIXMA E410 @ DietPi"
DriverGutenprint 5.3.4 ships a native Canon PIXMA E410 PPD — Canon's proprietary cnijfilter2 is not needed and should not be attempted on arm64
Cost to AP127cupsd idles at ~15 MB of the board's 969 MB; a rasterizing job can briefly slow one scrape cycle, which the thresholds above already absorb
This changed the Pi's boot config, so it is load-bearing for AP127 too. The Zero 2W's USB-C data port ships as dr_mode = "peripheral" with its companion EHCI/OHCI disabled — a USB printer plugged in there never enumerates and emits no error explaining why. Only H5 usbhost overlays ship with the kernel, so a custom H616 one was written: /boot/overlay-user/usb-otg-host.dtbo, enabled via user_overlays=usb-otg-host in /boot/dietpiEnv.txt. If a kernel upgrade ever loses that overlay, USB dies silently — that file is the first place to look.
Installing anything on this board must be staged. A single apt-get update && apt-get install running alongside the persistent headless Chromium crash-rebooted the Pi (watchdog or brownout on a 1 GB board) and left nothing installed — /var/log is RAMlog, so the apt history was gone too. One package group per invocation. This is a property of the board under load, not of printing.

Two CUPS settings that must stay set: IdleExitTimeout 0 — Debian's 60 s default lets socket-activated cupsd exit, which silently drops the Bonjour advert (the classic "AirPrint worked yesterday" failure); and systemctl enable cups for the service, not just the socket. Sharing itself is cupsctl --remote-admin --share-printers — note --remote-printers is not a valid cupsctl option and passing it aborts the whole command.

usblp0: removed / re-added pairs in dmesg bracketing a print job are normal — the CUPS usb backend claims the interface and releases it afterwards. Not a flapping cable.

06Telegram notification system — Watchdog

The Telegram system is entirely the ap127-watchdog Cloudflare Worker (repo CMDV2/watchdog/). No Telegram code exists anywhere else in the repos. It fetches AP-127 flights, diffs vs the previous snapshot, and sends one Telegram message per change to the affected student. Since 2026-09-06: cron */2 (was */5) plus a POST /notify push endpoint the Pi + CI call after a data publish, so a change reaches Telegram in seconds rather than on the next cron tick. It reads flight-data-recent.js straight from raw.githubusercontent.com (NOT the ap127-data Worker — a Worker cannot fetch a same-account *.workers.dev URL, CF 1042).

flowchart LR classDef w fill:#15110a,stroke:#f4a050,color:#ffe3c2; classDef k fill:#0e1219,stroke:#2a3340,color:#cdd6e0; RAW[("raw.githubusercontent
CMD_CTR/flight-data-recent.js")]:::k NOTIFY[["POST /notify (Pi + CI)"]]:::w --> F CRON[["cron */2 (backstop)"]]:::w --> F RAW --> F["fetch + filter
batch=AP-127"]:::w PREV[("KV watchdog:snapshot")]:::k --> DIFF F --> DIFF["diffSnapshots()"]:::w DIFF --> EV{"event type"}:::w EV -->|ADDED| TG EV -->|REMOVED| TG EV -->|STATUS| TG EV -->|CHANGED| TG TG["formatMessage +
sendTelegram()"]:::w --> BOT(["Telegram Bot API"]):::k DIFF --> SNAP[("write snapshot
if changed")]:::k TG --> LOG[("KV watchdog:log:YYYY-MM")]:::k

Tracked fields

Any change fires an event:
datestartendstatusinstructortaillesson
Not tracked: actuals (tkoff, ldgTime, airborne, to, ldg, inst), cond, isSim, isStandby, durMin, duration.

Event → message

ADDED✈️ New flight scheduled
REMOVED❌ Flight cancelled
STATUS🔄 Status update Pending→Completed
CHANGED⚠️ Flight updated (time/date/aircraft/FI/lesson)
SP name → @username via roster config; 1s pause between sends (rate-limit).

HTTP API (Watchdog tab)

GET/status · /config · /log?month=
POST/config · /test · /notify (need X-API-Key; /notify uses NOTIFY_KEY)
CORS allow-list: ap127-ngt2.pages.dev

KV keys (AP127_WD)

watchdog:snapshot — diff base
watchdog:config — roster + prefs
watchdog:status — heartbeat
watchdog:log:YYYY-MM[-A/B] — sharded @20MB

Message format (src/telegram.js)

✈️ New flight scheduled
SP: @username
📅 10 Jun 2026  08:00–09:30
📖 Lesson: CDGL 04
🛩 HS-NGT  |  FI: ITTIPOL P.

07Deployment

Sites — confirmed live (all 4 main = CF Pages Git build on push to main)

Pages projectRepoNotes
ap127-cmd-ctrCMD_CTRfetch_schedule.yml is a data job, not a deploy
ap127-db001DB001Git-integrated CF Pages. Data commits carry [CI Skip] so the build is skipped; a real code push builds normally. deploy-pages job in update-cache.yml is legacy/dead.
ap127-dataflight-schedule-feed/data-workernpx wrangler deploy — stateless raw.github proxy, no bindings/secrets. Serves the data files to browsers so data commits need no rebuild.
ap127-dashboardr1DB_Share (private)content written by sync-dashboardr1.js; redeploys ~hourly
ap127-ngt2CMDV2the live unified SPA + watchdog CORS origin
GitHub PagesPortalstatic.yml → ap127cmd.github.io/Portal/
Drift resolved: every main site shows Git Provider: Yes in Cloudflare — CF Pages is authoritative; the in-repo GitHub-Pages steps are dead code. Two extra projects ap127-cmdv2 / ap127-cmdv2-ngt-imp1 appear to be old/staging CMDV2 deploys.

Workers

WorkerHow deployed
ap127-dispatcherdeploy-dispatcher.yml on push to dispatcher/**: wrangler deploy + wrangler secret put GITHUB_PAT. Token CF_WORKERS_TOKEN = Workers Scripts:Edit + Account:Read.
ap127-data-apiManual wrangler deploy (no in-repo workflow), 2026-05-21. Bind KV AP127_STUDENT_DATA (c5c88c81…), set ALLOWED_ORIGIN.
ap127-watchdogwrangler deploy from CMDV2/watchdog/; KV bound by id; 3 secrets via wrangler secret put.

Confirmed Cloudflare ids & URLs

ResourceValue
Accountae38e04e56d0ae52d3ec47ad29977587 · anusorn.tanmetha@gmail.com
workers.dev subdomainanusorn-tanmetha.workers.dev
KV AP127_STUDENT_DATAc5c88c813d8d4f668f6081506ad98bcd
KV ap127-watchdog-AP127_WDb42f3202c5364f91aef3837132d6ccd5
Worker ap127-data-apiap127-data-api.anusorn-tanmetha.workers.dev (= CF_WORKER_URL)
Worker ap127-watchdogap127-watchdog.anusorn-tanmetha.workers.dev
Worker ap127-dispatcherap127-dispatcher.anusorn-tanmetha.workers.dev (cron-only)

08Secrets & environment inventory

Values are never stored — this is where each secret lives and what it does.

GitHub Actions — DB001

SecretPurpose
RELAY_URLApps Script relay (CSV CORS bridge); injected into index.html
ADMIN_PASSWORD_HASHSHA-256 of admin password; injected into index.html
CF_ACCOUNT_IDCloudflare account id
CF_KV_NAMESPACE_IDKV namespace id for AP127_STUDENT_DATA
CF_API_TOKENCF token w/ KV write (push-to-kv.js)
CF_WORKER_URLap127-data-api URL; injected into student.html + DB_Share
GH_PAT_DASHBOARDR1fine-grained PAT, Contents R/W on DB_Share only
CF_WORKERS_TOKENCF token, Workers Scripts:Edit + Account:Read (dispatcher deploy)
GH_PAT_DISPATCHERPAT uploaded to dispatcher worker as GITHUB_PAT

GitHub Actions — CMD_CTR

SecretPurpose
GH_PAT_WORKFLOWPAT to workflow_dispatch CMDV2 refresh-data.yml after a fetch

Cloudflare Worker secrets/vars

WorkerSecrets / vars
ap127-dispatcherGITHUB_PAT
ap127-data-apiALLOWED_ORIGIN (var) · KV binding
ap127-watchdogTELEGRAM_BOT_TOKEN · TELEGRAM_CHAT_ID · WATCHDOG_API_KEY · KV

09Reproduce from scratch

Order matters — later steps depend on earlier IDs/URLs.

1Google data layer — progress Sheets published as CSV; deploy the relay Apps Script (→ RELAY_URL). The Flight Ops Portal (→ /exec in fetch_schedule.py) is a third-party academy system — point the scraper at the real portal or your own equivalent.
2GitHub org + repos — create AP127CMD/{CMD_CTR,DB001,DB_Share,CMDV2,Portal}; push code.
3Cloudflare KV — create AP127_STUDENT_DATA and AP127_WD; note ids.
4Workers — deploy ap127-data-api, ap127-watchdog, ap127-dispatcher; bind KV + set secrets.
5Telegram — @BotFather bot → TELEGRAM_BOT_TOKEN; chat id → TELEGRAM_CHAT_ID; generate WATCHDOG_API_KEY; map SP names → @usernames in watchdog:config.
6GitHub secrets — add every secret from the Secrets tab to the right repo.
7Cloudflare Pages — create 4 Pages projects connected to their repos, building on push to main.
8Portal — enable GitHub Pages on AP127CMD/Portal (static.yml deploys).
9Kick the pipeline — run update-cache.yml + fetch_schedule.yml once; verify cache.json, flight-data.js, KV ap127_slice, and a Telegram /test message.
Ops-feed caveat: step 1's Flight Ops Portal is a third-party academy system — a clone can't recreate it; you must get portal access or substitute your own schedule source exposing an equivalent flightCache. The relay (progress feed) and everything else is fully reproducible from these docs.
Open items (full detail in System health): 1. ⚠ Now the #1 risk. ap127-watchdog-monitor still has no TELEGRAM_BOT_TOKEN (re-verified 2026-09-02, open 5+ weeks). As of 2026-09-02 two detectors route through that gate — watchdog-down and the new flight-data-staleness alert — so nothing in this ecosystem can page a human. The value is not recoverable from this machine. · 2. The Pi runs the CMD_CTR scrape only — DB001's cache rebuild + KV push have no Pi equivalent, so the progress pipeline is still fully cloud-dependent. Needs RELAY_URL + CF KV creds in pi-native/.env. · 3. The Pi's PAT cannot comment on or close issues (verified 403s), so a Pi-side failure opens an unlabelled issue nothing can auto-close — matters more now that the Pi is primary. · 4. 60 upstream ID collisions in the frozen pre-migration archive (data correct, identity broken; outside the watchdog window, so notifications are unaffected). · 5. Security: pi-native/mac-monitor/config.json (tracked in a public repo) contains a plaintext VNC password — still unfixed, and it is in git history. (The AP127_Portal/.git/config plaintext-PAT note is resolved 2026-09-07 — dead token, remote switched to keychain auth.) · Closed 2026-09-02: the two ap127-cmdv2* duplicate Pages projects were deleted. Still unconfirmed: any WAF rate-limit rule on the data worker · DB_Share index.html manual drift.