/health returns ok. The simpler launch mode returns an app id. A human clicks Connect on the path that needs a vendor configuration id. The API answers 400: that id was never in the process env.
That is not a flaky CDN. It is reachability ≠ merge. Unit tests mocked the client. The sprint row said code done. Production still could not start the mode that depends on a secret outside git.
Invariant: if launch mode M requires config C, then either C is present in the running process and visible in the launch payload, or the API refuses M with a clear error — never a silent empty field for the UI to discover.
Why “shipped” lied
The wrong model is: tests green and the PR merged means production is ready.
Env and vendor dashboard config live outside the repo. Health only proves the process is alive. A basic launch mode that needs only an app id still works. The extended onboarding mode — Login-for-Business-style Embedded Signup that needs a config_id — is the one that fails. Checklist theater closes the code row and leaves “env filled” open.
Lab: health is not the gate
I ran node lab/launch-config-reachability.mjs in this repo on 10 August 2026. In-process launcher. No network.
| Path | What we need | Result |
|---|---|---|
| empty env + extended mode | fail closed | ok: false, reason: "missing_config_id" |
| empty env + basic mode | still works | ok: true (no config required) |
| health alone as ship gate | would ship | healthOnlyWouldShip: true while extended dead |
| config set + extended | reachable | non-empty configId; featureType + session version set |
Trimmed report:
{
"emptyEnv": {
"healthOk": true,
"basicOk": true,
"extended": { "ok": false, "reason": "missing_config_id" },
"healthOnlyWouldShip": true,
"reachable": false
},
"filledEnv": {
"configIdLength": 18,
"featureType": "whatsapp_business_app_onboarding",
"sessionInfoVersion": "3",
"reachable": true
},
"allPassed": true
}
The interesting failure is not that extended returns 400. It is that health and basic launch stay green, so a health-only post-deploy check would have declared victory.
Smoke the live payload
After deploy, one authenticated call to the launch endpoint for the mode you actually ship. Assert shape, not vibes:
modeis the one under testconfig_id(or equivalent) is non-empty — publish length, not the secretfeature_type/ session version match what the vendor wizard promised
Fail closed on the API when config is missing. Do not return an empty id and let the popup fail opaquely.
function launch(mode, env) {
if (mode === "extended" && !env.configId) {
return { ok: false, reason: "missing_config_id" };
}
// build vendor launch payload from env — never invent a blank configId
}
Filling env is not more application code. The cheap check is asserting the payload once under a real token.
Decision boundary
Rejected: “we’ll fill env when someone clicks Connect,” health as the only post-deploy gate, and empty config_id passed through to the UI.
This note is the wrong tool for proving webhook subscriptions or a full human onboard — those are separate gates. Creating the vendor config itself (which checkboxes freeze at create time) is a sibling problem; this note stops at env reachability after the id exists.
What this does not solve
Irreversible Login-for-Business knobs. End-to-end vendor App Review. Whether an existing messaging account has finished human onboarding (enabled flags can stay false until someone completes the wizard).
Merged code needs a smoke that proves the mode you sell can start. Boring on purpose. Green health with a dead Connect button is the exciting bug.