Railway (Sep 30): For ~5 minutes, every Railway domain in every region returned 404 to new connections. Railway had recently split its routing control plane out of the monolith. That ran the schema migration and the binary rollout in parallel, which the monolith used to do in order. The new binary came up first and queried columns that didn't exist yet. The proxies then treated "lookup failed" the same as "no routes" and returned 404. The same change had also altered fallback behaviour, so they couldn't fall back to the last route they knew.
Firebase (outage Sep 28, post Oct 2): A 2-line cleanup removed a stale remote-config flag, but the iOS SDK still referenced it. GA4F apps crashed once on launch. The kill-switch channel itself ships globally, so the bad config reached a wide set of apps and stayed live for 2h11m. Dashboards stayed green because they only track server-side health. Existing tests passed.
- A failed dependency must not look like an authoritative empty answer. Keep a last-known-good fallback.
- Splitting a monolith removes ordering guarantees you got for free. Railway's fix is a binary that refuses to start until its schema is in place.
- Config and kill switches need canaries and phased rollouts, like code.
- Health checks must include client-side signals, or your status page will be wrong.