Daily Engineering Intelligence

Technology Newspaper

← Back to Home

Engineering Intelligence β€” Tuesday

October 6, 2026

πŸ”₯ What Changed

Light Tuesday. No Go items today: four open compiler/runtime bugs are on hold until their fixes merge. Today's picks: two config/control-plane postmortems, a K8s RBAC study across 65k clusters, and an OAuth issuer-binding advisory from the MCP SDKs. One of the bugs in that advisory (LiteLLM) is being actively exploited.

Railway + Firebase: two postmortems, one lesson

Railway (Sep 30): For ~5 minutes, every Railway domain in every region returned 404 to new connections. Railway had recently split its routing control plane out of the monolith. That ran the schema migration and the binary rollout in parallel, which the monolith used to do in order. The new binary came up first and queried columns that didn't exist yet. The proxies then treated "lookup failed" the same as "no routes" and returned 404. The same change had also altered fallback behaviour, so they couldn't fall back to the last route they knew.

Firebase (outage Sep 28, post Oct 2): A 2-line cleanup removed a stale remote-config flag, but the iOS SDK still referenced it. GA4F apps crashed once on launch. The kill-switch channel itself ships globally, so the bad config reached a wide set of apps and stayed live for 2h11m. Dashboards stayed green because they only track server-side health. Existing tests passed.

Why you care: Board DistSys 48. Both outages are the same failure class, with rules worth stealing:
  • A failed dependency must not look like an authoritative empty answer. Keep a last-known-good fallback.
  • Splitting a monolith removes ordering guarantees you got for free. Railway's fix is a binary that refuses to start until its schema is in place.
  • Config and kill switches need canaries and phased rollouts, like code.
  • Health checks must include client-side signals, or your status page will be wrong.
Read on Railway Blog β†’  Β·  Read on Firebase Blog β†’

Dangerous RBAC on Kubernetes built-in principals, 65k clusters

Datadog Security Labs (Oct 5) studied RBAC bindings to system:anonymous, system:unauthenticated and system:authenticated across 65,000+ clusters from ~10,000 organisations. Of 320k bindings, 265k are upstream defaults. Of the rest, 11k still reference PodSecurityPolicy, which was removed in v1.25: a sign nobody maintains that RBAC. Removing those leaves 44k, and 3,500+ of them grant at least one dangerous permission. Distro defaults matter:

  • AKS runs --anonymous-auth=false.
  • EKS (β‰₯1.32) and GKE (β‰₯1.35) restrict which endpoints anonymous users can reach.
  • On GKE, any valid Google account counts as system:authenticated by default. You can only opt out when creating the cluster.

Namespace-scoped bindings to system:unauthenticated can expose the whole cluster when tenants get "namespace as a service".

Why you care: Board K8s 15 + IAM rising. "Authenticated" is not an authorization boundary. Quick audit: list bindings to these three principals, old PSP-era roles, and namespace-level grants to unauthenticated users. Also check your distro's anonymous-auth default.
Read on Datadog Security Labs β†’

MCP OAuth: clients trusted the server about where to send credentials

GHSA-qx49-fqc8-xw99 (Sep 28, High 7.5, no CVE): the MCP Python SDK's OAuth client did not check the authorization server's issuer on every discovery path. In 2.0–2.1.1 it was skipped on the legacy fallback and on the 403 insufficient_scope step-up. Stored credentials also weren't tied to their issuer. As a result, a malicious MCP server could receive the client_secret, the authorization code and the PKCE verifier. WorkOS (Oct 2) groups it with two related bugs:

  • rmcp CVE-2026-63127: the client didn't check that the RFC 9728 resource field matched the server it connected to.
  • LiteLLM CVE-2026-59822: the reverse direction. A made-up Authorization header got requests through on /mcp/. This one is in CISA KEV, so it is actively exploited.
Why you care: Auth/IAM rising. The lessons apply to any OAuth client or gateway:
  • Take the expected issuer and resource from config, never from the response you're validating.
  • Validate on every path, including fallbacks and step-up.
  • Store credentials keyed by issuer.
  • Treat a missing field as a failure.
Action:
  • Python SDK: upgrade to mcp β‰₯1.30.0 / β‰₯2.2.0. Also pass issuer= to ClientCredentialsOAuthProvider / PrivateKeyJWTOAuthProvider, or the upgrade changes nothing for them. Clear old stored registrations.
  • rmcp: upgrade to β‰₯2.0.0.
  • LiteLLM: upgrade to β‰₯1.84.0, or block /mcp/ until you can.
  • No CVE was issued for the Python SDK fix, so CVE-based scanners won't flag it.
Read the GitHub advisory β†’  Β·  WorkOS summary β†’

πŸ“š Worth Your Time

Scaling Kubernetes workloads with Node Swap

Kubernetes Blog (Oct 5): first upstream benchmarks for node swap, which has been GA since v1.34. The setup is failSwapOn: false + memorySwap.swapBehavior: LimitedSwap, with swap on Local SSD and Burstable pods (memory limit > request). Results:

  • A kernel build fits in 300 MB instead of 600 MB with no slowdown (374s vs 433s). At 200 MB it runs >40% slower, because the active working set ends up in swap.
  • runc pods per c4-standard-32 node: 512 β†’ 768.
  • gVisor headless Chrome: 80 β†’ 160.
  • gVisor Python sandboxes: 80 β†’ 240.
  • At peak density, latency comes from CPU contention rather than swap I/O.
Why you care: Board K8s 15. Clear guidance on when swap is safe: it covers idle and burst memory, not the working set. Swap used to be discouraged because cgroup v1 counted memory and swap as one combined limit. Ignore the agent-sandbox pitch (Google authors, GKE-flavoured) and keep the mechanics and the 200 MB cliff.
Read on Kubernetes Blog β†’