Articles
Worth reading
A short take on pieces worth your time — architecture, AI, system design, and engineering careers — each one linked back to its original source.
22 articles
Avoiding fallback in distributed systems
A genuinely counterintuitive argument from AWS's own playbook: fallback logic — the code path meant to save you when the primary system fails — is often the riskier system, precisely because it almost never runs and so almost never gets tested under real conditions. I've seen this bite teams that were proud of their "graceful degradation" story; the fix argued for here is investing that same engineering effort into making the primary path more reliable instead of building a second, untested one. Changed how I run resilience design reviews.
The Tail at Scale
The paper that explains why your p50 latency dashboard is lying to you: at scale, it's the tail — p99, p999 — that decides whether users experience your service as fast, because any single request has to survive every slow component in its path. Dean and Barroso's techniques for taming it, especially hedged and tied requests, are still the starting point for anyone designing a fan-out call pattern or arguing for stricter SLOs on a critical dependency. Dense, but every page earns its place — required reading before you design anything with fan-out on the critical path.