Articles
Worth reading
A short take on pieces worth your time — architecture, AI, system design, and engineering careers — each one linked back to its original source.
5 articles
In-House LLM Serving at Netflix
Most companies just call a hosted LLM API and move on; Netflix's AI Platform team explains why and how they run the full inference stack themselves instead, inside their existing production environment rather than spinning up a separate ML silo. What's genuinely useful here is the decision-making, not just the architecture — the tradeoffs they weighed on cost, latency, and operational ownership before choosing to build rather than buy. Good grounding if you're evaluating whether your org actually needs to run its own inference infrastructure.
Netflix Tackles Data Deletion at Scale with Centralized Platform Architecture
InfoQ's coverage of a QCon talk on a problem almost nobody designs for up front: how do you actually delete data, correctly and completely, across dozens of heterogeneous storage systems, at a scale where 76.8 billion row deletions across 1,300 datasets is a normal workload? Netflix's centralized deletion platform is a reminder that "delete" is a distributed systems problem with its own consistency and observability requirements, not an afterthought bolted onto each service. Especially relevant if you're anywhere near GDPR or CCPA compliance work and have been treating deletion as someone else's problem.
How Discord Stores Trillions of Messages
Bo Ingram's account of Discord's migration off vanilla Cassandra to a Rust-based data service layer is one of the more honest scaling stories out there — no rewrite-everything triumphalism, just a clear-eyed walk through where their original design started to show cracks at trillion-row scale and what they changed. The details on data modeling for a chat workload (message ordering, hot partitions, compaction pressure) are transferable to basically any high-write-volume, append-heavy system. One of the better "how it actually broke and what we did about it" posts I've read.
Avoiding fallback in distributed systems
A genuinely counterintuitive argument from AWS's own playbook: fallback logic — the code path meant to save you when the primary system fails — is often the riskier system, precisely because it almost never runs and so almost never gets tested under real conditions. I've seen this bite teams that were proud of their "graceful degradation" story; the fix argued for here is investing that same engineering effort into making the primary path more reliable instead of building a second, untested one. Changed how I run resilience design reviews.
The Tail at Scale
The paper that explains why your p50 latency dashboard is lying to you: at scale, it's the tail — p99, p999 — that decides whether users experience your service as fast, because any single request has to survive every slow component in its path. Dean and Barroso's techniques for taming it, especially hedged and tied requests, are still the starting point for anyone designing a fan-out call pattern or arguing for stricter SLOs on a critical dependency. Dense, but every page earns its place — required reading before you design anything with fan-out on the critical path.