Wittyden
Long-form technical notes, written by operators for operators.
Structuring DNS for Reliability
July 19, 2026
Nameserver outages are devastating because they paralyze systems globally: infrastructure monitors report operational green across every host while end users cannot locate your endpoints. Mitigating this catastrophic blind spot requires standard multi-provider redundancy backed by distinct network backbones.
Designing cache lifetimes requires careful trade-off analysis. Overly aggressive short TTLs increase resolver query volume and amplify upstream blips, whereas prolonged caching windows shelter clients during outages but immobilize scheduled traffic migrations. Differentiating durations by record volatility—short timers for active endpoints, extended timers for static assets—provides balance.
Cache Invalidation Patterns That Survive Traffic
May 7, 2026
Everyone quotes the two hard things joke; fewer people ship caches engineered against stampede. The pattern that matters most is request collapsing: when a hot key expires, exactly one worker rebuilds while everyone else serves slightly stale data, which turns a cache miss from a database incident i…
Why Edge Caching Still Matters in 2026
August 17, 2026
Every few years someone declares the edge cache obsolete: bandwidth is cheap, compute is fast, so why bother? Yet p99 latency keeps telling a different story. The round trip from a user in Sao Paulo to an origin in Frankfurt costs around 200 ms on a good day, and no amount of application optimisatio…
When to Choose a Queue Over a Request
May 1, 2026
Traditional RPC calls remain popular due to straightforward causality: the client makes an invocation, waits for the response, and monitors latency directly. Asynchronous message queuing becomes essential when background tasks outlast active connections or when sudden volume surges threaten to overw…
The Operator's Guide to Load Testing
August 19, 2026
Load tests fail to predict production for a consistent reason: the traffic shape is wrong. Uniform random requests against one endpoint tell you about that endpoint. The production killer is the burst of clients retrying in sync after a thirty-second blip, or the cache-cold crawl at 5 a.m. when the …
More reading
- The Hidden Cost of Chatty Microservices — Engineering, July 20, 2026
- A Field Guide to Graceful Degradation — Engineering, July 27, 2026
- Multi-Region Failover Planning — Operations, June 9, 2026
About us
Our contributors have spent years on-call for large platforms. This site collects the playbooks, postmortems and reference material we wish someone had handed us earlier.