Articles

Why Fewer Syscalls Don't Make io_uring Faster Than epoll
io_uring can reduce system calls without improving throughput or tail latency. Compare its completion model, batching, …

10M-Token Context Windows: What Fits Is Not What Works
Maximum context length, effective context, and economically servable context are three different things. The U-shaped …

Redis Distributed Locks: Fine for Efficiency, Not Enough for Correctness
Redis distributed locks explained: SET NX PX basics, TTL expiry and stale clients, watchdog limits, the Redlock debate, …

mmap() vs read(): When Zero-Copy Is Actually Slower
Is mmap() faster than read()? Compare copying, page faults, readahead, TLB pressure, tail latency, and the workloads …

fork() in Multithreaded Programs: The Child Gets the State, Not the Threads
When a multithreaded process calls fork(), the child keeps locked mutexes, half-flushed buffers, and shared file …
Claude Opus 4.8 Released, But Not Showing Up in Claude Code? Here's the Fix
Anthropic released Claude Opus 4.8, but the model picker in Claude Code doesn't show it. Here's why it happens, how to …
Hermes in Practice: Two Saves and a Sharp Edge
A few months into running Hermes day-to-day, here are the moments worth writing down — the diagnosis from a machine I'd …
OpenClaw Telegram Multi-Agent Routing: Preventing Cross-Talk
Two Telegram bots, two isolated agents, zero cross-talk — but only if you respect OpenClaw’s default-agent and …
Claude Code with LongCat-Flash-Chat: Surprising Performance
After testing Meituan's open source LongCat-Flash-Chat model, I was impressed by its generous token allocation, …
Claude Code Multi-Model Launcher: Unified AI Provider Access
A smart launcher script for Claude Code that simplifies switching between multiple AI providers including Zhipu, …