Aug 22, 2026Price the Reward Function— I measured what a code-execution reward costs a GRPO training loop in GPU-seconds. The median service time turns out to be the wrong number to look at.
Aug 19, 2026Map the Failure Boundary— I froze a Go1 locomotion policy and measured where it fails across 6,400 combinations of floor friction and lateral push, then retrained on the two gaps the map exposed and measured again.
Jan 16, 202610,924x: The Instability Bomb at 1.7B Scale— Part 2 of the mHC reproduction series. I scaled from 10M to 1.7B parameters and watched Hyper-Connections hit 10,924x signal amplification. That's not a typo.
Jan 11, 2026DeepSeek's mHC: When Residual Connections Explode— DeepSeek replaced standard residuals with massive Hyper-Connections, and then watched them explode. I reproduced the 9x signal amplification to understand their fix: the Sinkhorn-Knopp algorithm. Part 1 of 2.