Taylor Kolasinski

Founder of Poisson Labs. ML systems & research, reinforcement learning, edge AI. Previously Matter AI, Better.com, WeWork. Brooklyn.

Writing about reinforcement learning, large-scale model training, and simulation.

Notes

View all
  • Aug 22, 2026 Price the Reward Function — I measured what a code-execution reward costs a GRPO training loop in GPU-seconds. The median service time turns out to be the wrong number to look at.
  • Aug 19, 2026 Map the Failure Boundary — I froze a Go1 locomotion policy and measured where it fails across 6,400 combinations of floor friction and lateral push, then retrained on the two gaps the map exposed and measured again.
  • May 22, 2026 Monte Found a Decorative Channel and a Reward Exploit in One of Our Benchmarks — Monte's adversarial training exposed two failure modes in a MARL benchmark: a communication channel that looked load-bearing but wasn't, and a reward function that incentivized trivial policies.

Viz

View all

Visual evidence from reinforcement learning, simulation, and large-scale model training.

Logs

View all