Writing
Wind Spirit: balance is measured, not argued
Before the UI, before the LLM chiefs, a headless harness ran a civilization sandbox for hundreds of years across many seeds. Four ways the villages collapsed, the fixes, and what the numbers said about the game I thought I was designing.
Wind Spirit: how much intelligence does a stone-age chief need?
Thirty-four models, 96 sampled game situations and about $22 of API credit: where an AI village chief needs a good model, where a cheap one is just as good, and how format failures looked like stupidity in the first pass.
Grammar Garden: evolving plants from tiny recipes, thirty years later
A Saturday-night return to L-systems: a browser game where plants grow from a few rewrite rules, fall over if they can't hold themselves up, and get bred by bees and butterflies under two different climates. I wrote no code.
What time is "last week"? Tools for AI agents to stop fumbling dates
A benchmark for how well LLM agents handle date/time in the tool calls they make and the answers they give — and what happened across seven models when I gave them a deterministic time library instead.