2026-09-18
Leaderboards
Your Design Picks Your Estimator — Two kinds of model leaderboard need almost entirely different statistics, and which one you are in is decided by your design rather than your taste.
2026-09-17
Leaderboards
How Do You Rank an AI? — Three years of the Arena leaderboard's changelog, read as a history of how AI got measured.
2026-07-26
Life
A Game That Cannot Be Gamed — Finishing the San Francisco First Half Marathon, six months after a 10K got me hooked.
2026-06-01
Evaluation
Beyond Error Bars — a nine-part series on the statistics of LLM evaluation.
- Part 1Every Eval Era Created the Next Measurement Problem
- Part 2An Eval Is a Decision Instrument
- Part 3Error Bars, Applied
- Part 4How Many Examples — and When Can You Stop?
- Part 5Humans Are Instruments — Rater Ops Is the Eval
- Part 6Your LLM Judge Is a Biased, Noisy Instrument. Debias It.
- Part 7Rankings Under Uncertainty: What Arena Numbers Mean
- Part 8Agent Evals: the Statistics of pass@k and pass^k
- Part 9Building a Product Eval, Instrumented
2026-01-01
Career
What Five Years at a Startup Taught Me — Sixteen things I believe now that I didn't believe as a new data scientist.