Short note

Evals are the product

14 September 20261 min read

If you can't measure it, you're demoing it. The fastest teams I've worked with treat their eval suite as the real spec.

  • Write the eval before the prompt
  • Track regressions like you track latency
  • Small, sharp datasets beat big vague ones

In practice

Choose an acceptance condition before a run, keep the baseline comparable and inspect the complete trace when a score changes. A lower reasoning cost is useful only when task quality holds up.

My model experiments use that distinction: an architecture implementation, a training run and a useful checkpoint are separate outcomes.