Short note
Evals are the product
If you can't measure it, you're demoing it. The fastest teams I've worked with treat their eval suite as the real spec.
- Write the eval before the prompt
- Track regressions like you track latency
- Small, sharp datasets beat big vague ones
In practice
Choose an acceptance condition before a run, keep the baseline comparable and inspect the complete trace when a score changes. A lower reasoning cost is useful only when task quality holds up.
My model experiments use that distinction: an architecture implementation, a training run and a useful checkpoint are separate outcomes.