Research note

What makes a useful coding model?

On this page

Start with a completed task

The outcome I care about is whether the model solves the task under the same conditions as a baseline. A shorter response or a lower token count is not enough.

Keep the trace

Tool calls, edits, failures and final verification explain why a run succeeded or failed. Aggregate scores are easier to trust when the underlying runs remain inspectable.

Compare the right things

Hold the task, tool environment and acceptance criteria constant. Record model and checkpoint versions. Check supervised fine-tuning before assuming preference optimisation adds value.

Explore the projects and the evaluation note.