Research note
What makes a useful coding model?
On this page
Start with a completed task
The outcome I care about is whether the model solves the task under the same conditions as a baseline. A shorter response or a lower token count is not enough.
Keep the trace
Tool calls, edits, failures and final verification explain why a run succeeded or failed. Aggregate scores are easier to trust when the underlying runs remain inspectable.
Compare the right things
Hold the task, tool environment and acceptance criteria constant. Record model and checkpoint versions. Check supervised fine-tuning before assuming preference optimisation adds value.
Explore the projects and the evaluation note.