Language model experiments
On this page
Research questions
Can recursive computation use a model’s parameters more effectively? Can uncertainty guide how much computation each token needs? Does a coding fine-tune improve completed tasks rather than merely shorten reasoning?
The projects
OSRT-Ostinato investigates sparse recursive transformer architectures. LANTERN explores adaptive computation based on token uncertainty. My coding-model fine-tuning repository covers supervised fine-tuning, preference optimisation and coding-agent evaluation.
How I assess the work
Architecture code, a training run and a useful checkpoint are different results. Comparisons need a named baseline, the same task conditions, complete execution traces and a task-success measure.
Reasoning cost is interesting only alongside quality. A model that uses fewer tokens but fails more tasks is not an improvement for a coding assistant.
Evidence and status
These personal research projects are in progress. The repositories contain architecture implementations, experiment setup and evaluation material for the relevant versions.