Language model research

Ostinato

Research in progress
On this page

The question

Can a language model get more useful computation from its parameters by revisiting the same decoder blocks?

What I’m exploring

Ostinato is my OSRT, or Optimized Sparse Recursive Transformer, project. It combines depth recurrence, sparse mixture-of-experts and Muon optimisation. Physical decoder blocks are reused across loops, increasing effective depth without adding a separate set of parameters for every step.

Token context passes through reused decoder blocks to predict the next token

Research direction

The work brings together architecture experiments, training and evaluation. Repeated computation is a design choice to investigate; its value depends on the quality and cost of the resulting model.

Evidence and status

This is ongoing research. The repository records the implementation and evolving experiments.

Parameters and computation are separate budgets

A conventional decoder gives each layer its own weights. Recurrence applies the same physical blocks repeatedly, allowing the hidden representation to develop over more steps. The parameter budget can stay smaller, while the amount of computation still grows with the number of loops.

Sparse mixture-of-experts adds another choice: routing a token to a subset of feed-forward experts. Total parameters, active parameters, training cost and decoding latency therefore describe different aspects of the model. A useful comparison needs to say which budget is held constant.

The engineering questions

Repeated blocks need a way to distinguish one pass from the next. Attention and caching must remain consistent with the recurrence schedule. Expert routing must distribute useful work without turning theoretical sparsity into expensive overhead. The optimiser and data pipeline must support a stable training run before architecture quality can be assessed.

The project uses Muon for selected matrix parameters alongside AdamW for other parameter groups. The exact shape and settings belong to a named experiment configuration, rather than a permanently fixed claim on this page.

What would establish an improvement?

I want to compare quality at an explicit computation and latency budget. A smaller parameter count alone does not establish faster decoding, and a completed training run does not establish better task performance. Fixed-depth controls help separate the value of recurrence from simply spending more compute.

Parameter counts should come from the instantiated model. The repository includes a budget script and a runbook so architecture settings, training setup and observed results can be read together. Check the current experiment notes for version-specific outcomes; this portfolio describes the research direction without presenting an unverified checkpoint as a result.

LANTERN and coding-model experiments explore adaptive computation and post-training. The coding-model evaluation note describes the task-level evidence I look for when deciding whether a model is useful.