EluteDiff
On this page
The problem
A retention-time estimate helps identify compounds in chromatography, but a single number hides uncertainty. I’m investigating whether a model can predict a distribution over the time axis instead.
What I’m exploring
EluteDiff conditions a discrete diffusion model on a molecule to generate an HPLC retention-time density. The project uses METLIN SMRT data and converts scalar retention-time labels into Gaussian density targets on a quantised time axis.
Model and evaluation
The training approach uses DiffusionGemma with Unsloth and LoRA. Target construction, tokenisation, parsing and metrics are implemented; model training and comparisons remain research work.
The question is whether a generated density can support useful uncertainty estimates, tolerance windows and candidate ranking. It connects my analytical chemistry background with generative modelling.
Turning a time into a modelling target
The data provides a scalar retention-time label for each molecule. EluteDiff constructs a Gaussian density around that label on a fixed, quantised time axis. Fixed-width tokens encode the density for discrete text diffusion, and a strict parser checks whether sampled output can be recovered as a valid numerical target.
This target construction is an experimental modelling choice. A Gaussian created from one measured time is not itself a measured uncertainty distribution. Whether the model learns calibrated uncertainty needs separate evaluation.
Why generate the whole density?
A density can allocate probability over the chromatographic canvas rather than only returning a centre. That representation could support a tolerance-window score or rank candidate molecules against an observed time. Bidirectional denoising may suit a fixed-axis output whose bins need to remain consistent with one another.
The interesting question is whether that representation adds useful information beyond a scalar predictor or a centre-bin control. Strong graph and molecular-fingerprint regressors remain relevant comparison points for point prediction.
Evaluation questions
- Does the parser accept the output, and does it represent a valid density?
- How accurate is the centre compared with the same baseline and data split?
- Does probability within a chosen tolerance window match observed outcomes?
- Does the density improve candidate ranking under the same conditions?
- What changes when the target width, quantisation or sampling procedure changes?
What is implemented
The CPU core includes target construction, tokenisation, parsing and evaluation metric families. METLIN/RDKit featurisation and classical/GNN baselines are typed scaffolds. The diffusion training and sampling path uses the reference Unsloth approach with LoRA and needs the training dependencies and suitable GPU resources.
The repository includes notebooks that begin with synthetic demo data, plus a density-first proposal explaining the controls and ablations. These make the experiment inspectable; they do not establish a completed accuracy result.