Skip to content

Blog

Long-form notes on the papers and systems I am trying to keep up with.

The Variance of Importance Sampling

Post-training a language model involves an awkward fact: the samples you have were usually not drawn from the distribution you care about. The rollouts came from a policy that is a few gradient steps stale, or from a replay buffer, or from an inference engine whose numerics do not exactly match the trainer's. Call the distribution that actually produced the tokens \(q\), and the policy you want to reason about \(p\).

Importance sampling is the textbook repair, and it has one very attractive property: it is unbiased. But unbiasedness is the cheapest property in statistics. An estimator can be unbiased and still be wrong by a factor of three every single time you run it — and, worse, report a tight confidence interval while doing so.

This post works through two discrete examples small enough to compute by hand, and deliberately shaped like a language model, to show how fast this goes wrong. All numbers below come from simulations you can re-run in a few lines of Python.

Policy Gradients and the Arithmetic of Variance

The policy gradient theorem is about four lines of algebra, and almost none of the work in making it function is in those four lines. The derivation gives you an estimator that is unbiased for any policy, any reward, any dynamics — including dynamics you cannot differentiate and rewards you cannot write down. Then you run it and nothing happens, because the estimator's variance is large enough to bury the signal it is carrying.

This post derives the gradient in Pieter Abbeel's notation, then treats variance as the actual subject. The claim I want to make precise is the one usually waved at: subtracting a baseline reduces variance. That is a theorem with an explicit form, an exactly optimal \(b^\star\), and a characterisation of which baselines help and which hurt — and the standard choice, the mean return, is not the optimal one.

Every number below comes from a simulation or an exact enumeration, and the closed forms are checked against Monte Carlo before they are stated.

Score Centering Is a Straight-Through Estimator

Marek and Ryabinin (2026) introduce score centering, an additive correction that stabilises RL on language models when the inference engine that produced the rollouts does not exactly match the trainer that computes the gradients. Their derivation is a covariance identity. Writing \(\bar s\) for the expected score under the sampler, the expected update at a prefix splits as

\[\underbrace{\mathbb{E}_q\!\left[R\, s_{y_t}\right]}_{\text{policy gradient}} = \underbrace{\mathbb{E}_q[R]\, \bar s}_{\text{drift}} + \underbrace{\mathrm{Cov}_q\!\left(R, s_{y_t}\right)}_{\text{signal}},\]

and the drift term — which knows nothing about which rollouts succeeded — is the negative gradient of a cross-entropy loss with the sampler as teacher. Vanilla policy gradient under mismatch is quietly distilling the trainer toward a biased copy of itself, and since that copy is resynced from the trainer every step, the error compounds instead of converging. Score centering subtracts \(\bar s\) and the drift cancels exactly.

I have nothing to correct in that derivation. What I want to add is a name for the object, because the name comes with a literature. The estimator that drift afflicts is a straight-through estimator, and the setting it lives in is quantization-aware training. The paper never uses either term, and it never specialises its scores to a softmax. Doing both is worth the trouble, for three reasons:

  • it turns \(\bar s\) from "a nonzero expected score" into a vector you can write down — exactly \(q - p\), the sampler minus the trainer;
  • it turns score centering into a statement about what the corrected estimator is: \(A(e_y - q)\), the sampler's own score, evaluated on the trainer's logits;
  • it reframes the paper's stated limitation — that the corrected update measures the covariance under the sampler rather than the trainer — as a question about which gradient you wanted, which turns out to predict the one place their experiments split.

I write \(A\) for the advantage and \(\mu\) for \(\bar s\). Everything below is checked numerically: the closed forms are verified against the released implementation to machine precision, and the simulations compute exact expectations over a vocabulary rather than Monte Carlo estimates, except where noted.

Video Generation Models Explosion 2024

Video generation models exploded onto the scene in 2024, sparked by the release of Sora from OpenAI. This blog post is my way of keeping track of the progress of this fascinating field. I will review all the key techniques that are used in building state-of-the-art video generation models (1).

    • A comprehensive review of all text-to-image/text-to-video models is beyond the scope of this blog post. I will focus on research that has been published, productionized, or open-sourced.
    • All of the videos and images are reproduced from the cited projects and papers, and the copyright belongs to the authors or the organization that published their papers. Below I adapted key figures for each paper under the fair use clause of copyright law.