Research

I work on generative models across modalities. Recent work covers conditioning and control in diffusion models (IPAdapter-Instruct), aligning frozen text-to-audio models to video for Foley (Foley Control), and fine-tuning image layer decomposition with reinforcement learning from vision-language model feedback (Stable-Layers).

More recently I have been looking at interpretability: what happens to an internal state when a language model is told to suppress it, and why activation monitors trained in one behavioural regime can quietly stop working in another (J-Space Hacking).

Research interests

  • Generative models
  • Diffusion & flow matching
  • Controllable image generation
  • Video-to-audio
  • Reinforcement learning from model feedback
  • Mechanistic interpretability

Elsewhere

Code and experiments live on GitHub. Papers are collected on the publications page.