When a language model is instructed to suppress an internal state, the state is not erased — it is re-encoded into a basis that ordinary expression-aligned readouts miss. Targeted interventions separate decodability from output accessibility, showing how activation monitors trained in one behavioural regime can fail in another.
Publications
Papers and preprints, most recent first.
A reinforcement learning framework that refines layer decomposition models from vision-language model feedback instead of paired data. Flow-GRPO with LoRA samples candidate decompositions, and a two-stage scoring pipeline counters the tendency of VLMs to cluster their scores into a narrow band.
A lightweight bridge for video-synchronised Foley that leaves both pretrained models frozen. V-JEPA2 video embeddings enter Stable Audio Open through compact cross-attention placed after the existing text cross-attention, so prompts set global semantics while video refines timing and local dynamics.
Image conditioning is ambiguous: the same reference image can mean style transfer, object extraction, or something else again. Pairing natural-image conditioning with "Instruct" prompts lets one model switch between those interpretations, learning several tasks at once with little quality loss against dedicated single-purpose models.