A young man with short black hair, wearing round glasses, a white dress shirt, and a blue suit jacket, poses against a plain grey background.

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation 

Date and time: Tuesday 25 August 2026, 15:15-16:15 CEST
Speaker: Se-Young Yun, the Kim Jaechul Graduate School of AI at KAIST
Title: Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Where: Digital Futures hub, Osquars Backe 5, floor 2 at KTH main campus OR Zoom
Directionshttps://www.digitalfutures.kth.se/contact/how-to-get-here/
OR
Zoomhttps://kth-se.zoom.us/j/69560887455

Host: Alexandre Proutiere alepro@kth.se

A young man with short black hair, wearing round glasses, a white dress shirt, and a blue suit jacket, poses against a plain grey background.

Bio: Se-Young Yun is an Associate Professor in the Kim Jaechul Graduate School of AI at KAIST, where he leads the Optimization and Statistical Inference (OSI) Lab. He received his B.S. and Ph.D. degrees in Electrical Engineering from KAIST. Prior to joining the KAIST faculty, he held research and postdoctoral positions at KAIST, KTH Royal Institute of Technology in Sweden, the MSR–INRIA Joint Research Center in Paris, Microsoft Research Cambridge, and Los Alamos National Laboratory in the United States. 

His research spans machine learning theory and modern foundation models, with current interests in large language model reasoning and alignment, reinforcement learning, efficient foundation-model training and inference, generative and multimodal AI, and statistical learning and optimization.

Yun’s recent work investigates topics including reinforcement learning and feedback-based alignment for LLMs, reasoning and evaluation, multi-agent learning, efficient language-model architectures and inference, diffusion and generative models, multimodal learning, and foundational problems in statistical learning and reinforcement learning.

Abstract: Scaling remains the dominant recipe for improving large language models, yet training each larger model from scratch discards the capability already accumulated in existing checkpoints. Reusing a trained model is by far the most economical path to a stronger one—but the models available to reuse are typically smaller than the model being trained, inverting the usual distillation setting: the student must surpass its teacher rather than merely imitate it.

Weak-to-strong knowledge distillation supplies dense, token-level supervision that makes early learning fast and stable, but it is bounded by the teacher’s own competence, and in the limit the student inherits the teacher’s ceiling. Reinforcement learning with verifiable rewards, such as GRPO, imposes no such ceiling, but its supervision is sparse and sequence-level, making it sample-inefficient precisely where the student is still weak. We argue that these two signals are complementary rather than competing, and present a training framework that unifies on-policy distillation with GRPO in a single loop.

Events & seminars