Making tools understandable, along with the thinking behind research.
Concept illustrationQuestions, evidence, tools and understanding begin with a clear question.
Course design and materials development
AI4Research
Course materials on human–AI collaboration and AI-assisted research, covering how to ask questions, build context, use research skills, manage agents, and organize research workflows.
Four areas of focus
01
Questions & Inquiry
Turning vague needs into questions with clear conditions, goals, and room for further inquiry.
02
Context & Evidence
Organizing literature, materials, research history, and sources that can be checked.
03
Tools & Workflows
Understanding how skills, agent management, and execution workflows support research.
04
Verification & Human Understanding
Checking outputs while explaining the methods, their limits, and why those checks matter.
The learning process I want to preserve
Research tools should help people develop their own understanding. I design courses around clear explanations, further inquiry, and feedback, connecting technical practice with the ideas behind it. This page introduces my course materials and areas of knowledge sharing.
These are prepared slides, speaker notes, and lesson plans. Dates record materials preparation.
Introductory materials and speaker notes ·
Post-training: from a methods map to a research choice
For students entering research, I organize a map of post-training: supervised fine-tuning, preference optimization, reinforcement learning with verifiable rewards, agent training, and the role of rewards and evaluation. The materials move from what the field does to what a student might investigate, bringing open questions, compute and evaluation requirements into the choice of a research direction.
Explore chapters & exercises
Where post-training sits in the training pipeline
Demonstrations, preferences and online feedback
Reasoning RL, agentic RL, rewards and verifiers
Open questions, research requirements and entry paths
60-minute lesson plan · 17 slides with speaker notes ·
Three post-training paradigms and how research questions arise
These materials compare the learning signals in SFT, offline DPO and online GRPO, then connect the feedback structure of language models to visual generation. Group-advantage calculations, a reward-scaling counterexample and a two-dimensional point-cloud demonstration distinguish overall diversity, task-relevant variation and actual training gains. The lesson ends by turning an intuition into a research question with controls, measurements and a possible failure.
Explore chapters & exercises
Training and generation: three roles for noise
SFT demonstrations, DPO preferences and GRPO group feedback
From language to visual generation and sampling-time randomness
From a paper’s claim to a falsifiable study sketch
Classroom exercise
A two-dimensional demonstration gives two outputs the same mean, total variance and response singular values, but different variation along the task direction. Switching the reward from the horizontal to the vertical axis reverses that ordering. This is a constructed teaching counterexample, rather than a measured SFT or GRPO result.
After-class tasks
Explain the three methods’ learning signals and online properties
Calculate standardized advantages for rewards [4,5,6,7] and test scaling
Swap the task direction in the 2D example and explain the reversal
State a controlled, falsifiable prediction in no more than 150 words
Materials for a roughly 55-minute introduction · 37 slides ·
World models: from predicting futures to making decisions
The introduction begins with a cup-pushing scenario, connecting state, observation, action and an imagined feedback loop. It then compares generative, representation-based and decision-oriented world models through architectures including World Models, Dreamer, MuZero and JEPA. Interaction, counterfactuals, persistence and uncertainty lead to a central question: does the simulation help a decision, beyond producing a convincing image?
Explore chapters & exercises
State, observation, action and three modelling routes
Classic architectures and video interaction
Counterfactuals, persistence and uncertainty
Testing action conditioning, decision value and efficiency
Classroom exercise
From the same state, would a different action change the prediction and the next decision? Counterfactual branches in the cup scenario help distinguish coherent imagery from a model that actually uses action information.