Machine Learning and AI Alignment Researcher
Occasional notes from experiments, debugging, and ideas around reinforcement learning, alignment, multi-turn agents, and human–AI interaction.
August 2026 · PPO · multi-turn RL · value functions
A weekend debugging note on turn-level credit assignment, critic representations, and what finally made a multi-turn PPO baseline learn.