Left: Performance across different training stages. Right: Performance compared to closed frontier models (Claude Sonnet 4.6 and Gemini Flash family). R = task reward, EM = exact match. Breakdown shows the reward components: abstention, company and time score.