Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks,
such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. arxiv.org/pdf/2607.07508
Professor @ Tsinghua, Founder of Z.ai.
AGI, LLM.
“The value of a man should be seen in what he gives and not in what he is able to receive.”―Einstein





