JustRL II (critic-based RL)
training_methodsJustRL II is OpenBMB's critic-based reinforcement-learning algorithm, used in the RL stage of MiniCPM5-2B post-training and described in "JustRL II: Scaling Small LLMs to 128K Reasoning with a Critic". Instead of critic-free policy-gradient methods (GRPO-family), it trains with a learned critic that provides value estimates, which substantially improves training stability and enables long-context reasoning at 128K for small (2B-class) models. In the MiniCPM5-2B pipeline the RL teachers for math, code, agentic tasks and writing are trained with JustRL II before being merged back into the release model via On-Policy Distillation; combined RL+OPD lifted reasoning/general benchmarks by an average +10.96 points and agentic capabilities by +6.96 points.