Xiaomi's MiMo team states that MiMo-V2.6 may represent one of the largest single reinforcement learning training runs ever conducted by an open-source model team. Despite compute constraints, the team dedicated dozens of personnel over a sustained period to advance RL scaling, utilizing MixRL training for verifiable tasks like code and merging capabilities via MOPD for complex or hard-to-verify scenarios. To support agent reinforcement learning research, the team released a Qwen model distilled from MiMo RL trajectories, 7,000 diverse environments, and a complete RL training framework. Xiaomi described MiMo-V2.6 as a starting point and committed to continued investment in time, compute, and R&D resources to develop sustainable self-improving intelligent systems.