cs.LG updates on arXiv.org
・arXiv:2609.08368v1 Announce Type: new Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training.
・Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable.
・With accuracy, efficiency, reliability, and scalability as first-class goals, Miles a