Miles
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
8 tools in Post-Training.
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
An agentic-first RL framework for research. Ray · vLLM · NVIDIA AutoModel — the smallest PyTorch-native stack for 1T-class fully-async, multimodal, multi-turn agentic RL.
An interface library for RL post training with environments.
PorTAL generates portable task specific LoRA adapters that can efficiently transfer across language models.
Continual learning infra for self-improving agents
Tinker is a training API for researchers and developers.
TRL is a cutting-edge library designed for post-training foundation models using advanced techniques like Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Built on top of the Transformers ecosystem, TRL supports a variety of model architectures and modalities, and can be scaled-up across various hardware setups.
Build continually improving models on your agent traces by distilling frontier open models
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.