TRL is a cutting-edge library designed for post-training foundation models using advanced techniques like Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Built on top of the Transformers ecosystem, TRL supports a variety of model architectures and modalities, and can be scaled-up across various hardware setups.

Categories

articleHugging Face3 Sept 2026

Training a coding model to paint watercolours with TRL and OpenEnv

An open reproduction of the viral watercolour-painting model: GRPO in TRL over an OpenEnv environment, with the reward split between HPSv3 and a pairwise judge scored against 178 hand-rated paintings — RL over taste.

Sergio Paniego
huggingface.co

More in Post-Training

View all tools
Miles logo

Miles

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Post-Training
Molt logo

Molt

An agentic-first RL framework for research. Ray · vLLM · NVIDIA AutoModel — the smallest PyTorch-native stack for 1T-class fully-async, multimodal, multi-turn agentic RL.

Post-Training
OpenEnv logo

OpenEnv

An interface library for RL post training with environments.

Post-Training

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.