articleHugging Face3 Sept 2026

Training a coding model to paint watercolours with TRL and OpenEnv

An open reproduction of the viral watercolour-painting model: GRPO in TRL over an OpenEnv environment, with the reward split between HPSv3 and a pairwise judge scored against 178 hand-rated paintings — RL over taste.

Sergio Paniego
huggingface.co

More in Post-Training

View all tools
Miles logo

Miles

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Post-Training
Molt logo

Molt

An agentic-first RL framework for research. Ray · vLLM · NVIDIA AutoModel — the smallest PyTorch-native stack for 1T-class fully-async, multimodal, multi-turn agentic RL.

Post-Training
PorTAL logo

PorTAL

PorTAL generates portable task specific LoRA adapters that can efficiently transfer across language models.

Post-Training

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.