Training a coding model to paint watercolours with TRL and OpenEnv
An open reproduction of the viral watercolour-painting model: GRPO in TRL over an OpenEnv environment, with the reward split between HPSv3 and a pairwise judge scored against 178 hand-rated paintings — RL over taste.
