What the paper is about
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.
What it covers
Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts Shuyue Stella Li Xiaochuang Han Yulia Tsvetkov Luke Zettlemoyer Affiliation: University of Washington Affiliation: [email protected] Affiliation: https://github.com/stellalisy/VVRBench https://huggingface.co/datasets/stellalisy/VVRBench Abstract Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench , with 10,000 tasks over 32 constraint types, and VVRBench -Challenge, with 720 more complex tasks; the strongest model we evaluate—GPT-Image-2.5—solves 21.4% of VVRBench -Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes. 1 Introduction Reinforcement learning with verifiable rewards has improved how precisely language models follow instructions: constraints such as use the word X at least three times are checked by code and used directly as rewards ( Zhou et al., 2023 ; Lambert et al., 2025 ; Pyatkin et al., 2025 ) , but image generation has no equivalent reward. Text-to-image generators often fail to follow instructions precisely: given three red circles to the left of two blue squares , they draw the wrong counts, colors, or positions, and fail more often as a prompt combines more requirements ( Ghosh et al., 2023 ; Huang et al., 2023 ; Kamath et al., 2025 ) . Post-training rewards for instruction following come from learned evaluators: preference models ( Kirstain et al., 2023 ; Xu et al., 2023 ) , vision-language models (VLMs) that answer questions about the image ( Hu et al., 2023 ; Cho et al., 2024a ; Lin et al., 2024 ) , and object detectors ( Ghosh et al., 2023 ) . These evaluators make errors on the judgments that instruction following depends on ( Saxon et al., 2024 ; Wiles et al., 2025 ; Kajić et al., 2024 ; Chen et al., 2025b ; Kamath et al., 2025 ) , and policies trained on them exploit these errors ( Zhang et al., 2024 ; Kim et al., 2024 ; Hong et al., 2026 ) . We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, in which the generated image is scored deterministically by verifiers : Python functions over pixels, with no learned detector, OCR system, embedding model, or VLM. VVR covers instructions with clear, objective requirements combining color, count, shape, and spatial relations (Figure 1 ). Unlike constraints in text instruction following, which govern mostly separate properties of the output (e.g., length, keyword, format) and can be excluded pair by pair ( Pyatkin et al., 2025 ) , visual constraints lead to more complicated compatibility conflicts. For example, in “ A contains B, B contains C, and C contains A ,” every pair of constraints can be satisfiable, but the three together are not. Therefore, we propose a generator that guarantees constraint satisfiability under compositions, and compose natural-language instructions from the valid constraint sets. With our generator and constraint taxonomy, new VVR tasks can be generated in any number and at any chosen complexity, for evaluation or for training. Program verifiers of each VVR constraint also enable fine-grained diagnosis of generator capability over different types of instructions. Figure 1: Example VVRBench tasks from different complexity ranges ( C 1 C_{1} – C 5 C_{5} ) and Challenge ( C ∗ C^{} ). Each panel shows the prompt, its complexity, the number of object instances, the number of constraints, the active constraint families, and a reference image that satisfies every constraint. VVRBench . We instantiate VVR with colored geometric shapes and 46 constraint types over counts, attributes, and spatial relations and release VVRBench , with 10,000 tasks across five complexity ranges, and VVRBench -Challenge , with 720 more complex tasks to discriminate among frontier models (§ 3 ). The strongest open-weight model, FLUX.2-dev, achieves 19.2% accuracy on VVRBench , and the strongest model overall, GPT-Image-2.5-Sunburst, solves 21.4% of VVRBench -Challenge. Failures concentrate in the Cardinality (e.g., “twice as many A as B” ) and Topology (e.g., “each A is inside a different B” ) constraint families. The benchmark can be updated with higher complexity as frontier models evolve. RLVVR. We use the verifiable VVR scores as rewards for reinforcement learning (RLVVR) to train image generators to follow instructions precisely. To study how training complexity affects generalization, we procedurally generate two training corpora: VVR-Easy contains only low complexity tasks of at most one constraint family, and VVR-Matched matches the VVRBench distribution. We show that 1) training on easy distribution generalizes to harder tasks, 2) training on harder tasks teaches compositionality, 3) training on colored shapes transfer to out-of-domain natural prompts to improve position and counting, and, most importantly, 4) mixing VVR with existing post-training objectives (e.g., GenEval2, OCR, PickScore) improves general benchmark performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes. Our contributions are: 1. VVR, the first framework for programmatically verifiable image rewards, whose generator composes constraints on color, count, shape, and spatial relations into satisfiable instructions at any chosen complexity, each constraint checked by its own verifier (§ 2 ). 2. VVRBench and VVRBench -Challenge, benchmark that reveals capability gap of image generators to follow instructions, on which even frontier image generators fail most complex tasks, with failures tracable to specific constraint types (§ 3 ). 3. RLVVR, reinforcement learning with VVR rewards, which improves precise instruction following on tasks harder than those seen in training, transfers from synthetic scenes to natural prompts, and, mixed with existing post-training objectives, improves general benchmark performance and human preference (§ 4 ). 2 Verifiable Visual Rewards 2.1 VVR Task Representation A VVR task specifies the requirements of a scene of colored shapes: s = ( 𝒢 , ℬ , 𝒜 , ℱ , p ) . s=(\mathcal{G},\mathcal{B},\mathcal{A},\mathcal{F},p). (1) Here 𝒢 \mathcal{G} is a set of object groups, ℬ \mathcal{B} is a set of background constraints, 𝒜 \mathcal{A} is a set of active constraints on the object groups, ℱ \mathcal{F} is a set of forbidden-content constraints, and p p is the natural-language instruction. In VVRBench , ℬ \mathcal{B} and ℱ \mathcal{F} are the same in every task: a plain background color and ℱ = { \mathcal{F}={ no_unrequested_objects } } , so we focus the rest of the section on the constraints in 𝒜 \mathcal{A} . Each constraint in 𝒜 \mathcal{A} is instantiated with a constraint type from the constraint library and one or more object groups. Each constraint type has a predefined number of object groups that it operates on, and a set of supported values (Table 1 ). For example, exact_count(g1; 3) is a unary constraint that requires group g1 to contain three objects, and left_of(g1,g2) requires group g1 to appear left of group g2 . Every group has exactly one color constraint and one shape constraint; all other constraints are optional. Appendix A.2 shows a complete task. A valid task must satisfy three desiderata: 1) Well-formed: every constraint uses a defined type with supported parameter values, such as one of eight colors or three shapes, and refers to as many object groups in 𝒢 \mathcal{G} as its type requires; 2) Jointly satisfiable: at least one placement and sizing of the specified objects satisfies all constraints in ℬ \mathcal{B} , 𝒜 \mathcal{A} , and ℱ \mathcal{F} simultaneously; and 3) Faithfully expressed: p p states every constraint in ℬ \mathcal{B} , 𝒜 \mathcal{A} , and ℱ \mathcal{F} without any addition or omission. Table 1: Constraint library. VVR groups 46 constraint types into five families. Appendix A.1 lists all types and their supported values, and Appendix C.2 defines their verifiers. Family Visual property Example constraint types Grounding Object identity and attribute binding Color; shape; color–shape binding Cardinality Quantities and count comparisons Exact count; equal, greater, or fewer counts; count ratios (X times as many) Spatial Position and arrangement Image regions; relative order; alignment; grids; distance comparisons Size Relative visual extent Pairwise and groupwise size; within-group variation; extrema Topology Contact and enclosure Touching; separation; containment; distinct containment 2.2 Task generation We now walk through the stages of the VVR generator that produces well-formed, jointly satisfiable, and faithfully expressed tasks. Appendix B provides an example generation and validation details. 1. It first creates a scene by sampling background constraints ℬ \mathcal{B} and objects. It randomly assigns every object a color, shape, position, and size, forming object groups 𝒢 \mathcal{G} , and samples forbidden-content constraints ℱ \mathcal{F} that no object in the scene violates. 2. For each constraint type in the library, it lists all object group tuples with size corresponding to the type’s arity. The constraint type and its input tuple form an instantiated constraint . 3. All instantiated constraints that are true under the constructed scene, checked by program verifiers, form satisfiable constraint set 𝒜 ∗ \mathcal{A}^{} . From 𝒜 ∗ \mathcal{A}^{} , multiple valid active constraint sets 𝒜 \mathcal{A} can be sampled such that 𝒜 ⊆ 𝒜 ∗ \mathcal{A}\subseteq\mathcal{A}^{} . 4. A template τ \tau with phrasing variants transforms each task requirement into natural language, p = τ ( 𝒢 , ℬ , 𝒜 , ℱ ) p=\tau(\mathcal{G},\mathcal{B},\mathcal{A},\mathcal{F}) , forming s = ( 𝒢 , ℬ , 𝒜 , ℱ , p ) s=(\mathcal{G},\mathcal{B},\mathcal{A},\mathcal{F},p) . Every constraint instantiates a library type on a tuple of sampled groups whose length equals the type’s arity, so every task is well-formed by construction. The scene satisfies ℬ \mathcal{B} , ℱ \mathcal{F} , and every constraint in 𝒜 ∗ \mathcal{A}^{} , so it satisfies the task formed with any 𝒜 ⊆ 𝒜 ∗ \mathcal{A}\subseteq\mathcal{A}^{} , making the task jointly satisfiable . Finally, in the template τ \tau , every requirement of s s has a fixed phrase in p p , and every phrase in p p comes from a requirement of s s , guaranteeing expression faithfulness . Structural complexity estimates task difficulty. We define the structural complexity of a task as C ( s ) = ∑ a ∈ 𝒜 c ( a , s ) C(s)=\sum_{a\in\mathcal{A}}c(a;s) , where c ( a , s ) c(a;s) is the complexity contribution of one constraint a a in task s s . c ( a , s ) c(a;s) follows a fixed rule for each constraint type and grows with the number of object instances that a a evaluates in s s . Appendix A.1 gives the complexity contribution of every constraint type. In the specific instantiation of tasks that produces datasets in Table 2 , each task has one background-color constraint and one forbidden-content constraint, so the constraints in ℬ \mathcal{B} and ℱ \mathcal{F} are excluded. Datasets. Given a target distribution 𝒯 \mathcal{T} over constraint families and structural complexity range, the generator can retain task candidates to fit 𝒯 \mathcal{T} . Thus datasets can be built to evaluate or learn specific constraint types at specified difficulty. VVR datasets used by this paper and their complexity distribution are listed in Table 2 . 2.3 Deterministic, reference-free constraint verification VVR is an open-ended image generation task, where any image that satisfies all constraints receives full credit, so the verifier has to be reference-free. It takes in the generated RGB image x x and the formal constraints ( 𝒢 , ℬ , 𝒜 , ℱ ) (\mathcal{G},\mathcal{B},\mathcal{A},\mathcal{F}) of the task and produces a correctness decision. Object extraction from pixels. VVR first extracts candidate objects from the generated image using deterministic pixel-level operations. 1) It produces a binary mask for each supported color, with fixed hue and contrast thresholds. 2) Connected-component analysis assigns the same label to foreground pixels connected by a path of edge- or corner-adjacent pixels; each labeled region is a candidate object. 3) Fixed contour measurements classify each candidate into one of the supported shapes (circle, square, or triangle) based on its aspect ratio, bounding-box coverage, and convexity. 4) Candidates are then matched to object groups by the color and shape constraints of each group. Appendix C.1 visualizes the extraction pipeline, including the color map and shape classifier, as well as the handling of ambiguous colors, irregular contours, fragmented objects, and blurred boundaries. Constraint verifier library. The verifier library 𝒱 \mathcal{V} contains one program verifier for each constraint type: a Python function that applies the type’s requirement to the extracted objects. Verifier decisions use fixed comparisons of object counts, positions, extents, or boundary distances. For a constraint a ∈ ℬ ∪ 𝒜 ∪ ℱ a\in\mathcal{B}\cup\mathcal{A}\cup\mathcal{F} , the verifier v a v_{a} returns a pass-or-fail decision d a ( x , s ) ∈ { 0 , 1 } d_{a}(x,s)\in{0,1} and a partial-credit score q a ( x , s ) ∈ [ 0 , 1 ] q_{a}(x,s)\in[0,1] . ⬇ def left_of ( g1 , g2 , m ): g1x = np . mean ([ c . centroid [0] for c in g1 ]) g2x = np . mean ([ c . centroid [0] for c in g2 ]) delta = g2x - g1x partial = np . clip ( delta / max ( m , 1), 0, 1) return delta >= m , partial Consider the constraint a = left_of(g1,g2) a=\texttt{left_of(g1,g2)} , its program verifier computes δ \delta , the mean horizontal centroid coordinate of g2 minus that of g1 , and compares it with a separation margin m m . It returns two values: 1) the decision d a d_{a} , which passes when g1 lies to the left of g2 by at least the margin; and 2) the partial-credit score q a = min ( 1 , max ( 0 , δ / m ) ) q_{a}=\min(1,\max(0,\delta/m)) , the fraction of the required separation that the image achieves. The score is 0 when g1 is at or to the right of g2 , rises linearly as g1 moves left, and reaches 1 at the margin, where the decision also passes. Appendix C.2 gives the verifier code for every constraint in an example task, Appendix C.3 provides verifier validation details. Scores. A generated image succeeds only if every constraint in ℬ \mathcal{B} , 𝒜 \mathcal{A} , and ℱ \mathcal{F} passes: r exact ( x , s ) = ∏ a ∈ ℬ ∪ 𝒜 ∪ ℱ d a ( x , s ) . r_{\mathrm{exact}}(x,s)=\prod_{a\in\mathcal{B}\cup\mathcal{A}\cup\mathcal{F}}d_{a}(x,s). (2) VVRBench accuracy is the mean of r exact r_{\mathrm{exact}} across tasks. For training, we design a dense reward r dense ∈ [ 0 , 1 ] r_{\mathrm{dense}}\in[0,1] that gives partial credit through the verifier partial-credit scores: r dense ( x , s ) = ψ ( x , s ) ∑ a ∈ ℬ ∪ 𝒜 ∪ ℱ w a q a ( x , s ) . r_{\mathrm{dense}}(x,s)=\psi(x,s)\sum_{a\in\mathcal{B}\cup\mathcal{A}\cup\mathcal{F}}w_{a},q_{a}(x,s). (3) The weights w a w_{a} are fixed by constraint type, and ψ ( x , s ) ∈ [ 0 , 1 ] \psi(x,s)\in[0,1] is a multiplicative penalty factor that prevents the model from exploiting any single easy-to-learn constraint while ignoring others ( Zhang et al., 2024 ; Hong et al., 2026 ) . Appendix C.4 provides more details on w a w_{a} and ψ \psi . 3 VVRBench Table 2: Training and evaluation datasets produced by the VVR task generator. Appendix B.5 provides details on target distribution. Dataset Use Size Complexity VVRBench evaluation 10,000 3–48 VVRBench -Fast evaluation 820 3–44, 20 each VVRBench -Challenge evaluation 720 45–80, 20 each VVR-Easy training 100,000 ≤ \leq 20 VVR-Matched training 100,000 ∼ \sim VVRBench Benchmark splits. We evaluate on three benchmark splits (Table 2 ). 1) VVRBench , the main benchmark, contains 10,000 tasks of complexity 3 to 48, which we report in five ranges C 1 C_{1} to C 5 C_{5} of about 2,000 tasks each. 2) VVRBench -Fast is an 820-task subset covering the same range with 20 tasks at each integer complexity for evaluating image APIs at a twelfth of the generation cost. 3) VVRBench -Challenge contains 720 tasks of complexity 45 to 80 and adds 14 more difficult, group level constraint types, above the VVRBench range, to separate the strongest generators. Models. We evaluate ten open-weight models: FLUX.2-dev ( Black Forest Labs, 2025 ) , HunyuanImage-2.1 ( Tencent Hunyuan Team, 2025 ) , Qwen-Image-2512 ( Wu et al., 2025 ; Qwen Team, 2025 ) , HiDream-I1-Full ( Cai et al., 2025 ) , FLUX.1-dev and FLUX.1-schnell ( Black Forest Labs, 2024 ) , Stable Diffusion 3.5 Medium and Large ( Esser et al., 2024 ; Stability AI, 2024 ) , SDXL ( Podell et al., 2024 ) , and Sana 1.6B ( Xie et al., 2025 ) , and seven API based models: GPT-Image-2.5-Sunburst ( OpenAI, 2026b ; OpenAI, 2026a ) , GPT-Image-2 ( OpenAI, 2026c ) , GPT-Image-1-mini ( OpenAI, 2025 ) , Gemini-3.1-Flash-Image ( Google, 2026 ) , Gemini-3.1-Flash-Lite-Image ( Google DeepMind, 2026 ) , Gemini-3-Pro-Image ( Google DeepMind, 2025 ) , and Gemini-2.5-Flash-Image ( Google, 2025 ) . Appendix D.1 gives the additional evaluation details. 3.1 Precise instruction following is far from solved Table 3: Accuracy (%) on the 10,000 VVRBench tasks, overall and by complexity range. C 1 C_{1} to C 5 C_{5} split the tasks by structural complexity C ( s ) C(s) into five ranges of about 2,000 tasks each: 3–16, 16–21, 21–26, 26–31, and 31–48. Accuracy generally falls with complexity. GPT-Image-2 drops from 97.65% in C 1 C_{1} to 65.72% in C 5 C_{5} , and no open weight model exceeds 20% overall. Model Accuracy (%) ↑ \uparrow C 1 C_{1} C 2 C_{2} C 3 C_{3} C 4 C_{4} C 5 C_{5} GPT-Image-2 86.86 ±0.68 97.65 ±0.74 98.18 ±0.70 91.00 ±1.34 81.83 ±1.75 65.72 ±2.10 GPT-Image-1-mini 26.40 ±0.87 74.06 ±1.93 36.78 ±2.18 12.47 ±1.53 4.70 ±1.02 2.39 ±0.76 FLUX.2-dev 19.15 ±0.78 49.86 ±2.15 24.26 ±1.96 12.42 ±1.52 5.91 ±1.12 2.24 ±0.74 HunyuanImage-2.1 18.79 ±0.78 45.53 ±2.15 19.64 ±1.83 14.84 ±1.63 8.61 ±1.31 4.29 ±0.98 Qwen-Image-2512 5.79 ±0.47 19.16 ±1.75 5.87 ±1.14 2.11 ±0.73 1.10 ±0.56 0.15 ±0.29 HiDream-I1-Full 4.22 ±0.41 16.47 ±1.65 3.06 ±0.87 0.80 ±0.50 0.20 ±0.31 0.00 ±0.19 FLUX.1-dev 3.87 ±0.40 15.51 ±1.62 2.18 ±0.75 0.96 ±0.53 0.15 ±0.29 0.00 ±0.19 FLUX.1-schnell 2.88 ±0.35 11.96 ±1.46 1.66 ±0.67 0.25 ±0.34 0.10 ±0.26 0.00 ±0.19 SD3.5 Medium 2.81 ±0.34 12.01 ±1.47 1.19 ±0.59 0.35 ±0.37 0.05 ±0.23 0.00 ±0.19 SD3.5 Large 2.45 ±0.32 10.47 ±1.39 1.14 ±0.58 0.20 ±0.32 0.05 ±0.23 0.00 ±0.19 SDXL 1.0 0.02 ±0.05 0.10 ±0.25 0.00 ±0.20 0.00 ±0.19 0.00 ±0.19 0.00 ±0.19 Sana 1.6B 0.00 ±0.04 0.00 ±0.18 0.00 ±0.20 0.00 ±0.19 0.00 ±0.19 0.00 ±0.19 As shown in Table 3 , the strongest open-weight model, FLUX.2-dev, solves 19.15% of VVRBench tasks, and only 2.24% in the high complexity bin C 5 C_{5} . Every other open-weight model solves less than 19%. GPT-Image-2 solves 86.86% of all tasks, but its accuracy falls from 97.65% in C 1 C_{1} to 65.72% in C 5 C_{5} . On VVRBench -Fast (Figure 2 ), GPT-Image-2.5-Sunburst and GPT-Image-2 solve 84.51% and 82.20% of the tasks, respectively, leading other API models by a large margin (exact scores in Appendix D.2 ). Figure 2: API models on VVRBench -Fast. Table 4: VVRBench -Challenge acc. (%). The best model solves 21.39% overall and 7.92% at complexity 69 to 80. Model Accuracy ↑ \uparrow 45 to 56 57 to 68 69 to 80 GPT-Image-2.5-Sunburst 21.39 ±3.14 31.67 ±6.13 24.58 ±5.82 7.92 ±4.12 GPT-Image-2 10.28 ±2.43 17.50 ±5.31 10.83 ±4.57 2.50 ±2.85 Gemini-3.1-Flash-Lite-Image 7.36 ±2.14 10.00 ±4.45 4.17 ±3.33 7.92 ±4.12 Gemini-3-Pro-Image 4.58 ±1.78 7.08 ±3.97 4.17 ±3.33 2.50 ±2.85 Gemini-3.1-Flash-Image 3.89 ±1.67 3.33 ±3.11 3.75 ±3.22 4.58 ±3.44 Gemini-2.5-Flash-Image 1.11 ±1.07 2.50 ±2.85 0.42 ±1.91 0.42 ±1.91 GPT-Image-1-mini 0.28 ±0.73 0.83 ±2.15 0.00 ±1.58 0.00 ±1.58 VVRBench -Challenge separates frontier models. With tasks in the complexity range of 3–44, VVRBench -Fast barely separates the strongest frontier models, GPT-Image-2.5-Sunburst and GPT-Image-2, with a 2.3-points margin. Therefore, we create VVRBench -Challenge by sampling tasks from the uniform complexity distribution of 45–80 over a wider range of constraints using the VVR generator (Appendix B.5 ). As shown in Table 2 , VVRBench -Challenge discriminates among frontier models and exposes new failure modes. GPT-Image-2.5-Sunburst solves 21.39% of Challenge tasks, twice the 10.28% of GPT-Image-2, and its accuracy falls from 31.67% at complexity 45–56 to 24.58% at 57–68 and 7.92% at 69–80. Every other model solves at most 8% of VVRBench -Challenge, suggesting that there is still large room for improvement. Interestingly, we observe occasional abstention behaviors from all Gemini models, stating the instruction is unsatisfiable, demonstrating failure in spatial reasoning (Appendix D.3 ). The uniform drop of model accuracy across increasing complexity bins validates the design of the structural complexity score as a model-independent heuristic to generate tasks with controlled difficulty. Appendix B.4 provides more details on complexity as a predictor of failure. Finding 1. Open-weight generators fail most VVRBench tasks, and even the strongest API models lose accuracy sharply as complexity grows. 3.2 Failures concentrate in counting and object matching The binary pass-fail score (Eq. 2 ) is composed of individual verifier decisions from each of the active constraints in each task. Figure 3 presents the constraint-level pass rate of the API models on VVRBench -Challenge. Organized by constraint families (Table 1 ), 95% of Grounding constraints are satisfied, while only 56% of Topology constraints are rendered, averaged across models. The hardest constraint types concern counts or relations across object groups: same count passes in 31% of checks, times as many in 33%, and each contains , which requires each object of one group to contain a different object of another group, in 34%. Appendix D.4 gives the pass rate of every constraint type, Appendix D.5 correlates each family to accuracy across all models at matched complexity, and Appendix D.6 shows typical failures in which a model adds objects that the prompt excludes. Finding 2. Generators satisfy requirements on individual objects and pairs but fail requirements that constrain whole sets of objects. Figure 3: Pass rates of individual constraint for API models on VVRBench -Challenge, by family (left) and for the four constraint types with the lowest and the four with the highest average pass rates among those with at least 100 checks per model (right). 4 RLVVR: VVR for Diffusion Post-Training VVR scores images with program verifiers, avoiding error propagation from unreliable learned evaluators, and is not limited to fixed prompt sets, so training can be scaled to any desired data size and difficulty distributions. These properties allow us to improve image generation instruction following by using VVR as a reward in reinforcement learning (RLVVR). In this section, we post-train image generators with RLVVR to answer the following research questions: RQ1. Does RLVVR teach precise instruction following, and how does the complexity of the training tasks shape what is learned? RQ2. Do the skills learned from synthetic scenes transfer to natural prompts beyond VVR? RQ3. Is supervision from synthetic scenes complementary to existing post-training rewards? 4.1 Experimental setup Data. We generate two training corpora using the VVR generator (§ 2.2 ). VVR-Easy contains tasks that contain at most one constraint family and have complexity of at most 20, and VVR-Matched matches the VVRBench distribution (complexity 3–48). Each dataset contains 100K VVR tasks after decontamination from benchmark data (Table 2 ). For reward-mixture experiments, we train with GenEval2 ( Kamath et al., 2025 ) , OCR ( Liu et al., 2025a ) , and a five-reward objective that combines GenEval ( Ghosh et al., 2023 ) , GenEval2, OCR, PickScore ( Kirstain et al., 2023 ) , and UnifiedReward ( Wang et al., 2025 ) . Each of these objectives is trained alone and mixed with VVR-Easy, with equal number of prompts per objective. Training. We train Stable Diffusion 3.5 Medium ( Esser et al., 2024 ; Stability AI, 2024 ) with Flow-GRPO ( Liu et al., 2025a ) . Reward for each rollout is assigned by the scorer of the task objective that its prompt comes from. We use the VVR dense score r dense r_{\mathrm{dense}} for VVR prompts (Eq. 3 ). We name each trained model after its training data. Appendix E reports training details. Evaluation. We evaluate trained models on VVRBench , GenEval, GenEval2, OCR, PickScore, HPSv2.1, CLIPScore, aesthetic score, ImageReward, HPSv3, and UnifiedReward (Appendix E.1 ). 4.2 RQ1: RLVVR teaches precise instruction following Training on VVR-Easy raises VVRBench accuracy from 2.81% to 28.27% (Figure 5 ). Every task in C 3 C_{3} – C 5 C_{5} is more complex than any VVR-Easy task, and on these ranges accuracy still rises by 17.16, 8.31, and 1.35 points. Training on data from the benchmark distribution with VVR-Matched raises accuracy to 46.60% overall and to 45.62%, 38.39%, and 21.82% on C 3 C_{3} – C 5 C_{5} (Appendix F.1 ). Figure 4: VVRBench accuracy by complexity range. Shaded ranges lie above the complexity of every VVR-Easy training task, and VVR-Easy improves them. Training on harder generated tasks (VVR-Matched) closes more of the gap. Figure 5: VVR-Easy closes the gap between pretrained and VVR-Match more effectively on partial scores (individual constraint) than all-satisfy scores (compositionality) on both count and relation constraints. We separate how reliably a model satisfies individual constraints from how well it satisfies compositional requirements, using the two kinds of constraints that nearly every complex task contains: counts and relations . We compare the models’ partial scores q a ( x , s ) q_{a}(x,s) on these constraints with how often they satisfy every count or every relation (Appendix F.2 ). On tasks outside of its training complexity range, VVR-Easy closes 79% and 64% of the gap between the pretrained model and VVR-Matched in the partial scores of counts and relations, respectively, but only 55% and 46% in how often all counts or all relations in a task are satisfied (Figure 5 ). Easy tasks thus make individual constraints reliable, and satisfying many constraints in the same image is learned from large scenes. Finding 3. Training only on easy tasks makes individual constraints reliable, including on harder tasks, and training on large scenes teaches compositionality. 4.3 RQ2: Skills learned from synthetic scenes transfer to natural prompts Trained only on colored shapes, VVR-Easy improves over the pretrained reference on eight of ten non-VVR metrics, including GenEval by 0.113 and OCR by 0.111 (Table 5 ). Human annotators confirm the transfer: VVR-Easy is preferred by annotators over the pretrained SD3.5-M on their generations from 160 natural prompts outside VVR with a win rate of 71.6% with 83.8% pairwise agreement (Table 6 ). Appendix G reports annotation details. VVR-Easy also scores higher on 6 out of 9 natural prompt benchmarks than the model trained with the GenEval2 reward, whose training prompts name real objects—especially GenEval (0.729 vs. 0.688) and OCR (0.587 vs. 0.501). Notably, the GenEval gain comes from position ( + 0.150 +0.150 ), counting ( + 0.103 +0.103 ), and color attribution ( + 0.025 +0.025 ), all skills that VVR trains, while the single object, two object, and colors categories are comparable to the GenEval2-trained model. Finding 4. Skills learned from synthetic VVR scenes transfer to natural prompts, especially in position and counting. Table 5: RLVVR with VVR-Easy and reward mixtures transfers to most benchmarks and metrics. Training reward VVR GenEval GenEval2 OCR PickScore HPSv2.1 HPSv3 CLIPScore Aesthetic ImageReward UnifiedReward Pretrained 0.028 0.616 0.237 0.476 0.841 0.300 7.689 0.956 5.517 0.929 0.636 + + VVR-Easy 0.283 0.729 0.268 0.587 0.849 0.294 8.275 0.979 5.483 1.114 0.641 Δ \Delta +0.255 +0.113 +0.031 +0.111 +0.008 − 0.006 -0.006 +0.586 +0.023 − 0.034 -0.034 +0.185 +0.005 GenEval2 0.039 0.688 0.454 0.501 0.848 0.297 8.289 0.974 5.523 1.110 0.635 + + VVR-Easy 0.218 0.718 0.478 0.532 0.848 0.302 8.511 0.976 5.533 1.155 0.637 Δ \Delta +0.180 +0.030 +0.025 +0.030 +0.0004 +0.005 +0.222 +0.002 +0.009 +0.045 +0.0019 + + VVR-Matched 0.335 0.712 0.491 0.510 0.848 0.301 8.498 0.972 5.516 1.148 0.637 Δ \Delta +0.296 +0.024 +0.038 +0.009 − 0.0002 -0.0002 +0.004 +0.209 − 0.001 -0.001 − 0. The google story also surfaces in Google Research Releases ToolGrad Framework With..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!