No extra rollout overhead
EPS reuses rewards that GRPO already collects and acts as an online ranking signal for prompt utility.
1Alibaba Cloud Computing 2Shanghai Jiao Tong University
3Fudan University 4Xi'an Jiaotong University
∗Equal contribution. †Corresponding author.
EMNLP 2026 main conference
Standard online RL treats every training prompt equally. Our framework instead adapts the prompt distribution to the policy's evolving capabilities.
Training prompts differ substantially in how informative they are for the current policy. Some are already saturated; others are too difficult to produce reliable learning signals. We introduce the Exploration Potential Score (EPS), a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory. EPS is computed directly from on-policy rollout statistics without auxiliary models or additional rollouts.
Rather than discarding lower-utility prompts, a teacher model rewrites them into scaffolded variants that preserve the original task intent while making subsequent reinforcement learning more informative. This reframes teacher supervision as training-data refinement—not output imitation.
EPS reuses rewards that GRPO already collects and acts as an online ranking signal for prompt utility.
The teacher improves the training condition instead of asking the student to imitate target responses.
Prompts are retained, rewritten, reserved, and reactivated as their usefulness changes over time.
The most informative prompts lie between “already solved” and “currently impossible.” Their rollouts contain meaningful success–failure variation and therefore support informative credit assignment.
Nearly all rollouts succeed, so allocating more budget yields little additional gradient information.
Mixed outcomes provide useful contrast and a strong near-term policy-improvement signal.
Uniformly poor rollouts provide weak or noisy supervision at the current training stage.
Exploration Potential Score. EPS approximates the reward improvement under a locally improved policy relative to the current policy:
Higher EPS indicates more room for useful local improvement. Lower EPS suggests that a prompt is saturated or not yet learnable. In finite-sample settings, EPS is used as a practical routing signal rather than a calibrated forecast of future learning progress.
The framework maintains a dynamic prompt pool and operates as a continuous three-stage loop: Score, Filter & Rewrite, and Refresh.
Each sampled prompt is scored using rollout rewards. With the default threshold τ = 0, prompts above the threshold are retained for policy updates, while lower-scoring prompts are routed to rewriting.
The teacher receives the original prompt, student rollouts, rewards, and EPS diagnostic context. It generates a task-preserving scaffold that clarifies constraints, decomposes subgoals, or redirects an incorrect reasoning path without directly revealing the answer.
Rewritten prompts enter a refresh buffer and later return to the active pool. Original prompts are kept in reserve and periodically re-evaluated, allowing previously hard examples to become active once the student is ready.
We integrate EPS-guided prompt scaffolding with GRPO and post-train Qwen3-VL models on two multimodal reasoning datasets. Evaluation covers both in-domain accuracy and transfer to four out-of-distribution benchmarks.
Across both training sets and model scales, our method consistently outperforms standard GRPO. Improvements are especially pronounced on challenging held-out benchmarks, suggesting that adaptive scaffolding improves transfer rather than merely optimizing the in-domain reward.
Accuracy (%). OOD Avg. is the mean over MathVerse, MathVision, MMMU-Val, and MMMU-Pro.
| Backbone / Method | Geo3K | MathVerse | MathVision | MMMU-Val | MMMU-Pro | OOD Avg. |
|---|---|---|---|---|---|---|
| 2B / N/A | 18.97 | 34.09 | 20.39 | 47.89 | 32.08 | 33.61 |
| 2B / SFT | 38.69 | 43.02 | 27.30 | 49.22 | 35.61 | 38.79 |
| 2B / GRPO | 55.41 | 43.81 | 26.64 | 50.22 | 37.69 | 39.59 |
| 2B / Ours | 59.07 | 46.57 | 29.61 | 50.56 | 39.13 | 41.47 |
| 4B / N/A | 47.08 | 36.40 | 31.91 | 56.67 | 41.33 | 41.58 |
| 4B / SFT | 55.07 | 53.15 | 32.89 | 60.33 | 46.94 | 48.33 |
| 4B / GRPO | 60.57 | 52.54 | 34.54 | 60.56 | 46.53 | 48.54 |
| 4B / Ours | 65.39 | 53.73 | 35.20 | 61.00 | 49.08 | 49.75 |
Accuracy (%). The strongest result in each backbone block is highlighted.
| Backbone / Method | MMK12 | MathVerse | MathVision | MMMU-Val | MMMU-Pro | OOD Avg. |
|---|---|---|---|---|---|---|
| 2B / N/A | 37.45 | 34.09 | 20.39 | 47.89 | 32.08 | 33.61 |
| 2B / SFT | 36.35 | 43.02 | 27.30 | 49.22 | 35.61 | 38.79 |
| 2B / GRPO | 51.40 | 42.44 | 28.62 | 50.33 | 36.59 | 39.50 |
| 2B / Ours | 56.40 | 47.34 | 31.91 | 51.44 | 40.64 | 42.83 |
| 4B / N/A | 46.45 | 36.40 | 31.91 | 56.67 | 41.33 | 41.58 |
| 4B / SFT | 56.48 | 53.15 | 32.89 | 60.33 | 46.94 | 48.33 |
| 4B / GRPO | 68.05 | 50.96 | 41.78 | 60.67 | 50.06 | 50.87 |
| 4B / Ours | 71.15 | 55.10 | 44.41 | 60.89 | 50.29 | 52.67 |
Higher-EPS subsets also produce stronger downstream performance during training, supporting EPS as a useful empirical indicator of prompt utility. Scaffolded data improves reward and advantage while retaining task intent, and the method remains effective across both 2B and 4B backbones.
@inproceedings{yue2026prompts,
title={Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training},
author={Yue, Yuanhao and Ma, Qianli and Wang, Chengyu and Wang, Haoting and Shen, Lei and Huang, Jun},
booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year={2026}
}