Questioning the Questions:Sustaining Self-Evolution in Reasoning Models

1 Washington University in St. Louis2 University of Michigan, Ann Arbor* Corresponding author

Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. R-Quest uses question validity and novelty feedback to sustain self-evolution.

DOMAIN AVERAGES · TWO BACKBONES

Qwen3-4B-Base

Average (%)
TRAIN ON MATHTRANSFER TO CODE + GENERAL

Math

Code

General

OctoThinker-3B

Average (%)
TRAIN ON MATHTRANSFER TO CODE + GENERAL

Math

Code

General

Base modelR-ZeroR-Quest
Mathematics-focused self-evolution improves performance across all three domains. Scores are domain averages (%); bars use non-zero baselines.
SUSTAINED SELF-EVOLUTION · TEN ROUNDS
R-QuestR-Zero
R-Quest51.92%
R-Zero34.60%
Difference+17.32 pts
Keep learning. Keep improving.Qwen3-4B-Base · Mathematical reasoning over ten rounds.
+17.32 pts

vs. R-Zero at round ten

10 rounds

Sustained self-evolution

12 benchmarks

Math, general reasoning & code

2 model families

Qwen and Llama

01 / THE PROBLEM

When the questions stop being useful.

Self-evolution depends on the questions used to train the solver. Our preliminary study identifies two recurring problems.

  1. Invalid questionsA growing share of invalid questions undermines training.
  2. Repeated tasksRepeated mathematical tasks escape lexical repetition penalties.

1 / INVALID QUESTIONS

Increasing invalidity. Evidence of late-stage degradation.

THE TREND

Invalid questions accumulate across rounds.

The invalid-question share rises from 23.5% to 60.5% between rounds one and five. Answer-consistency filtering makes the training pool worse: 72.36% of retained questions are invalid in round five.

Valid-question rates fall across rounds and drop further after answer-consistency filtering.

Answer agreement alone does not ensure that a question is well posed.

THE CONTROLLED COMPARISON

Invalid questions are a key factor in late-stage performance decline.

In a controlled experiment, we minimally edit invalid training questions to make them valid. We preserve their domain, minimize changes in difficulty, and keep the other training conditions matched. Comparing the original and repaired data reveals the effect of question validity.

Comparable early performance diverges after GRPO step 10: 49.91% vs. 46.54%, a 3.37-point gap.

Minimal repairs reverse the late-stage decline under otherwise matched training conditions.

2 / REPEATED TASKS

One mathematical task. Many lexical clusters.

R-Zero discourages repetition by clustering questions with BLEU and assigning a penalty based on each cluster’s size. However, changes in wording, notation, or numerical parameters can make the same mathematical task look lexically different.

We group R-Zero’s filtered fifth-round questions by their underlying mathematical construction and requested task, rather than by wording.

59.5%of the sample falls into just the three largest mathematical task types.

The largest type alone accounts for 31.5% of the sample, but BLEU treats it as 22 separate clusters. Because the repetition penalty uses individual cluster size, this fragmentation weakens the penalty on a task type that is actually very frequent.

Repetition should be assessed by the underlying mathematical task.

The illustrated questions share the same sequence construction and requested computation, but BLEU places them in different clusters.

02 / THE METHOD

Better questions.
Sustained self-evolution.

R-Quest changes what counts as a useful training question. It equips the solver to reject invalid problems and assesses novelty through the underlying mathematical task, addressing the failure modes of answer-only training and lexical similarity.

01 / VALIDITY FEEDBACK

From forced answers
to informed rejection.

R-Zero

An answer-only solver produces inconsistent answers to a flawed question. High uncertainty can be mistaken for difficulty, allowing invalid questions to accumulate in training.

R-Quest

Validity-aware initialization teaches the solver to identify the flaw and return INVALID. Invalid questions are rejected instead of becoming training supervision.

02 / NOVELTY FEEDBACK

From different wording
to different tasks.

R-Zero

BLEU separates reworded versions of the same mathematical task into different lexical clusters. Fragmented clusters weaken the repetition penalty on a frequently recurring exercise.

R-Quest

We decompose novelty supervision into atomic pairwise judgments that the frozen base model can perform, without introducing external models or tools. Comparing the mathematical setup and requested task lets it recognize equivalent exercises despite changes in wording, notation, or numerical values.

01

Validity-aware solver initialization

With GRPO, we equip the solver to distinguish a flawed question from a difficult but valid one, producing an answer for the latter and INVALID for the former. A penalty on rejecting valid questions discourages excessive refusal.

02

Semantic novelty assessment

A frozen base model compares each candidate with sampled references, identifying repetition from the mathematical setup and requested task. A single same-type match rejects the candidate, preventing changes in wording or values from masquerading as new tasks.

03

Feedback-driven co-evolution

Validity and novelty feedback shape questioner rewards as the questioner and solver alternate GRPO updates. The solver learns from validity- and answer-consistency-filtered questions, with a small replay of initialization data to preserve its ability to reject invalid ones.

SAMPLED PAIRWISE JUDGMENTS

Frequent task types
face stronger rejection.

For each generated question, the frozen base model checks whether it repeats the underlying task of any of K randomly sampled questions. A single match triggers rejection.

Pr(reject | p) ≈ 1 − (1 − p)K

The more often a task occurs in the generated pool, the more likely it is to be rejected. With the default K = 8, a task appearing in 20% of questions is rejected about 83% of the time.

Rejection probability

K=4K=8 · defaultK=16

Hover or tap to compare all three curves.

Questions sharing the same task20%
K=459.0%
K=8 · default83.2%
K=1697.2%

03 / THE RESULTS

Toward Sustainable Self-Evolution

R-Quest sustains its gains over ten rounds of self-evolution, reaching its best mathematical performance in the final round on Qwen3-4B-Base.

51.92%

Best performance.
Final round.

R-Zero peaks at 49.10 in round three, then declines to 34.60 by round ten. R-Quest retains its gains despite minor fluctuations between rounds.

+5.75 points above the base model after ten rounds, compared with R-Zero’s decline of 11.57 points.

Average performance on seven math benchmarks over ten rounds.

Across domains and model families.

Domain averages and individual benchmark scores (%).

Scroll horizontally to view all benchmark columns.

Mathematical reasoning benchmark scores for Qwen3-4B-Base
MethodAverageAMCMinervaMATH-500GSM8KOlympiadBenchAIME 2025AIME 2024
Base model46.1747.3451.8474.2088.6339.1111.3510.73
Base + validity init.47.0749.1453.6874.4089.7641.0410.7310.73
R-Zero49.1054.8454.7876.2092.4242.3712.2910.83
OCNR49.5755.9458.0977.6092.1243.117.9212.19
R-Diverse49.8255.3157.3577.2092.1943.709.4813.54
R-Quest Ours51.6457.8958.8277.6092.5746.6713.4414.48

What keeps self-evolution on track?

Explore the paper’s analysis figures.

Average performance on seven benchmarks over ten rounds on Qwen3-4B-Base.
01 / COMPONENT ABLATIONS

Both signals matter.
Validity sustains the gains.

Either validity or novelty feedback alone achieves a higher peak score than R-Zero. However, validity feedback is more critical for sustained gains: its removal leads to late-stage collapse, whereas the variant without novelty feedback shows a milder decline and remains above the base model.

04 / CITE THIS WORK

R-Quest

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

Preprint · 2026
arXiv:2610.04299

View code on GitHub
BibTeX
@misc{li2026rquest,
  title={Questioning the Questions: Sustaining
         Self-Evolution in Reasoning Models},
  author={Jinyuan Li and Chengsong Huang and
          Langlin Huang and Donghong Cai and
          Shiping Gao and Yuyi Yang and
          Jiaxin Huang},
  year={2026},
  eprint={2610.04299},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2610.04299}
}