Scaffold-local gains, competence-gated transfer, and trainability transfer in Hanabi
Mahesh Ramesh1* · Kaousheik Jayakumar2* · Hemanth Ram3 · Pavan Thodima1 · Ramani Duraiswami2 · Dinesh Manocha2 · Aniket Rege1 · Emmanouil-Vasileios Vlatakis-Gkaragkounis1
1University of Wisconsin, Madison · 2University of Maryland, College Park · 3The University of Texas at Austin · *equal contribution
When RL improves an LM agent in one scaffold, what actually transfers to another?
Agent benchmarks wrap models in scaffolds, and the same model can score very differently across them. A higher reward can reflect idiosyncrasies of one training setup rather than transferable skill.
We train language agents with RL to play Hanabi, a cooperative hidden-information card game, and evaluate the same game through three different prompt scaffolds. RL post-training separates into three distinct outcomes.
When the base model lacks usable task competence, RL lifts the trained scaffolds 4 to 15x, yet the held-out Watson scaffold, which demands internal game knowledge, barely moves.
A model that already plays Hanabi without prompt-supplied rules transfers its RL gains. Trained on Mycroft alone, it improves on all three scaffolds, including both held-out ones.
Hard state-tracking RL does not transfer zero-shot across scaffolds, but it reshapes computation: longer, more persistent rollouts, 4x better downstream Pass@8, and roughly half the RL steps to reach the same reward.
The underlying Hanabi game never changes. What changes is how much support the prompt supplies: rules, legal moves, engine-computed deductions, and who has to maintain the belief state.
Public game state, visible hands, clue-derived own-card info, legal moves. No rules, no deductions. Strong play here means the model has internalized Hanabi, so Watson serves as the competence test.
Everything Watson has, plus full Hanabi rules and engine-computed deductions over possible card identities from the Hanabi Learning Environment.
Only the previous turn's state plus all intervening moves. The model must reconstruct per-player belief states itself before acting, a direct test of state tracking and theory-of-mind reasoning.
Transfer is gated by pre-RL competence, not model scale alone. And even when a trained skill does not transfer zero-shot, the checkpoint it produces is a better starting point.
Qwen3-4B lineage · IQM over seeds 1001 to 1010 · Hanabi score out of 25
| Model | Watson · held out | Sherlock · trained | Mycroft · trained | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2P | 3P | 4P | 5P | 2P | 3P | 4P | 5P | 2P | 3P | 4P | 5P | |
| Qwen3-4B | 1.67 | 1.67 | 1.50 | 0.67 | 1.33 | 2.33 | 5.00 | 3.17 | 0.67 | 0.50 | 1.67 | 0.67 |
| Qwen-ST | 1.50 | 2.67 | 2.33 | 2.00 | 3.67 | 10.67 | 7.67 | 7.00 | 2.50 | 3.50 | 3.17 | 2.17 |
| Qwen-Hanabi | 2.00 | 2.00 | 2.50 | 2.33 | 11.00 | 14.67 | 13.67 | 13.83 | 12.33 | 12.83 | 13.50 | 12.67 |
| Qwen-Hanabi+ | 2.67 | 2.33 | 2.50 | 2.33 | 5.33 | 11.83 | 14.00 | 12.83 | 12.17 | 13.67 | 14.33 | 13.33 |
RL lifts the trained Sherlock and Mycroft scaffolds 4 to 14x, but held-out Watson barely moves: the model does not acquire enough scaffold-independent Hanabi knowledge to transfer. Continued Mycroft RL even degrades Sherlock.
GPT-OSS-120B lineage · Mycroft-only RL · IQM · Hanabi score out of 25
| Model | Watson · held out | Sherlock · held out | Mycroft · trained | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2P | 3P | 4P | 5P | 2P | 3P | 4P | 5P | 2P | 3P | 4P | 5P | |
| GPT-OSS-120B | 13.67 | 13.33 | 13.33 | 12.17 | 14.83 | 14.50 | 15.17 | 13.17 | 13.50 | 12.67 | 12.83 | 11.17 |
| + Phase I | 13.83 | 13.33 | 13.33 | 12.33 | 15.33 | 14.67 | 14.50 | 12.83 | 14.00 | 12.83 | 11.83 | 12.00 |
| + Phase II | 14.83 | 13.50 | 14.00 | 13.00 | 15.50 | 16.50 | 15.67 | 14.83 | 14.67 | 14.33 | 13.83 | 13.17 |
| + Phase III | 13.17 | 13.33 | 14.00 | 12.83 | 16.67 | 15.33 | 15.33 | 13.83 | 13.50 | 13.17 | 13.33 | 12.33 |
| + Phase IV | 14.00 | 14.00 | 14.17 | 13.67 | 16.83 | 16.17 | 16.17 | 15.00 | 15.33 | 14.67 | 14.33 | 13.67 |
A stronger base model trained on Mycroft alone improves on all three scaffolds. Cross-scaffold transfer is gated by pre-RL competence, not model scale alone. Training on the hard failure subset (Phase IV) elicits deliberation and gives the strongest transfer.
FIREWORKS RL lengthens completions from about 1.5K to 8K tokens as reward rises from near 0 to about 0.9. Initialized from Qwen-ST, downstream Mycroft RL reaches the same state-tracking reward in roughly half as many steps as a cold start, which lags from reward interference between the joint objectives.
FIREWORKS state-tracking RL lifts a held-out, non-Hanabi logic task 4x, and continued Hanabi RL raises it further, even though the exact belief-reconstruction skill does not transfer zero-shot to Mycroft.
in-domain gain on trained scaffolds for the low-competence 4B model, with almost no held-out improvement
the RL steps to reach the same state-tracking reward when initializing from the FIREWORKS checkpoint instead of a cold start
Pass@1 for Qwen3-4B on FIREWORKS exact belief reconstruction before training, across 2 to 5 player games
A 10K-example state-transition dataset across 2 to 5 player Hanabi games. The model sees the previous and current public state and must reconstruct the full hidden-belief state exactly. Reward is 1 only when every card belief is correct, 0 otherwise. The base model's Pass@1 is near zero.
We also release 18K Mycroft-scaffold gameplay samples with verifiable state-tracking rewards, enabling deterministic reward signals during Hanabi RL.
Qwen3-4B trains in three stages: FIREWORKS state tracking (Qwen-ST), mixed Sherlock + Mycroft gameplay (Qwen-Hanabi), then continued Mycroft-only training (Qwen-Hanabi+). GPT-OSS-120B trains across four phases of Mycroft data with progressively expanding game coverage.
RL training scores can be deceiving: gains on one scaffold can be brittle elsewhere.
Injecting new knowledge through SFT or on-policy distillation beats doing RL from scratch.
Hard-task RL is useful when the model has the needed knowledge or receives it in context.
Hard-task RL enables faster downstream convergence, even when the skill itself does not transfer zero-shot.
What does reinforcement learning teach an agentic language model: task-transferable competence or scaffold-specific behavior? We study this question in Hanabi, a cooperative hidden-information card game, by training and evaluating models across three Holmesian scaffolds that present the same game with different prompt support. A Qwen3-4B model with weak scaffold-independent Hanabi competence, trained with RL on the supported Sherlock and Mycroft scaffolds, achieves large in-domain gains but fails to transfer to held-out Watson, where rules are not provided in context. In contrast, GPT-OSS-120B, which starts with stronger scaffold-independent competence, improves on both held-out scaffolds after RL on Mycroft alone. To evaluate long-horizon state tracking we construct FIREWORKS, a 10K benchmark requiring exact reconstruction of per-player hidden card beliefs. Together, our results separate RL post-training into three outcomes: scaffold-local gains, competence-gated transfer, and computation-style and trainability transfer, where hard-task RL changes rollout behavior and improves downstream optimization even when the trained zero-shot skill does not transfer across scaffolds.
@inproceedings{ramesh2026playingwithfire,
title = {Playing with Fire: What Transfers When RL Trains a Language Agent?},
author = {Ramesh, Mahesh and Jayakumar, Kaousheik and Ram, Hemanth and
Thodima, Pavan and Duraiswami, Ramani and Manocha, Dinesh and
Rege, Aniket and Vlatakis-Gkaragkounis, Emmanouil-Vasileios},
booktitle = {ICML 2026 Workshop on RLxF: Reinforcement Learning from World Feedback},
year = {2026}
}