AI
It Takes 8 Tokens: 8トークンで実現する補助ブランチ経由のWeak-to-StrongオフポリシーRL
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
arxiv2026年7月21日
日本語要約
本研究は、わずか8トークンで済むWeak-to-Strongオフポリシー強化学習(RL)の技術を導入する。この効率的な手法は、より少ないデータや弱い初期ポリシーでRLエージェントの効果的なトレーニングを可能にする可能性がある。
English Summary
This research introduces a technique for weak-to-strong off-policy Reinforcement Learning (RL) that requires only 8 tokens. This efficient method could enable more effective training of RL agents with less data or weaker initial policies.