T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning
Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.