Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.