Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.