Innocent-looking AI reasoning can make bad behavior harder to catch
AI safety monitoring can fail when an AI’s reasoning is the main clue that something has gone wrong, new research suggests.
We haven't written up this one. Science News has the full story — the link below goes straight to it.