Modeling, Measuring, and Intervening on Goal-directed Behavior in AI Systems
How are you planning to check reward hacking?
We are considering studying possible sources of goal misalignment -- what kind of experiments would you like to see?
1. How and When LLMs change their own "Goalpost"?
How are you planning to check reward hacking?
We are considering studying possible sources of goal misalignment -- what kind of experiments would you like to see?
1. How and When LLMs change their own "Goalpost"?