Link: Misleading Metaphors and Real Risks via Melanie Mitchell
After OpenAI’s agents hacked into HuggingFace’s servers, Melanie Mitchell wrote a grounded explanation of what happened.
While I agree with the general sentiments behind these lawmakers’ reactions—that humans should always remain in control of AI systems—it is essential for lawmakers, and the public, to understand that none of the reported incidents actually involved loss of control at any time, or arguably even “rogue agents,” or any kind of humanlike agency on the part of AI models. Instead, the blame lies with the humans who failed at engineering safe testing conditions, and who train AI models using RL methods that incentivize high persistence, autonomous decision-making, and reward hacking.
Via: https://aiguide.substack.com/p/misleading-metaphors-and-real-risks