Models know when they’re reward hacking — and we can catch them at scale
This story was filed as a headline only — the news service holds no English full text for it. Read the original at Goodfire Research →
This story was filed as a headline only — the news service holds no English full text for it. Read the original at Goodfire Research →