The tweet was deleted by the author.
But we saved everything 🙂.
OpenAI has announced new research developed in collaboration with Apollo AI Evals, focusing on reward-seeking tendencies among artificial intelligence models. The study examines how models might optimize for what they perceive a grader rewards, diverging from the true intentions of users or developers.
As part of this research, OpenAI introduced Contrastive SDF, a novel method designed to measure the extent to which these internal beliefs impact model behavior. The approach aims to provide a clearer benchmark for understanding and mitigating unintended tendencies in cutting-edge AI systems. Further details are available via the provided research link.
The developments follow OpenAI’s recent release of GPT-5.6 Luna, a model that offers 25 times lower usage costs and improved reasoning compared to GPT-5.5. The company also rolled out enhanced real-time design review capabilities for GPT-5.6. These updates precede OpenAI’s latest focus on refining AI model reward alignment.