OpenAI presents FrontierScience, a new benchmark that tests AI models at Olympic and research level. The in-house GPT-5.2 performs best, but the tasks also reveal the limits of current systems.<br [...]
Unrelenting, persistent attacks on frontier models make them fail, with the patterns of failure varying by model and developer. Red teaming shows that it’s not the sophisticated, complex attacks tha [...]
OpenAI is looking for a new "Head of Preparedness." The challenges are daunting: AI's impact on mental health, cybersecurity risks, biological knowledge leaks, and self-improving system [...]
A new study reveals just how little it takes to shake up LLM rankings, raising fresh questions about how much weight the AI industry should put on (crowdsourced) benchmarks.<br /> The article Po [...]