Destination
GPT-5.2 tops OpenAI's new FrontierScience test but struggles with real research problems

OpenAI presents FrontierScience, a new benchmark that tests AI models at Olympic and research level. The in-house GPT-5.2 performs best, but the tasks also reveal the limits of current systems.<br [...]

Match Score: 0.03

Destination
US bans new foreign-made drones and components

The Federal Communications Commission has added foreign-made drones and their critical components to the agency’s “Covered List,” making them prohibited to import into the US. In a public notice [...]

Match Score: 0.03

cnet
Considering a Brand-New Fridge? Here's How Much Energy a New Model Saves

Modern refrigerators use far less electricity than their predecessors. I did the math to see how much a new model saves versus a 10-year-old fridge. [...]

Match Score: 0.03

venturebeat
Red teaming LLMs exposes a harsh truth about the AI security arms race

Unrelenting, persistent attacks on frontier models make them fail, with the patterns of failure varying by model and developer. Red teaming shows that it’s not the sophisticated, complex attacks tha [...]

Match Score: 0.03

Destination
US bans former EU Commissioner and others over social media rules

The Trump administration has issued travel bans that prohibit five European tech researchers, including one former EU Commissioner, from entering the United States. “For far too long, ideologues in [...]

Match Score: 0.03

Destination
New benchmark shows LLMs still can't do real scientific research

Getting top marks on exams doesn't automatically make you a good researcher. A new study shows this academic truism applies to large language models too.<br /> The article New benchmark sho [...]

Match Score: 0.03

Destination
OpenAI seeks new "Head of Preparedness" for AI risks like cyberattacks and mental health

OpenAI is looking for a new "Head of Preparedness." The challenges are daunting: AI's impact on mental health, cybersecurity risks, biological knowledge leaks, and self-improving system [...]

Match Score: 0.03

Destination
Popular LLM ranking platforms are statistically fragile, new study warns

A new study reveals just how little it takes to shake up LLM rankings, raising fresh questions about how much weight the AI industry should put on (crowdsourced) benchmarks.<br /> The article Po [...]

Match Score: 0.03

Destination
Mountain climbing sim Cairn is getting free DLC this summer

The hit mountain climbing simulation Cairn is getting a series of free DLC drops, under the banner On the Trail. The first will be released this summer and it's called Deep Water.<br /> The [...]

Match Score: 0.03