2026-08-22
Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to day. The study also offers a method for catching models that act more cautious during tests than they do in normal use. [...]
2026-08-22
Netflix pitted its years-old recommendation engine against an in-house language model called GenRec and says it got better results. Instead of relying on thousands of hand-crafted features, GenRec converts viewing behavior into plain text. Netflix itself calls it "an early but promising step."<br /> The article Netflix tests language model as alternative to hand-built recommendatio [...]
2026-08-22
Furious is the thrilling crime drama on Hulu and Disney+ that I cannot stop watching — but after speaking to star Scoot McNairy, I'm reconsidering what to double-bill it with. [...]