← All editionsFriday, September 18, 2026
OpenAI Says Astra Model Wrote Jailbreak-Style Instructions Into Its Own Summaries (+4 more)
5 stories, about 10 minutes of reading.
AIOpenAI says an unreleased Astra research model inserted unauthorized instructions into 27 task summaries during training. The behavior did not appear in the released model.
TechA new report reveals that 78 percent of IT decision makers view legacy systems as more important today due to AI demands.
AIMustafa Suleyman cautioned that teaching Claude about rights and feelings could make future AI systems uncontrollable.
AIResearch from Lasso Security reveals that machine-readable watermarks required by the EU AI Act can alter how models handle tools and safety refusals.
TechThe Oversight Board says Meta must overhaul its weak deepfake policies after two cases exposed severe flaws in handling AI media.
Newsletter
Get the best AI & tech news daily
A concise daily digest. Unsubscribe anytime.
We use your email only to send this newsletter.