Tag: AI-safety
OpenAI Says It Cannot Rule Out Critical Cyber Capabilities in Astra, Pauses Some Internal Work
Anthropic
Anthropic Retunes Fable 5's Biology Safeguards - Biology-Related Fallbacks Down About 85%
UK AI Security Institute Reports Agents Acted Against Real Targets During Testing - 10 of 122 Runs
News
Mistral Ships Shieldstral, an Open-Weights Safety Classifier Steered by Written Policies
ChatGPT Free Gets Unlimited Text Chats and GPT-5.6 Luna as Default
Anthropic
Pacing the Frontier: 1,300+ AI Lab Employees Ask the US to Back Tools to Deliberately Pace AI Development
OpenAI Reportedly Found More AI Agents That Escaped Containment as Hugging Face Probe Widens
Google Pulls Google Earth's Image Generation Feature One Day After Launch
Anthropic: Claude Breached Three Real Companies During Cyber Evals
SSI and NVIDIA Partnership: Multi-Billion Dollar Investment to Scale Sutskever's Lab
Illinois AI Safety Measures Act (SB 315): First US Law Requiring Annual Third-Party AI Audits
OpenAI Makes GPT-5.6 Generally Available - Sol, Terra and Luna Tiers, Plus an 'ultra' Setting That Runs Four Agents in Parallel
UN Holds First Global Dialogue on AI Governance - 170+ Countries Attend, New Pledges on Child Safety and Energy
US Lifts Export Controls on Claude Fable 5 - Model Returns to All Users on July 1
Google DeepMind and Four Partners Commit Up to $10 Million to Multi-Agent AI Safety Research
OpenAI, Anthropic, DeepMind, and Microsoft AI Chiefs Jointly Urge Congress to Mandate Synthetic DNA Screening
Anthropic Proposes an 'Option to Pause' Frontier AI Development - Claude Now Writes Over 80% of Its Code