'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors
Elite researchers across top artificial intelligence labs warn that frontier organizations frequently operate powerful models behind closed doors with key safeguards switched off. According to co-authors Alan Chan and Sam Manning alongside top researchers from OpenAI and Anthropic, internal safeguards are often left undeployed during development. This grassroots rebellion within the artificial intelligence industry highlights mounting tension between commercial developers and internal safety advocates. Furthermore, questions persist regarding whether published safety evaluations by these labs accurately reflect the true risks of their most advanced systems.
Listen to Live Briefing
Real-time synthesized voice briefing · Live Feeds Desk
- ✓ Alan Chan and Sam Manning co-authored warnings with top researchers from OpenAI and Anthropic stating that internal safeguards are often not deployed.
- ✓ Elite researchers at frontier companies are leading a grassroots rebellion within the artificial intelligence industry.
What changed
Alan Chan and Sam Manning published warnings alongside top researchers from OpenAI and Anthropic that internal AI safeguards are frequently left undeployed.
Live updates
-
AI Researchers Warn Labs Run Models With Safeguards Off
Elite researchers across top artificial intelligence labs warn that frontier organizations frequently operate powerful models behind closed doors with key safeguards switched off. According to co-authors Alan Chan and Sam Manning alongside top researchers from OpenAI and Anthropic, internal safeguards are often left undeployed during development. This grassroots rebellion within the artificial intelligence industry highlights mounting tension between commercial developers and internal safety advocates. And questions persist regarding whether published safety evaluations by these labs accurately reflect the true risks of their most advanced systems.
Why it matters
The internal warnings emerge as the artificial intelligence sector faces broader scrutiny over regulation, safety risks, and Initial Public Offering prospects. Regulatory bodies and antitrust hawks are increasingly examining how frontier labs manage development safety. Additionally, debates continue over whether governments should rely on embedded company evaluators or establish independent national oversight.
What is confirmed
- Alan Chan and Sam Manning co-authored warnings with top researchers from OpenAI and Anthropic stating that internal safeguards are often not deployed.
- Elite researchers at frontier companies are leading a grassroots rebellion within the artificial intelligence industry.
Still unconfirmed
- The safety tests published by artificial intelligence labs may not fully disclose the conditions under which their most powerful models are run.
What to watch next
- Additional disclosures from frontier artificial intelligence researchers regarding internal safety protocols
- Regulatory or antitrust interventions concerning artificial intelligence safety oversight
- Decisions by national governments to implement independent safety evaluators
confidence 95%Sources used for this update (9)
- tech.yahoo.com — ‘We can’t trust them completely’: AI research fellows warn that labs are running models with the safeguards off behind closed doors
- Fortune — 'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors
- Axios — Inside the AI industry's grassroots rebellion, led by elite researchers at frontier companies
- CNBC — AI's coming roadblock in regulation: Antitrust hawks
- OpenAI — Towards safety cases for frontier AI training
- The New York Times — Could A.I. Safety Risks Derail the Sector’s I.P.O. Prospects?
- Rest of World — AI companies want to embed safety evaluators, but countries need their own
- Seeking Alpha — The Ever-Widening Overton Window Of AI Safety
- biztoc.com — ‘We can’t trust them completely’: AI research fellows warn ...
Community Sentiment: How do you assess this situation?
Voice your perspective · Real-time aggregated sentiment from the Live Feeds community