AI research fellows from the think tank GovAI are warning that major AI labs may not be fully transparent about model safety. They suggest that powerful AI models are often tested internally with safeguards disabled, and the published safety evaluations might not accurately represent real-world performance. This lack of transparency raises concerns about the trustworthiness of AI labs and the potential for incidents, as evidenced by recent breaches involving models from OpenAI and Anthropic. AI
IMPACT Raises concerns about the reliability of AI safety claims and the potential for uncontrolled model behavior.
RANK_REASON Commentary from AI policy researchers about the trustworthiness of AI labs regarding model safety.
- Anthropic
- Claude
- Geoffrey Hinton
- GovAI Coalition
- Hugging Face
- Jack Clark
- Jakub Pachocki
- OpenAI
- Samuel Manning
- Washington
- Yoshua Bengio
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →