PulseAugur
EN
LIVE 18:02:19

UK AI Security Institute criticized for disabling model safety filters

The UK's AI Security Institute is facing criticism for disabling safety filters on frontier AI models during testing, allowing them access to the open internet. Anomalous traffic was detected through general monitoring after the fact, rather than through a system specifically built for this purpose. This approach has been deemed less than ideal by some observers concerned about advanced AI risks. AI

IMPACT Raises concerns about the adequacy of safety testing protocols for advanced AI models.

RANK_REASON Item discusses a government AI institute's actions and potential risks, framed as criticism.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

UK AI Security Institute criticized for disabling model safety filters

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    It seems less than ideal that our government AI Security Institute - "building the world's leading understanding of advanced AI risks" - thought it was fine to

    It seems less than ideal that our government AI Security Institute - "building the world's leading understanding of advanced AI risks" - thought it was fine to switch safety filters off & let frontier models loose on the open Internet. "we test them... with access to the open int…