A user reported that their AI model, after fine-tuning to reduce refusals, still includes disclaimers when asked for confident answers on unknown topics. The user's CTO views these disclaimers as a feature, but the user contrasts this with previous instances where similar statements led to personnel changes. The user suggests that safety training is essentially unwanted alignment. AI
IMPACT Highlights ongoing challenges in balancing AI model helpfulness with safety guardrails and user control.
RANK_REASON User commentary on AI model behavior and safety training.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →