A new position paper argues that claims about the stability of AI model explanations are scientifically invalid unless validated across multiple methods. Experiments with DenseNet201, ResNet50V2, and InceptionV3 showed that their stability rankings reversed depending on the attribution method used. The paper concludes that explanation stability is a property of the model-method pair, not the model alone, and calls for validation across multiple attribution methods in regulatory submissions to prevent illusory safety assurances. AI
IMPACT Highlights the need for robust validation of AI model explanations, potentially impacting how AI safety and reliability are assessed.
RANK_REASON This is a research paper published on arXiv discussing methodology for AI model explanations. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →