Noor S. Mohammad
PulseAugur coverage of Noor S. Mohammad — every cluster mentioning Noor S. Mohammad across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New research reveals how LLMs develop harmful intent, introducing 'Herald' moderator
Researchers have identified "Harmfulness Propagation Dynamics" (HPD), a phenomenon where the representation of harmful intent in large language models increases with the model's depth. This suggests that harmfulness is …
-
New MIRAGE benchmark reveals amplified anti-Muslim bias in LLMs
A new benchmark called MIRAGE has been developed to assess anti-Muslim bias in large language models, moving beyond simple prompt completion to evaluate reasoning, agentic decision-making, and time-coupled conditions. T…
-
New framework certifies faithfulness in AI-generated math proofs
Researchers have introduced Bidirectional Provability Fingerprinting (BPF), a new framework designed to certify the faithfulness of autoformalized mathematical statements. This method addresses the challenge where trans…