A survey of 49 operational sources reveals a significant gap in evidence supporting the deployment of Large Language Models (LLMs) in fraud detection and trust-and-safety workflows. While LLMs are increasingly proposed for these tasks, much of the existing literature focuses on model performance rather than practical operational constraints like latency, cost, and adversarial risk. The survey, which coded sources on fraud detection, investigation support, and content moderation, found that fraud-related papers often report offline task performance instead of crucial per-decision metrics. To address this, the research introduces FORTE, a framework for organizing LLM roles in these workflows, and a minimum deployment-evidence checklist to guide future research needed for robust LLM integration. AI
IMPACT Highlights critical evidence gaps for deploying LLMs in sensitive operational workflows, guiding future research toward practical deployment considerations.
RANK_REASON The item is a survey paper analyzing existing research and identifying evidence gaps for LLM deployment. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →