Researchers have developed a new framework called Targeted Counterfactual Fingerprinting (TCF) to verify ownership of large language models (LLMs) when they are only accessible via black-box APIs. TCF converts open-ended generation comparisons into constrained-answer targeted counterfactual transfers, reducing ambiguity in verification scores. The method optimizes prompt perturbations to achieve a counterfactual target and uses a protected-model-only quantity called the source-model counterfactual margin (SCM) to certify the target's likelihood. Across four LLM families, TCF demonstrated an average AUC of 0.9861, outperforming existing methods like TRAP, ProFLingo, and ZeroPrint. AI
IMPACT This research offers a more robust method for verifying LLM ownership in black-box scenarios, potentially impacting intellectual property protection and model security.
RANK_REASON The cluster contains a research paper detailing a new method for LLM ownership verification. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- LLMs
- ProFLingo
- source-model counterfactual margin
- Targeted Counterfactual Fingerprinting
- ZeroPrint
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →