A new research paper examines the reproducibility of knowledge graph extraction from threat reports, highlighting inconsistencies in how systems match predicted triples to gold annotations. The study found that stated matching rules could only be reimplemented for a fraction of inspected systems, and re-scoring outputs under different protocols altered pairwise orderings. An LLM judge achieved higher agreement with human adjudication than mechanical matchers, and a new pipeline called CTIForge was developed to isolate component effects, revealing that validation layers can impact precision differently based on whether the backbone is hosted or offline. AI
IMPACT Highlights the need for standardized evaluation and validation in AI systems used for threat intelligence, impacting the reliability of security analysis.
RANK_REASON The cluster contains a research paper detailing an audit of knowledge graph extraction systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- Connected Papers
- CTIForge
- DagsHub
- Gotit.pub
- GRID
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →