Researchers have introduced FirmCORe, a new benchmark designed to evaluate the ability of large language models (LLMs) to identify and reason about collaboration opportunities between companies. The benchmark consists of 2,805 labeled firm pairs, assessing not only the detection of potential collaborations but also their strength, type, and role direction. Initial experiments with various LLMs demonstrated that while models can detect opportunities with a 74.51% macro-F1 score, their performance drops significantly when identifying specific collaboration details, achieving only a 61.57% exact match across all output fields. The benchmark also includes parallel Chinese and English evaluation sets to analyze language sensitivity. AI
IMPACT This benchmark could drive improvements in LLMs' ability to understand complex business relationships and facilitate strategic partnerships.
RANK_REASON The item describes a new academic benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- English
- FirmCORe
- Gotit.pub
- Hugging Face
- Influence Flower
- large language models
- ScienceCast
- Standard Chinese
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →