PulseAugur
实时 14:35:16
English(EN) InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

新基准InvestPhilBench测试LLM在投资理念中的程序化推理能力

研究人员推出了InvestPhilBench,这是一个旨在评估大型语言模型(LLM)在专家投资理念领域程序化推理能力的新基准。该基准的v0.6版本包括了经过验证的投资原则卡、决策框架卡和问答题,以及一个自动评分管道(BASP)和故障模式检测协议。对四种模型的初步测试显示,前沿模型与其他模型之间存在显著的性能差距,综合得分在前沿模型上趋于饱和,但门重构准确率(GRA)等特定指标仍表明存在程序化缺陷。 AI

影响 该基准旨在通过专门测试LLM的程序化推理能力来提高其在金融分析中的可靠性。

排序理由 该集群描述了一个用于评估LLM的新学术基准和方法的发布。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新基准InvestPhilBench测试LLM在投资理念中的程序化推理能力

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Mingguang Chen, Bo Qu ·

    InvestPhilBench:一个多层动态基准,用于评估大型语言模型在专家投资理念中的程序推理能力

    arXiv:2606.25984v1 Announce Type: cross Abstract: Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision frameworks of expert investors. We introd…

  2. arXiv cs.AI TIER_1 English(EN) · Bo Qu ·

    InvestPhilBench:一个多层动态基准,用于评估大型语言模型在专家投资理念中的程序推理能力

    Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision frameworks of expert investors. We introduce InvestPhilBench, a multi-layer dynamic benchma…