A new research paper explores the effectiveness of an agentic harness, which combines a 27B language model with financial calculations and validation checks, for generating presentations in corporate and investment banking. The study found that when evaluated by human judges, the system consistently scored higher on development deliverables compared to direct generation from a short prompt. However, the paper notes that the use of LLM judges to guide engineering changes raises questions about whether higher scores reflect genuine document improvement or shifts in grading criteria. The research also observed variability in judge agreement on final rankings and noted that repeated grading could alter scores on unchanged decks, complicating the assessment of small improvements. AI
IMPACT This research suggests potential for AI to streamline complex financial document generation, though further validation is needed to distinguish genuine improvement from judge bias.
RANK_REASON Research paper published on arXiv detailing an AI system's performance. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →