A user on r/LocalLLaMA tested the lexical knowledge and storytelling capabilities of various local AI models, specifically focusing on the game Chrono Trigger. The user developed a detailed prompt to assess how accurately models could recount the game's opening sequence up to the point where the main character, Crono, meets Marle. The evaluation involved checking for specific details like character names, locations, and plot points, while penalizing models for including information beyond the requested scope or for hallucinating details. AI
IMPACT Highlights the challenges local LLMs face with nuanced storytelling and factual recall, indicating areas for improvement in lexical knowledge and adherence to prompt constraints.
RANK_REASON User-generated benchmark and commentary on local LLM performance.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →