A user benchmarked two "System One" decision models, Jev and Laya, for an LLM router. Initially, Jev was usable out-of-the-box, while Laya required fine-tuning. After fine-tuning Laya on custom labels, its performance improved significantly, matching Jev in complexity and risk assessment, though it still struggled with confidence. Further testing on real-world home automation decisions showed Laya's stock version provided overly simplistic answers. AI
IMPACT Provides insights into the practical performance and fine-tuning needs of specific LLM decision models for routing tasks.
RANK_REASON User benchmark and fine-tuning of specific LLM decision models.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →