Tokenless, a new startup backed by Y Combinator, has launched a service designed to reduce AI inference costs. The tool acts as a drop-in replacement for API calls, intelligently routing requests to the most cost-effective model while maintaining quality. By analyzing model performance in real-time, Tokenless cancels redundant model executions, potentially halving inference expenses. The service offers an OpenAI and Anthropic compatible endpoint and is built by researchers from Google DeepMind, Princeton, and UC Berkeley. AI
IMPACT Could significantly lower operational costs for AI applications by optimizing model selection.
RANK_REASON This is a product launch for a tool that integrates with existing AI models, rather than a core AI release from a frontier lab.
Read on HN — claude cli stories →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →