A demonstration showcases voice conversations between two models, Gemma4 12B and E2B, running on different hardware. Gemma 4 12B operates on an RTX PRO 4500 Blackwell, while Gemma 4 E2B utilizes a Jetson Orin NX 16GB, with similar performance anticipated on a Jetson Orin Nano Super 8GB. Both setups employ a reSpeaker microphone array and a 3W speaker, with inference managed by Cortexist Little Gemma, a C-based LLM engine for CUDA devices that reportedly outperforms llama.cpp on the Jetson Orin and supports features like lip sync, expressions, and gestures. AI
IMPACT Shows potential for running advanced conversational AI on diverse hardware, including edge devices.
RANK_REASON Demonstration of LLM models running on consumer and embedded hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →