A new 102M parameter model called Recursive BitNet N-Gram 102M has been released, featuring a combination of ternary weights, shared transformer layers, and hashed n-gram embeddings. This experimental model was trained from scratch on approximately 4.7 billion tokens, with a significant portion of its training focused on a 64K context window. While benchmarks show modest results, the model demonstrates an average score of 40.59% across several evaluations, with the developers open to feedback. AI
IMPACT This experimental model showcases novel techniques for efficient training and long context windows, potentially influencing future small-scale model development.
RANK_REASON Release of a new, albeit experimental, model with novel architectural features and benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →