ZAYA1-8B: An 8B Moe Model with 760M Active Params Matching DeepSeek-R1 on Math

(firethering.com)

22 points | by steveharing1 3 hours ago ago

20 comments