Simon Willison· @simonw · X·· 8 小时前AI 评分18
AI 导读
在不到60GB内存里,最适合编程的开源权重Mixture-of-Experts LLM是哪个? 我觉得在我能用的硬件上,要想达到还算可交互的速度,MoE可能是必需的——我想要比12 tokens/秒更快的东西
正文
What's the best open weight Mixture-of-Experts LLM for coding that fits in less than 60GB of RAM?
I think MoE might be necessary to get reasonably interactive speeds on the hardware I have access to - I want something faster than 12 tokens/second
来源:Simon Willison · x.com