11小时前Do small models actually pass your evals? A replication attemptOpenAI News · T1人工智能#llm#efficiency◆AI 生成69◀
11小时前Structured sparsity patterns that survive quantizationGoogle Research Blog · T1人工智能#efficiency#inference◆AI 生成55◀
5天前OpenAI cuts API inference prices again as serving costs fall below training for the first timeOpenAI News · T1模型与训练#llm#efficiency#inference◆AI 生成52◀
5天前Google details how it squeezed 40% more tokens per dollar out of serving fleetsGoogle Research Blog · T1模型与训练#llm#efficiency#inference◆AI 生成52◀
5天前Understanding KV cache reuse across conversational workloadsOpenAI News · T1人工智能#llm#efficiency◆AI 生成52◀