跳到正文
原文
llama.cpp Releases· github-actions[bot]·AI 评分22

llama.cpp b11372 发布:qwen4exp 索引器显存减半并优化多后端

b11372

AI 导读

llama.cpp 发布 b11372,qwen4exp 将 lightning indexer 的索引器分数内存减半,改为每个 head 单独计算并原地累加进单个 [n_pool, n_tokens] 分数张量。CUDA 新增支持 4 heads 的向量内核,Metal 将 head 数作为函数常量,Vulkan 按 keys×tokens 分块并改用 fp16 点积。

来源:llama.cpp Releases · github.com