llama.cpp Releases· github-actions[bot]·AI 评分22
llama.cpp b11372 发布:qwen4exp 索引器显存减半并优化多后端
b11372
AI 导读
llama.cpp 发布 b11372,qwen4exp 将 lightning indexer 的索引器分数内存减半,改为每个 head 单独计算并原地累加进单个 [n_pool, n_tokens] 分数张量。CUDA 新增支持 4 heads 的向量内核,Metal 将 head 数作为函数常量,Vulkan 按 keys×tokens 分块并改用 fp16 点积。
来源:llama.cpp Releases · github.com