Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

vulkan: extend topk_moe fusion to support sqrt(softplus) ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#26124 opened Jul 25, 2026 by jeffbolznv Contributor Loading…
common: Add CLI > ENV > models-presets > INI precedence
#26118 opened Jul 25, 2026 by jcmdln Loading…
Fix speculative models failing to load when running llama-server
#26114 opened Jul 25, 2026 by sheldonrobinson Contributor Loading…
hexagon: enable quantized matmul for multi-sequence inputs ggml changes relating to the ggml tensor library for machine learning Hexagon
#26113 opened Jul 25, 2026 by w1049 Contributor Loading…
CUDA : add warp-per-row WKV7 kernel for single-token decode CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26111 opened Jul 25, 2026 by 123123213weqw Loading…
Feature/e2k support ggml changes relating to the ggml tensor library for machine learning
#26107 opened Jul 25, 2026 by png-tech Loading…
sycl: fix classification of iGPUs ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26105 opened Jul 25, 2026 by KyleHagy Contributor Loading…
ggml-cpu: skip unsupported ARM ISA variants under GGML_CPU_ALL_VARIANTS ggml changes relating to the ggml tensor library for machine learning
#26103 opened Jul 25, 2026 by ninihuang2026 Loading…
common: add subproc.h wrapper, disabled on android/ios build Compilation issues mtmd Related to multimodal functionality (video/image/audio) server testing Everything test related
#26102 opened Jul 24, 2026 by ngxson Collaborator Loading…
ui: stabilize the rendering of tool invocations server/ui
#26098 opened Jul 24, 2026 by zachwinter Contributor Loading…
ui: rendering performance follow-up server/ui
#26097 opened Jul 24, 2026 by allozaur Contributor Loading…
docs: use ROCM_PATH instead of HIP_PATH in linux HIP build command (#26060) documentation Improvements or additions to documentation
#26096 opened Jul 24, 2026 by amd-hasnasir Loading…
Add LlamaNet to the list of tools in the README. documentation Improvements or additions to documentation
#26095 opened Jul 24, 2026 by unixguru2k Loading…
opencl: fix fused RMS norm mul view offset ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend
#26085 opened Jul 24, 2026 by happyyzy Contributor Loading…
metal: fix memory leak if model is freed without any GPU operations Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26082 opened Jul 24, 2026 by nikwen Contributor Loading…
llama: add default load-mode auto, which avoids mmap on iGPUs AMD ZenDNN Issues related to the AMD ZenDNN backend Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning Hexagon IBM zDNN issues specific to IBM zDNN Accelerator OpenCL Issues specific to the OpenCL backend OpenVINO SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language Vulkan Issues specific to the Vulkan backend WebGPU
#26081 opened Jul 24, 2026 by 0cc4m Contributor Loading…
CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26079 opened Jul 24, 2026 by praneshgo Contributor Draft
kleidiai: Update KleidiAI Documentation documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning
#26078 opened Jul 24, 2026 by JonathanC-ARM Draft
kleidiai: Rework KleidiAI Build System/Integration ggml changes relating to the ggml tensor library for machine learning
#26077 opened Jul 24, 2026 by JonathanC-ARM Draft
kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai ggml changes relating to the ggml tensor library for machine learning
#26076 opened Jul 24, 2026 by JonathanC-ARM Loading…
ggml : handle graph buffer reservation failure ggml changes relating to the ggml tensor library for machine learning
#26070 opened Jul 24, 2026 by FaiChou Loading…
ggml-cpu: Enable tiled gemm for BF16 and FP16 ggml changes relating to the ggml tensor library for machine learning
#26068 opened Jul 24, 2026 by shalinib-ibm Contributor Loading…
ProTip! Updated in the last three days: updated:>2026-07-22.