Findings
-
[1] 2026-09-23 b11147
opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49630180 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA…
-
[2] 2026-09-23 b11146
llama.cpp : bump version to 0.5.0 (#29333) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49623059 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries…
-
[3] 2026-09-23 b11140
CUDA: enable sparse-fa for dsv4 prefill (again) (#29298) CUDA: enable sparse-fa for dsv4 prefill (again) CUDA: unroll the query loop of the sparse mask scan The query loop of flash_attn_mask_to_sparse_indices has a runtime trip count, which keeps the unrolled scan over the values of a lane from issuing its loads together. Template the kernel on ncols1 so the loop is…
-
[4] 2026-09-23 b11139
server: fix token counting API crash on sleep (#29309) server: wake up sleeping server correctly server: wake up sleeping server correctly (local aliases removed) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49582694 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64…
-
[5] 2026-09-23 b11138
jinja : parse unary +/- before variables (#29244) jinja : parse unary +/- before variables Lexer already emits unary_operator for -n / +n, and runtime executes unary -. Parse them at multiplicative precedence so slices like items[:-n] and GigaChat indent[:-indent_factor] work. jinja : keep filters/tests outside unary operands Unary +/- must bind only the primary/postfix operand so -n|abs is (-n)|abs,…
-
[6] 2026-09-23 b11136
server: accept OpenAI video_url content type and data: video URIs (#27921) The OpenAI chat completions API specifies content part type "video_url" with a {"url": …} object, and clients typically send data: URIs (e.g. data:video/mp4;base64,…). The llama-server only accepted the non-standard "input_video" type and rejected data: URIs for video (accept_base64_uri=false), so any OpenAI-conformant client failed with "unsupported content[].type" or "Invalid uri…
-
[7] 2026-09-23 b11135
server: Dedup the draft HF model via dedup-cache-models (#27934) server: Dedup the draft HF model via dedup-cache-models Fixes #27846 server: avoid capturing structured binding in lambda Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49545978 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan)…
-
[8] 2026-09-23 b11132
model : support Gemma4 DSpark draft backbone (#29226) dspark: add Gemma 4 draft support Add GGUF conversion and runtime support for full-attention and SWA Gemma 4 DSpark drafts, including tied output weights and boolean backbone metadata. Assisted-by: Codex dflash: infer Gemma draft features from metadata Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49534734 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)…
-
[9] 2026-09-23 b11130
make-release : update summary prompt Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49525233 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries Ubuntu arm64…
-
[10] 2026-09-23 b11126
vulkan: add IQ4_XS MMQ/MMV matmul kernels (#28415) vulkan: optimize IQ4_XS matmul kernels Assisted-by: OpenAI Codex vulkan: address IQ4_XS review nits drop the dead LOAD_VEC_A != 8 branch in the IQ4_XS shmem load; iq4_xs is in lut_load_vec_a()'s "8" list, so that path is never generated disable MMVQ for IQ4_XS on Intel (27.3% tg regression on A770) remove a stray empty line…
-
[11] 2026-09-23 v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)
Signed-off-by: Andreas Karatzas akaratza@amd.com Co-authored-by: OpenAI Codex noreply@openai.com
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.