Findings
-
[1] 2026-09-24 b11163
llama: add llama_batch_ext (#24669) (wip) add llama_batch_ext wip updated design updated impl change signature unused var demo common_prompt_batch_decode fix pos tmp disable test-batch-alloc fix compat nits: add const no more pos_max add comment about llama_batch_ext_set_embd_state handle n_embd_out properly rename api –> embd_token llama_embd stub llama_batch_ext_set_embd_state support both token + embd + state in batch llama_batch_ext_add_embd upstream some changes nits fix…
-
[2] 2026-09-24 b11160
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952) vulkan: add int8 coopmat quantized matmul shader apply scales inline use scalar sums probe and directly access coopmat values instead of going through shmem add q8_0 support add BK_STEP to shader, default to 2 use larger workgroups double buffering preload scales coopmat load first, then wmma use float for…
-
[3] 2026-09-24 b11159
vulkan: handle misalignment in conv_2d and conv_3d (#29365) vulkan: handle misalignment in conv_2d and conv_3d fix test-backend-ops print Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49846182 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) -…
-
[4] 2026-09-24 b11158
vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328) Enable coopmat support for Vulkan backend Fixed the mul_mat_s Removed the debug statement Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49831613 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan)…
-
[5] 2026-09-24 b11157
cuda : add conv3d with implicit GEMM (#29137) cuda : add conv3d with implicit GEMM cuda : refine conv3d implicit GEMM and handle empty kernels Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49789544 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu…
-
[6] 2026-09-24 b11156
model : add Ling 3.0 VL support (#29151) model : fold Ling 3.0 VL into the BailingMoeV3 architecture Assisted-by: Scout model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections Co-authored-by: aetherbird aetherbird@users.noreply.github.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49780861 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu…
-
[7] 2026-09-24 b11155
server,common : fix the GCC 12 stringop-overread false positive (again) (#29325) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49775037 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries…
-
[8] 2026-09-24 b11154
test-save-load-state : print a per-model results table in –models mode (#29316) test-save-load-state : print a per-model results table in –models mode in –models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model…
-
[9] 2026-09-24 b11153
hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49757881 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA…
-
[10] 2026-09-23 v0.5.0
Overview This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP binding, image outputs from function… Changelog since v0.4.1 7fe450e llama.cpp : bump version to 0.5.0 (#29333) 177cd8c sync : ggml e4e2f62 ggml : bump version to 0.25.1 (ggml/1637) 66fba63 CUDA: add a reserve to avoid spurious warning on older GCC builds (#29317) bddf826 common :… (#28450) 9b421fa ui : Accept WEBM video files (#28622) 348f853 jinja: use const for statement::execute and ::visit (#29271) 217f81c server: Add support for binding to multiple addresses (#28690) 828fdf2 spec : support DFlash for HunyuanOCR (#28890) bfd73a8 convert: add MiMo-V2.6… ded MoE work in mul_mm coopmat1 path (#25483) 817e5f8 sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (#28929) c57da6f opencl: fix various warnings (#28984) 79bfc1d docs: remove JG as CODEOWNER for test-llama-archs (#29003) 05f2dcf vulkan: fix buffer_reference alignment…
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.