Skip to content Skip to footer

Open-Source Models & Communities — September 24, 2026

Findings

  1. [1] 2026-09-24 b11163

    llama: add llama_batch_ext (#24669) (wip) add llama_batch_ext wip updated design updated impl change signature unused var demo common_prompt_batch_decode fix pos tmp disable test-batch-alloc fix compat nits: add const no more pos_max add comment about llama_batch_ext_set_embd_state handle n_embd_out properly rename api –> embd_token llama_embd stub llama_batch_ext_set_embd_state support both token + embd + state in batch llama_batch_ext_add_embd upstream some changes nits fix…

  2. [2] 2026-09-24 b11160

    vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952) vulkan: add int8 coopmat quantized matmul shader apply scales inline use scalar sums probe and directly access coopmat values instead of going through shmem add q8_0 support add BK_STEP to shader, default to 2 use larger workgroups double buffering preload scales coopmat load first, then wmma use float for…

  3. [3] 2026-09-24 b11159

    vulkan: handle misalignment in conv_2d and conv_3d (#29365) vulkan: handle misalignment in conv_2d and conv_3d fix test-backend-ops print Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49846182 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) -…

  4. [4] 2026-09-24 b11158

    vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328) Enable coopmat support for Vulkan backend Fixed the mul_mat_s Removed the debug statement Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49831613 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan)…

  5. [5] 2026-09-24 b11157

    cuda : add conv3d with implicit GEMM (#29137) cuda : add conv3d with implicit GEMM cuda : refine conv3d implicit GEMM and handle empty kernels Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49789544 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu…

  6. [6] 2026-09-24 b11156

    model : add Ling 3.0 VL support (#29151) model : fold Ling 3.0 VL into the BailingMoeV3 architecture Assisted-by: Scout model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections Co-authored-by: aetherbird aetherbird@users.noreply.github.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49780861 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu…

  7. [7] 2026-09-24 b11155

    server,common : fix the GCC 12 stringop-overread false positive (again) (#29325) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49775037 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries…

  8. [8] 2026-09-24 b11154

    test-save-load-state : print a per-model results table in –models mode (#29316) test-save-load-state : print a per-model results table in –models mode in –models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model…

  9. [9] 2026-09-24 b11153

    hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49757881 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA…

  10. [10] 2026-09-23 v0.5.0

    Overview This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP binding, image outputs from function… Changelog since v0.4.1 7fe450e llama.cpp : bump version to 0.5.0 (#29333) 177cd8c sync : ggml e4e2f62 ggml : bump version to 0.25.1 (ggml/1637) 66fba63 CUDA: add a reserve to avoid spurious warning on older GCC builds (#29317) bddf826 common :… (#28450) 9b421fa ui : Accept WEBM video files (#28622) 348f853 jinja: use const for statement::execute and ::visit (#29271) 217f81c server: Add support for binding to multiple addresses (#28690) 828fdf2 spec : support DFlash for HunyuanOCR (#28890) bfd73a8 convert: add MiMo-V2.6… ded MoE work in mul_mm coopmat1 path (#25483) 817e5f8 sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (#28929) c57da6f opencl: fix various warnings (#28984) 79bfc1d docs: remove JG as CODEOWNER for test-llama-archs (#29003) 05f2dcf vulkan: fix buffer_reference alignment…

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.

Sources

  1. [1] b11163
  2. [2] b11160
  3. [3] b11159
  4. [4] b11158
  5. [5] b11157
  6. [6] b11156
  7. [7] b11155
  8. [8] b11154
  9. [9] b11153
  10. [10] v0.5.0

Leave a comment

0.0/5