Skip to content Skip to footer

Open-Source Models & Communities — September 25, 2026

Findings

  1. [1] 2026-09-25 b11189

    opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (#29401) opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel opencl: fix s transpose – s only transposed for bin kernels Co-authored-by: Li He lih@qti.qualcomm.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50281199 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64…

  2. [2] 2026-09-25 b11188

    Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (#29373) (#29409) vulkan : fix build issue of legacy glslc version by adding GGML_VULKAN_COOPMAT_GLSLC_SUPPORT macro check for Intel FA shader compiling vulkan : add preprocess condition to filter out unsupported FA 2 phases kernels before creation. vulkan : move lock_guard for Intel FA shader pointer…

  3. [3] 2026-09-25 b11185

    common : extract shared unicode path/string helpers (#29415) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50264913 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA…

  4. [4] 2026-09-25 b11184

    metal: FWHT kernels for block widths above 512 (#29095) metal: FWHT kernels for block widths above 512 The Metal FWHT covers widths 64 to 512, one row per simdgroup with N/32 values per lane. Wider blocks need more registers per lane than that layout allows. kernel_fwht_tg runs one row per threadgroup with 256 threads, so each thread keeps N/256 values.…

  5. [5] 2026-09-25 b11183

    metal : split fa kernels into per-dtype libraries (#29329) metal : split fa kernels into per-dtype libraries Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp cont : minor fix comment Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50249107 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64…

  6. [6] 2026-09-25 b11182

    llama : add llama_prec_policy + model-driven W4A4 path (#24364) Rebase and update based on #26675 Signed-off-by: ynankani ynankani@nvidia.com CI failure fix(launh_bounds overload on HIP) and cleanup Signed-off-by: ynankani ynankani@nvidia.com Address review comments Signed-off-by: ynankani ynankani@nvidia.com Use ggml tensor instead of name in act policy map Signed-off-by: ynankani ynankani@nvidia.com Address review comments and cleanup Signed-off-by: ynankani ynankani@nvidia.com Address review comments Signed-off-by:…

  7. [7] 2026-09-25 b11181

    HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (#29231) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50216799 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu…

  8. [8] 2026-09-25 b11180

    rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (#29283) rpc : include nb in the get_alloc_size cache key and floor the result at ggml_nbytes cont : remove redundant comment cont : add TODO Co-authored-by: Georgi Gerganov ggerganov@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50207987 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS…

  9. [9] 2026-09-25 b11179

    [SYCL] support sparse FA (#28796) fix conflict fix format issue rm unused code Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50188435 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64…

  10. [10] 2026-09-25 b11178

    musa: fix PH1 (MTT S5000) operator failures and build issues (#29193) musa: use 16-byte copies for MUSA like sm_70+ ggml_cuda_get_max_cpy_bytes() derives the copy width from CUDA_ARCH. mcc never defines it, so MUSA fell into the generic branch and returned 8… topk_moe_cuda returns early for the rows past the end of the graph, but one block covers TOPK_MOE_ROWS_PER_BLOCK (8) rows, so the last block is only partially filled whenever n_rows is not a multiple of 8. On MUSA a warp that…

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.

Sources

  1. [1] b11189
  2. [2] b11188
  3. [3] b11185
  4. [4] b11184
  5. [5] b11183
  6. [6] b11182
  7. [7] b11181
  8. [8] b11180
  9. [9] b11179
  10. [10] b11178

Leave a comment

0.0/5