Skip to content Skip to footer

Open-Source Models & Communities — September 23, 2026

Findings

  1. [1] 2026-09-23 b11147

    opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49630180 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA…

  2. [2] 2026-09-23 b11146

    llama.cpp : bump version to 0.5.0 (#29333) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49623059 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries…

  3. [3] 2026-09-23 b11140

    CUDA: enable sparse-fa for dsv4 prefill (again) (#29298) CUDA: enable sparse-fa for dsv4 prefill (again) CUDA: unroll the query loop of the sparse mask scan The query loop of flash_attn_mask_to_sparse_indices has a runtime trip count, which keeps the unrolled scan over the values of a lane from issuing its loads together. Template the kernel on ncols1 so the loop is…

  4. [4] 2026-09-23 b11139

    server: fix token counting API crash on sleep (#29309) server: wake up sleeping server correctly server: wake up sleeping server correctly (local aliases removed) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49582694 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64…

  5. [5] 2026-09-23 b11138

    jinja : parse unary +/- before variables (#29244) jinja : parse unary +/- before variables Lexer already emits unary_operator for -n / +n, and runtime executes unary -. Parse them at multiplicative precedence so slices like items[:-n] and GigaChat indent[:-indent_factor] work. jinja : keep filters/tests outside unary operands Unary +/- must bind only the primary/postfix operand so -n|abs is (-n)|abs,…

  6. [6] 2026-09-23 b11136

    server: accept OpenAI video_url content type and data: video URIs (#27921) The OpenAI chat completions API specifies content part type "video_url" with a {"url": …} object, and clients typically send data: URIs (e.g. data:video/mp4;base64,…). The llama-server only accepted the non-standard "input_video" type and rejected data: URIs for video (accept_base64_uri=false), so any OpenAI-conformant client failed with "unsupported content[].type" or "Invalid uri…

  7. [7] 2026-09-23 b11135

    server: Dedup the draft HF model via dedup-cache-models (#27934) server: Dedup the draft HF model via dedup-cache-models Fixes #27846 server: avoid capturing structured binding in lambda Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49545978 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan)…

  8. [8] 2026-09-23 b11132

    model : support Gemma4 DSpark draft backbone (#29226) dspark: add Gemma 4 draft support Add GGUF conversion and runtime support for full-attention and SWA Gemma 4 DSpark drafts, including tied output weights and boolean backbone metadata. Assisted-by: Codex dflash: infer Gemma draft features from metadata Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49534734 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)…

  9. [9] 2026-09-23 b11130

    make-release : update summary prompt Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49525233 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries Ubuntu arm64…

  10. [10] 2026-09-23 b11126

    vulkan: add IQ4_XS MMQ/MMV matmul kernels (#28415) vulkan: optimize IQ4_XS matmul kernels Assisted-by: OpenAI Codex vulkan: address IQ4_XS review nits drop the dead LOAD_VEC_A != 8 branch in the IQ4_XS shmem load; iq4_xs is in lut_load_vec_a()'s "8" list, so that path is never generated disable MMVQ for IQ4_XS on Intel (27.3% tg regression on A770) remove a stray empty line…

  11. [11] 2026-09-23 v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)

    Signed-off-by: Andreas Karatzas akaratza@amd.com Co-authored-by: OpenAI Codex noreply@openai.com

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.

Sources

  1. [1] b11147
  2. [2] b11146
  3. [3] b11140
  4. [4] b11139
  5. [5] b11138
  6. [6] b11136
  7. [7] b11135
  8. [8] b11132
  9. [9] b11130
  10. [10] b11126
  11. [11] v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)

Leave a comment

0.0/5