Skip to content Skip to footer

Open-Source Models & Communities — September 28, 2026

Findings

  1. [1] 2026-09-28 b11236

    batch: migrate speculative, mtmd and server to batch_ext (#29385) adapt common add common_batch wip wip: spec cont common_speculative_process server_batch to use common_batch rm some stale calls Assisted-by: Claude Fable 5.1 migrate mtmd handle imrope, handle return val of add()/add_embd() add spec zeros vector add warning on zero fill path Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50870264 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…

  2. [2] 2026-09-28 b11235

    common : fix HF cache paths on Windows (#29475) Supersedes #29158 Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50839831 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries…

  3. [3] 2026-09-28 b11234

    webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (#29471) Fix: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor Clang formatting Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50830975 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries…

  4. [4] 2026-09-28 b11233

    tests : refactor test-recurrent-state-rollback (#29426) tests : use llama_context_ptr in test-recurrent-state-rollback Replace raw llama_context pointers with llama_context_ptr and drop the manual llama_free calls and cleanup lambda. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL tests : run test-recurrent-state-rollback over all dummy models Add a –models DIR mode that mirrors test-save-load-state: iterate every dummy model, report PASS/FAIL/SKIP in a table and fail only when a model fails.…

  5. [5] 2026-09-28 b11232

    ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (#29423) ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail ggml-cpu: fix FA softcap handling for padded KV tiles Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50796641 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…

  6. [6] 2026-09-28 b11229

    HIP: fix template skip for DKQ > 256 mfma kernels (#29559) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50767985 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13)…

  7. [7] 2026-09-28 b11228

    metal: support left and circular padding in GGML_OP_PAD (#29561) metal: support left and circular padding in GGML_OP_PAD Align Metal with CPU, CUDA and Vulkan: shift the source coordinates by the left paddings, wrap them around with the same wrap_around when circular, and read the source through nb00, which also fixes a right padding of a permuted source. A test case…

  8. [8] 2026-09-28 Holo4: powering generalist computer-use agents

  9. [9] 2026-09-28 b11227

    context : do not re-reserve the scheduler when toggling causal_attn (#28751) context : do not re-reserve the scheduler when toggling causal_attn llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive sched_reserve() passes per image.…

  10. [10] 2026-09-28 b11226

    Enables Windows ARM64 build with MSVC cl.exe (#28362) can reproduce the issue vlad sees fix fma issue drop volatile fix volatile runtime task add arm flag if needed fix hsum compile error fix syntax in quants strengthen sve probing make the syntax fixes one liners remove debug code formatting remove macro for float drive down gcc instruction count support armec…

  11. [11] 2026-09-28 b11225

    tests : fix ggml init (#29554) tests : init ggml for test-recurrent-state-rollback cont : same for test-save-load-state cont : add to test-state-restore-fragmented + add TODOs Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50700507 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu…

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.

Sources

  1. [1] b11236
  2. [2] b11235
  3. [3] b11234
  4. [4] b11233
  5. [5] b11232
  6. [6] b11229
  7. [7] b11228
  8. [8] Holo4: powering generalist computer-use agents
  9. [9] b11227
  10. [10] b11226
  11. [11] b11225

Leave a comment

0.0/5