Findings
-
[1] 2026-09-27 b11222
common : avoid side effects around params parsing (#29537) register –rpc unconditionally and call llama_supports_rpc() only from its handler print server "initialization …" log after args are parsed Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50580981 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x…
-
[2] 2026-09-27 b11221
common : make string_split throw on invalid values (#29518) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50573055 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64…
-
[3] 2026-09-27 b11218
jinja : add support for dict builtin (#29477) add support for dict builtin add tests Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50563209 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries…
-
[4] 2026-09-27 b11217
opencl: refine bin kernel loading condition (#29503) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50560500 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries…
-
[5] 2026-09-27 b11216
sycl: FWHT kernels for block widths above 512 (#29243) The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2)…
-
[6] 2026-09-27 b11215
CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#26289) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50554486 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13)…
-
[7] 2026-09-27 b11214
HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (#28907) HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes CI: hip-quality-check: ignore spills for very large mfma mma kernels Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50549501 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
-
[8] 2026-09-27 b11213
vulkan: fix argsort kernel selection for Adreno (#29469) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50544436 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4…
-
[9] 2026-09-27 b11212
common : throw instead of abort on grammar without llguidance (#29516) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50541746 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries…
-
[10] 2026-09-27 b11211
RPC: use RDMA completion channel to not spin (#29440) RPC: use RDMA completion queue to not spin add TODO for apple RDMA Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50534798 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu…
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.