Findings
-
[1] 2026-09-26 b11201
Revert "Change max context length for auto-fitting with unified KV (#28849)" (#29437) This reverts commit b04d4e5. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50433369 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8…
-
[2] 2026-09-26 b11200
jinja : implement sameas test (#29448) implement sameas test add tests Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50397675 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13)…
-
[3] 2026-09-26 b11199
jinja : fix compile error (#29468) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50395469 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries Ubuntu…
-
[4] 2026-09-26 b11195
ggml-cpu: tiled mul_mat for k-quants (#27851) Added tiled mul_mat. For each mul_mat_one_chunk, quants are unpacked into (max) 256×256 tiles of int8, one routine per quent. Then microkernel computes 16×16 tiles before writing out 256×256 float reults to main memory. Tests/benches in tests/test-tiled-mulmat.cpp. 3-6x speed improvement for large matmul, break even at 4096×64 * 64×4096, 80% performance (net loss) for GEMV.…
-
[5] 2026-09-26 b11194
opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50382689 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA…
-
[6] 2026-09-26 b11193
hexagon: find software divide calls using binary inspection tool (#29449) hex-scripts: fix table alignment hex-scripts: find sw div calls using binary inspection tool Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50359675 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan)…
-
[7] 2026-09-26 b11192
vendor : update cpp-httplib to 0.58.0 (#29407) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50337762 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4 libraries…
-
[8] 2026-09-25 b11191
common,rpc : simplify fs_create_directory_with_parents() (#29432) The original function was broken on Windows for some unicode paths Paths without a trailing separator now create the last directory too, matching the function name. All current callers already include a trailing separator, so this change does not affect them. Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50291720 macOS/iOS: macOS Apple Silicon (arm64) macOS…
-
[9] 2026-09-25 b11190
mtmd: fix mel preprocessor in LFM2 audio (#29403) which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: uses log(x + 2^-24) instead of clamping to the log floor uses a symmetric Hann window, equivalent to torch.hann_window(periodic=False) adds the normalization epsilon to the standard deviation…
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.