Findings
-
[1] 2026-09-29 b11259
common : stop accepting draft tokens at EOG (#29638) common : stop accepting draft tokens at EOG cont : remove the test Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51199980 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu…
-
[2] 2026-09-29 b11258
server : remove the built-in UI's service worker when the UI is not served (#29565) With –path or –no-ui, /sw.js returned 404, and a 404 does not remove a service worker, so browsers kept showing the cached built-in UI. Serve a worker that unregisters itself, clears its caches and reloads open tabs. A sw.js in the –path folder is still…
-
[3] 2026-09-29 b11257
common : use fs::path for config dir (#29649) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51176378 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA…
-
[4] 2026-09-29 b11256
tests : adjust server string regex to also match m2 utlra results (#29648) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51122498 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64…
-
[5] 2026-09-29 b11255
common : add fs_write_atomic() (#29642) Check for buffered write errors when closing downloaded files. Use UTF-8 paths when writing ETag files on Windows. Write in binary mode on Windows. Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51111496 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64…
-
[6] 2026-09-29 b11254
ggml : collect all input tensors into graph_inputs (#29634) graph_inputs was populated while splitting the graph, so it only contained the inputs that are used as srcs of some node. With pipeline parallelism (n_copies > 1), each graph input contributes n_copies leafs to graph_copy, so switching between batches that consume different inputs (e.g. token batches that do not use the…
-
[7] 2026-09-29 b11249
llama : fix init in several tools/examples (#29632) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51066139 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) – CUDA 12.8 libraries Ubuntu x64 (CUDA 13) – CUDA 13.4…
-
[8] 2026-09-29 b11247
vulkan : reuse descriptor sets when bindings are constant (#29280) vulkan : reuse descriptor sets when bindings are constant vulkan : bump buffer_destroy_count before destroying the buffer Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51033526 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64…
-
[9] 2026-09-29 b11246
chat : fix Muse Glimmer ignoring response_format json_schema with –jinja (#29615) chat : fix Muse Glimmer ignoring response_format json_schema with –jinja Fixes #29613 chat : accept json fences and clean up chat : fix choice parenthesis Co-authored-by: Alde Rojas hello@alde.dev Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51023533 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
-
[10] 2026-09-29 b11245
common : use fs::path for cache dirs (#29595) Avoid useless string conversions on Windows. No need for BSD or emscripten special cases. Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51016898 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan)…
-
[11] 2026-09-29 v0.31.0rc1: [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)
Signed-off-by: khluu khluu000@gmail.com Co-authored-by: Claude Opus 5.5 noreply@anthropic.com (cherry picked from commit aedaba8)
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our Machine Learning Development practice, or Scope an ML Project.