Findings [1] 2026-09-24 b11163 llama: add llama_batch_ext (#24669) (wip) add llama_batch_ext wip updated design updated impl change signature unused var demo common_prompt_batch_decode fix pos tmp disable test-batch-alloc fix compat nits: add const no more pos_max add comment about llama_batch_ext_set_embd_state handle n_embd_out properly rename api --> embd_token llama_embd stub llama_batch_ext_set_embd_state support both token + embd…
Findings [1] 2026-09-23 b11147 opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49630180 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12)…
Findings [1] 2026-09-22 b11112 server: support input_image in function_call_output (#20663) (#22575) server: support input_image in function_call_output (#20663) server: fix if statement spacing server: avoid repeated type lookup Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49330467 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64…
What Happened
Over the last update cycle ggml/llama.cpp received a set of operational, portability and performance changes that materially affect how open weights and inference engines are deployed in production:
llama-server gained environment‑variable control via new LLAMA_ARG_* mappings (e.g., LLAMA_ARG_TEMP, LLAMA_ARG_TOP_P, LLAMA_ARG_REPEAT_PENALTY), enabling systemd/EnvironmentFile driven configuration for runtime sampling parameters; documentation was regenerated…
What Happened
Kernel specialization for variable HC: A metal backend change allows dsv4_hc_pre kernels to accept arbitrary hc values (previously hardcoded to 4). The patch passes n_hc as a function constant, generates per-n_hc pipeline variants and adds tests for many hc values — fixing a fallback-to-CPU failure for models that vary hc by…
What Happened
Over the last development cycle the open-source inference ecosystem delivered two parallel waves of work: (1) broad, low-level backend and kernel hardening in the llama.cpp/ggml ecosystem that expands supported accelerators and fixes stability/performance bugs, and (2) a major inference/runtime release that adds schedulers, memory/residency improvements, new models and language/runtime tooling for high‑throughput production…
What Happened
Over the last set of commits to ggml / llama.cpp the community shipped multiple engineering changes that materially affect inference performance, portability and robustness for open‑weight models. Key changes include:
New binary matmul kernels (including A8 Q6_K non‑MoE and IQ3_S MMQ kernels) and layout fixes that broaden high‑performance kernel coverage across…
What Happened
Community maintainers of ggml/llama.cpp shipped a dense set of fixes and platform expansions focused on robustness, backend coverage and model-format correctness. Changes include:
Expanded CI/build matrix and attestations for multi‑platform binary builds (macOS/iOS, Linux x64/arm64/s390x, Android, Windows, openEuler) and many GPU/backends (Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL, OpenCL) [1][2][4][6].
…
What Happened
The llama.cpp community pushed multiple engineering and model-conversion changes that affect production inference stacks: kernel and backend improvements, new quant formats and Hexagon support, a novel causal-only HRM model conversion, GPU/CUDA optimizations, and a high‑severity RPC use‑after‑free fix.
Fixed fused QKV split-state for gemma4 and added fused full-attention handling for Qwen35…
What Happened
Over the last development cycle the ggml/llama.cpp ecosystem received a series of coordinated engineering changes that materially affect running open weights and community inference tooling:
MoE, OpenCL and GEMM fixes: OpenCL changes select MoE expert matmuls by batch size, gate the prebuilt q4_0 MoE GEMM on routing count, and stop writing…
What Happened
The llama.cpp project published a major v0.4.1 milestone and associated commits that together expand model support, backend coverage, server reliability and observability while hardening correctness across CPU/GPU/accelerator backends.
New models & formats: adds Maple 20B‑A1B (ternary MoE, CPU) and a Tencent Hy 4 preview; conversion and format flags (e.g., --fuse-qkv) to…
What Happened
Over the last set of upstream changes, the ggml/llama.cpp ecosystem pushed multiple engineering fixes and tooling improvements that matter for production inference deployments. Key items:
Expanded and hardened multi‑platform CI/build matrix (macOS/iOS, Linux x64/arm64/s390x, Android arm64, Windows, and openEuler variants) with many GPU/backends covered (CUDA 12/13, Vulkan, ROCm 10.0, OpenVINO, SYCL…