What Happened
The developments in these notes are concentrated in llama.cpp and ggml, not new open-weight model releases. llama.cpp added support for the pplx-decider model, while text, vision and audio support for embeddinggemma2 is tracked as a request rather than a confirmed release. [5] [8]
Inference changes include RPC support for tensor split mode, with…
What Happened
Two major open-source inference releases expand the choices for running models outside a managed API. llama.cpp v0.6.0 adds model and multimodal support, new processing and precision APIs, and performance improvements across several hardware backends. Its server now accepts typed vision, audio, and video inputs for embeddings and reports model input and output modalities…
What Happened
The documented activity centers on llama.cpp rather than new model weights. Recent changes include a fix for a chat tool-call parser use-after-free and double-free, a fix for a CUDA mixture-of-experts memory fault when expert count greatly exceeds microbatch size, improved Vulkan matrix-vector tuning for RDNA4 GPUs, and vectorized BF16, FP16, and FP32 tinyBLAS…
What Happened
Recent open-source activity in these notes centers on llama.cpp and its inference backends, not on a new open-weight model release. The changes add text-only support for the Clef decision model and fused MoE support for gemma-4, while improving Qwen3.5 MoE execution through the OpenVINO backend [8][6]. No new Hugging Face, Ollama, vLLM, SGLang,…
What Happened
SGLang 0.5.21 expands support for open-weight and other models, including DeepSeek-V4.1 Flash and several vision and diffusion models. It also makes its Rust radix-tree cache core the default, improves prefill/decode serving and KV-cache handling, and adds classification and candidate-scoring APIs. Its reported 22% improvement in first-token time for DeepSeek-V4.1 on long prompts and…
What Happened
Recent llama.cpp releases focus on inference reliability and model compatibility, rather than announcing new model weights. A direct-I/O change avoids making a second full-size copy of each tensor during memory mapping, while another fixes a workqueue race that could leave read and write state out of sync. Backend changes address a ROCm hardware-detection…
What Happened
Over the last set of community updates the llama.cpp / ggml ecosystem and adjacent tooling have added broad platform support, new numeric formats, and robustness fixes while vllm and Ollama pushed complementary standards and runtime features:
Numerical and precision work: BF16 support was expanded in ggml for unary, GLU, binary and…
Findings [1] 2026-09-29 b11259 common : stop accepting draft tokens at EOG (#29638) common : stop accepting draft tokens at EOG cont : remove the test Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/51199980 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU)…
Findings [1] 2026-09-28 b11236 batch: migrate speculative, mtmd and server to batch_ext (#29385) adapt common add common_batch wip wip: spec cont common_speculative_process server_batch to use common_batch rm some stale calls Assisted-by: Claude Fable 5.1 migrate mtmd handle imrope, handle return val of add()/add_embd() add spec zeros vector add warning on zero fill path Website:…
Findings [1] 2026-09-27 b11222 common : avoid side effects around params parsing (#29537) register --rpc unconditionally and call llama_supports_rpc() only from its handler print server "initialization ..." log after args are parsed Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50580981 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
Findings [1] 2026-09-26 b11201 Revert "Change max context length for auto-fitting with unified KV (#28849)" (#29437) This reverts commit b04d4e5. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50433369 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan)…
Findings [1] 2026-09-25 b11189 opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (#29401) opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel opencl: fix s transpose - s only transposed for bin kernels Co-authored-by: Li He lih@qti.qualcomm.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50281199 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS…