What Happened
Two release candidates, v0.40.0-rc4 and v0.40.0-rc5, include MLX-related changes. The rc4 notes report an MLX version bump and added unit-test scopes intended to reduce memory use [2]. The rc5 notes describe a fix to a patch following a recent MLX update, tracked as #18812, but give no implementation details [1]. The notes do…
What Happened
Ollama v0.40.0 runs models with MLX-supported architectures on MLX by default on Apple Silicon. The supported list includes qwen3.8, gemma4, qwen3.6 and qwen3.5, alongside several decision models. Its release candidate also refined MLX tokenization to match publisher behavior, covering Unicode boundaries, added tokens and BPE merges. The published changelog spans v0.34.4 through the…
What Happened
LiteLLM 1.103.3 makes the proxy migration check mandatory by default and prevents migration checks from creating hand-built SpendLogs indexes. It also updates dependencies, bumps litellm-proxy-extras to 0.4.100.post1, and backports fixes. Its Docker image is signed with cosign; LiteLLM recommends verifying it with the public key from immutable commit 0112e53046018d726492c814b3644b7d376029d0 [1].
LiteLLM 1.104.0 adds…
What Happened
Several releases affect application behavior, model serving, and gateway operations. Streamlit 1.65.0 adds on_change="ignore" to several widgets, broader alt-text support, required inputs, URL-bound tabs and expanders, and side-drawer dialogs. It also fixes browser navigation state, forms, dates, and widget behavior. The changelog covers changes since 1.64.0 [1].
Ollama 0.35.1 supports Clef decision models…
What Happened
LiteLLM v1.103.2 backports proxy and Anthropic fixes to its stable branch. LiteLLM v1.101.4 syncs its stable branch to v1.101.3 and backports three changes; the release notes do not identify those changes. Both releases state that LiteLLM Docker images are signed with cosign and recommend verifying against a public key pinned to commit 0112e53046018d726492c814b3644b7d376029d0,…
What Happened
Major model/runtime release (5.18.0) — Added new multimodal and diarization models (Nemotron3 Diarization, NemotronH Omni, HyperCLOVAX Vision V2, GTE embeddings), broad bug and tokenizer fixes, memory/offload and parallelism improvements (FSDP2, 2‑D device mesh, MoE GGUF support), hardware/kernel fixes (MPS/FP8, XPU kernels), and many CI/tooling updates [1].
…
Findings [1] 2026-09-30 v1.104.0-rc.2 Verify Docker Image Signature All LiteLLM Docker images are signed with cosign. Every release is signed with the same key introduced in commit 0112e53. Verify using the pinned commit hash (recommended): A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original…
Findings [1] 2026-09-29 v0.35.0 Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification. Available models: Nimble from Bespoke Labs Tev1 from Together AI ollama pull nimble…
Findings [1] 2026-09-27 v1.103.0 fix(proxy): unregister logging callbacks removed from the stored conf… [2] 2026-09-27 1.64.1.dev20260926 Streamlit nightly 1.64.1.dev20260926
Where Kimbodo Comes In Kimbodo builds and operates this in production for businesses — see our AI Application Development practice, or Estimate My AI Application.
Sources
[1] v1.103.0 [2] 1.64.1.dev20260926
Findings [1] 2026-09-26 1.64.1.dev20260925 Streamlit nightly 1.64.1.dev20260925
Where Kimbodo Comes In Kimbodo builds and operates this in production for businesses — see our AI Application Development practice, or Estimate My AI Application.
Sources
[1] 1.64.1.dev20260925
Findings [1] 2026-09-25 v0.40.0 What's Changed Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX. ollama pull qwen3.8 ollama run qwen3.8 During the pre-release we will be testing and enabling additional models. Full Changelog: v0.34.4...v0.40.0-rc0…
Findings [1] 2026-09-24 langchain-core==1.6.5 Changes since langchain-core==1.6.4 release(core): 1.6.5 (#40816) fix(core): abbreviate long tool IDs in XML buffer strings (#40792) [2] 2026-09-24 langchain-openai==1.6.6 Changes since langchain-openai==1.6.5 release(openai): 1.6.6 (#40800) fix(openai): raise on error events in stream path (#40791) [3] 2026-09-24 v0.34.4 What's Changed Structured outputs on thinking models now apply…