Findings
-
[1] 2026-09-23 v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550)
mlx: speed up Qwen 3.8 prompt processing Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU. address comments
-
[2] 2026-09-22 langchain-openai==1.6.4
Changes since langchain-openai==1.6.3 release(openai): 1.6.4 (#40775) chore(model-profiles): refresh openai model profile data (#40774)
-
[3] 2026-09-22 langchain-anthropic==1.7.3
Changes since langchain-anthropic==1.7.2 release(anthropic): 1.7.3 (#40773) chore(model-profiles): refresh anthropic model profile data (#40772) fix(anthropic): auto-route with_structured_output to method="json_schema" for fable and opus 5.5 (#40766) chore(anthropic): update docs for Opus 5.5 (#40765) feat(anthropic): send mid-conversation SystemMessages in place (#40622) chore(deps): bump anyio from 4.11.0 to 4.14.2 in /libs/partners/anthropic (#40643)
-
[4] 2026-09-22 v0.34.3
What's Changed GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}' { "thinking": { "values": ["low", "high", "max"], "default": "max" } } Also available on ollama.com directly for cloud models. Nemotron H vision models are…
-
[5] 2026-09-22 v1.102.0
Verify Docker Image Signature All LiteLLM Docker images are signed with cosign. Every release is signed with the same key introduced in commit 0112e53. Verify using the pinned commit hash (recommended): A commit hash is cryptographically immutable, so this is… feat(policy_engine): execute post_call guardrail pipelines on streaming responses by @mateo-berri in #38788 fix(policy_engine): apply post_call pipeline text rewrites on streams by @mateo-berri in #39233 fix(proxy): redact provider keys from pass-through failure tracebacks by @mateo-berri in #39964 fix(file_search): scope emulated file_search… open vector store details by @devin-ai-integration[bot] in #40150 feat(ui): make automatic auto-router setup discoverable and show what it configured by @tin-berri in #40146 fix(docs): fix stale file paths in ARCHITECTURE.md by @jrlprost in #40157 fix(guardrails): accept on_violation block and alert… efusing deployment on every router entrypoint by @devin-ai-integration[bot] in #40306 fix(proxy): default max_idle_connection_lifetime on componentized DB URLs by @devin-ai-integration[bot] in #40285 fix(a2a): forward caller identity headers on message/send and message/stream by @yassin-berriai in #40305 feat(team): let a team admin manage… bot] in #40271 fix(proxy): accept non-string callback vars in default_team_settings by @devin-ai-integration[bot] in #40458 fix(router): resolve team-scoped auto-routers by their public name by @tin-berri in #40432 fix(langfuse): give each call in a session header its own trace instead of upserting… ion[bot] in #40527 feat(mock): report admission-time input token count in mock_response usage by @devin-ai-integration[bot] in #40590 perf(proxy): reuse cached model group and deployment info in budget reservation by @devin-ai-integration[bot] in #40593 fix(router): strip Codex harness envelopes before classification by @moe-berri…
-
[6] 2026-09-22 langchain-fireworks==1.6.2
Changes since langchain-fireworks==1.6.1 fix(fireworks): use current completions model in LLM tests (#40740) hotfix(fireworks): use available model in LLM tests (#40737) release(fireworks): 1.6.2 (#40735) chore(model-profiles): refresh model profile data (#40665) chore(deps): bump anyio from 4.11.0 to 4.14.2 in /libs/partners/fireworks (#40639) chore(deps): bump urllib3 from 2.7.0 to 2.8.0 in /libs/partners/fireworks (#40587) chore(deps): bump langsmith from 0.12.1 to 0.12.6 in /libs/partners/fireworks (#40586) chore(deps):…
-
[7] 2026-09-22 1.64.1.dev20260921
Streamlit nightly 1.64.1.dev20260921
-
[8] 2026-09-22 langchain-deepseek==1.1.1
Changes since langchain-deepseek==1.1.0 fix(deepseek,infra): resolve compatible minimum OpenAI dependencies, bump min ver (#40738) release(deepseek): 1.1.1 (#40734) chore(deps): bump anyio from 4.11.0 to 4.14.2 in /libs/partners/deepseek (#40641) chore(model-profiles): refresh model profile data (#40399) fix(deepseek): route strict mode to the beta endpoint (#40249) chore(model-profiles): refresh model profile data (#39906) chore(model-profiles): refresh model profile data (#39844) fix(deepseek): map prompt_cache_hit_tokens to cache_read (#39668) chore(model-profiles):…
-
[9] 2026-09-22 langchain-openrouter==0.2.9
Changes since langchain-openrouter==0.2.8 release(openrouter): 0.2.9 (#40736) chore(model-profiles): refresh model profile data (#40705) chore(model-profiles): refresh model profile data (#40685) chore(model-profiles): refresh model profile data (#40665) chore(deps): bump anyio from 4.13.0 to 4.14.2 in /libs/partners/openrouter (#40627) chore(model-profiles): refresh model profile data (#40600) chore(model-profiles): refresh model profile data (#40541) chore(model-profiles): refresh model profile data (#40500) chore(model-profiles): refresh model profile data (#40452) chore(model-profiles): refresh…
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our AI Application Development practice, or Estimate My AI Application.