Skip to content Skip to footer

What New Ollama and LiteLLM Releases Mean for Production AI Deployments

What Happened

Ollama v0.40.0 runs models with MLX-supported architectures on MLX by default on Apple Silicon. The supported list includes qwen3.8, gemma4, qwen3.6 and qwen3.5, alongside several decision models. Its release candidate also refined MLX tokenization to match publisher behavior, covering Unicode boundaries, added tokens and BPE merges. The published changelog spans v0.34.4 through the v0.40.0 release candidates, so teams upgrading from older versions should review intermediate changes. [1][3]

LiteLLM v1.105.0-rc.1 adds a Microsoft 365 Graph server to its MCP catalog and scoped SQL queries with schema-aware tracing help. Reported fixes address proxy startup and logging, guardrail routing and verdict handling, Azure Storage file isolation, and migration recovery. Its Docker images are signed with cosign. This is a release candidate, not a stable release. [5]

Streamlit 1.65.1.dev20261003 is a nightly development build. No release notes or changes were supplied, so there is no basis to claim a new feature or fix. The Ollama v0.40.0-rc2 entry is a branch merge without change details. [4][2]

Why It Matters to Businesses

Ollama’s new Apple Silicon default can change the execution path without an application change. Teams should check model output, latency and memory use before promoting the version; the tokenizer work makes exact prompt and output regression tests particularly relevant. LiteLLM’s fixes touch operationally sensitive paths—guardrails, file isolation and recovery—while its new integrations expand what must be authorized and audited. [1][3][5]

Kimbodo Engineering Perspective

We would treat the Ollama default as a behavior change, not merely a performance feature. A faster local backend is useful only if it preserves application-level results within agreed tolerances. We would evaluate LiteLLM’s release candidate separately from stable upgrades: its fixes may justify early testing, but not automatic production rollout. A Streamlit nightly with no change notes belongs in a test environment until its behavior and release status are clear. [1][4][5]

How We Would Implement It

  • Pin current versions and capture baseline prompts, outputs, latency and resource use for each deployed model and hardware class. Test Ollama v0.40.0 on Apple Silicon with tokenization edge cases and representative workloads before changing the production image. [1][3]
  • Stage LiteLLM v1.105.0-rc.1 behind integration tests for guardrail decisions, Azure Storage isolation, migrations, proxy startup and logging. Enable the Graph MCP server or scoped SQL access only for explicitly approved use cases. [5]
  • Verify LiteLLM image signatures with cosign against a pinned commit hash, then promote immutable image digests through environments. Pin the Streamlit nightly only if a specific test requires it. [5][4]

Risks, Costs and Security

Changing Ollama’s default backend may shift memory demand and performance, while tokenizer differences can affect prompts, output and downstream tests. New LiteLLM connectors and query capabilities increase permission and data-exposure risk; limit credentials and query scope, and retain audit trails. Signature verification improves artifact provenance but does not replace vulnerability review. Budget for regression testing and rollback capacity, especially when assessing a release candidate or nightly build. [1][3][5][4]

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Application Development practice, or Estimate My AI Application.

Sources

  1. [1] v0.40.0
  2. [2] v0.40.0-rc2
  3. [3] v0.40.0-rc1: mlx: match publisher tokenizer semantics (#18779)
  4. [4] 1.65.1.dev20261003
  5. [5] v1.105.0-rc.1

Leave a comment

0.0/5