Findings
-
[1] 2026-09-29 How Elasticsearch Serverless hollow shards cut indexing-node shutdowns by 30%
Elasticsearch Serverless now unloads idle indexing shards from memory. We call them hollow shards; the shard stays allocated, but its Lucene IndexWriter and segment readers are gone until the next write arrives. Hollow shards build on thin indexing shards, which… It then loads a HollowIndexEngine.We keep ingestion blocked on the target so the first write cannot sneak into the hollow engine.Hollow recovery skips translog replay and cache prewarm because an idle shard doesn’t need them. If an indexing node restarts,… In addition to the authors, we would like to thank Francisco Fernández Castaño, Tanguy Leroux, Benjamin Lerer, Albert Zaharovits, Artem Prigoda, and Henning Andersen, who designed, implemented, and productionized the work. We would also like to thank the wider Elasticsearch…
-
[2] 2026-09-29 Constella Preview: Swap Query Models Without Re-Embedding
As query traffic grows, so does the compute bill for embedding it. On a low-power device, a large model may not fit in memory. A smaller query model could reduce that cost, but switching usually means re-embedding the collection. Constella lets you make that switch. It’s a family of models on Hugging Face built around Stella, a 400M-parameter English embedding…
Where Kimbodo Comes In
Kimbodo builds and operates this in production for businesses — see our RAG Development Services practice, or Estimate My RAG System.