What Happened
Google DeepMind’s EmbeddingGemma 2 can run on a phone, but serving its embeddings across a large document collection is a different resource problem. For 10 million documents, full-size float32 vectors alone require 30.7 GB of RAM, before indexes, metadata, replicas or application overhead [1].
In early Qdrant tests, quantized full-size vectors used 30…
What Happened
Weaviate v1.39.3 fixes a high-severity credential disclosure vulnerability in its Google-backed embedding and generation modules. Earlier versions could send a configured Google API key or OAuth token to an attacker-controlled host if an unvalidated apiEndpoint was supplied. An attacker could set that endpoint through collection configuration with schema-write access or, for generative-google, through…
What Happened
Recent work shows that retrieval — not the generator — is the primary failure point in retrieval-augmented generation (RAG) and agent workflows: missing or low-quality candidates cause hallucination, context rot, excessive tool calls and unpredictable latency. Improving search quality reduces downstream errors and tool usage, but “improve retrieval” is too vague: you must…
Findings [1] 2026-09-29 How Elasticsearch Serverless hollow shards cut indexing-node shutdowns by 30% Elasticsearch Serverless now unloads idle indexing shards from memory. We call them hollow shards; the shard stays allocated, but its Lucene IndexWriter and segment readers are gone until the next write arrives. Hollow shards build on thin indexing shards, which… It…
Findings [1] 2026-09-28 The best LLM writes correct Elasticsearch ES|QL 59% of the time. Here's what breaks the other 41%. We gave four models 500 natural-language questions from BIRD's Mini-Dev set, asked each for one Elasticsearch Query Language (ES|QL) query, and graded all 6,000 answers on whether the right rows came back. The best…
Findings [1] 2026-09-25 GPU-accelerated vector indexing in Elasticsearch with NVIDIA cuVS: 138M vectors in under 10 minutes Modern enterprise applications are ingesting terabyte- to petabyte-scale unstructured data to power semantic search, large language model–based (LLM-based) retrieval augmented generation (RAG), and recommender systems. At this scale, vector indexing on CPUs can take days or even…
Findings [1] 2026-09-24 Ask Elastic Agent Builder why it's slow: Natural-language trace analysis Ask Elastic Agent Builder how many tokens your agents burned today, and it writes the Elasticsearch Query Language (ES|QL) and runs it against your OpenTelemetry (OTel) trace data. Then it answers in the chat UI. The same holds for your… That…
Findings [1] 2026-09-22 Agent Memory with Engram: A Practical Guide Every tool you've ever set up has an onboarding screen you just straight click past. Usually the defaults are chosen by someone who thought about them harder than you have time to. Engram's setup has one of those screens too,… That one line made…
Findings [1] 2026-09-21 Agentic workflows in Elasticsearch: pause an AI agent for human approval, resume 72 hours later An AI agent receives a question and processes it within seconds. The agent responds before the session expires. That model works well for question and answer or code generation. And it works well for point-in-time analysis.…
What Happened
Recent practical advances
Two developments show where production RAG (retrieval‑augmented generation) systems are trending: an end‑to‑end, low‑latency video search pipeline that combines scene detection, proxying and dense vectors indexed in Elasticsearch; and a single layout‑aware OCR VLM (jina-ocr-v1) that extracts structured text, tables, math and handwriting across 100+ languages in one call [1][2].…
What Happened
Weaviate 1.39 introduced 4-bit Rotational Quantization (RQ4, plus an uncentered variant RQ4c) and a SIMD Fast Walsh–Hadamard Transform implementation that significantly reduces encoder latency, index size and import time while preserving much of RAG recall when combined with recovery techniques. SIMD distance kernels, nibble operations and prefetch improvements yield large per-query and import…
What Happened
Retrieval‑Augmented Generation (RAG) has moved from prototypes to production patterns: embedding pipelines, ANN vector stores, hybrid BM25+vector retrieval, and reranking are now standard. Frameworks and orchestration layers (LlamaIndex, LangChain, Haystack) provide connectors and orchestration for embedding, retrieval and LLM chaining. Vector database options now span fully managed services and open source projects (Pinecone,…