Skip to content Skip to footer

AI Research & Papers — September 23, 2026

Findings

  1. [1] 2026-09-23 Offloaded inference for real-world physical AI robotics

    At a glance Challenges a core assumption in robotics AI: Our research shows that running physical AI inference exclusively on onboard GPUs can limit robot performance, battery life, and scalability, and that offloading inference to edge or cloud GPUs can… AI Testing and Evaluation: Learnings from Science and Industry Discover how Microsoft is learning from other domains to advance evaluation and testing as a pillar of AI governance. Listen now Opens in a new tab Toolset for automatic offload We…

  2. [2] 2026-09-23 MIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow

    David Siegel SM ’86, PhD ’91, a computer scientist, entrepreneur, and philanthropist, will serve as the next MIT Innovation Fellow during the 2026-27 academic year. Working with the MIT Schwarzman College of Computing, Siegel will explore how artificial intelligence can… “His long-standing commitment to MIT, together with his vision for the future of AI and science, makes him an especially fitting Innovation Fellow. We look forward to the contributions he will make and the connections his work will foster across…

  3. [3] 2026-09-23 PAGE: Partition-Aware Gated KV-Cache Eviction

    arXiv:2609.22157v2 Announce Type: replace-cross Abstract: KV-cache eviction can do more than compress. In long-context LLMs, keeping only some cached tokens sometimes matches or exceeds full-cache accuracy, because many redundant prefill tokens otherwise dilute attention away from the tokens that carry the answer. This benefit is not uniform, and evicting the wrong tokens can drop accuracy to zero on tasks that…

  4. [4] 2026-09-23 FrontierMath ErdH{o}s

    arXiv:2609.25050v1 Announce Type: new Abstract: We introduce FrontierMath ErdH{o}s (FME), a benchmark of 68 ErdH{o}s problems that are open as of August 2026. To solve a task in FME, AI systems must resolve (prove or disprove) one of the 68 conjectures in the proof assistant Lean. Our 68 problems were selected by the second author among 652 open problems on…

  5. [5] 2026-09-23 FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research

    arXiv:2602.14285v2 Announce Type: replace-cross Abstract: Automated scholarly paper review (ASPR) has entered the coexistence phase with traditional peer review, where artificial intelligence (AI) systems are increasingly incorporated into real-world manuscript evaluation. In parallel, research on automated and AI-assisted peer review has proliferated. Despite this momentum, empirical progress remains constrained by several critical limitations in existing datasets. While reviewers routinely evaluate…

  6. [6] 2026-09-23 Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

    arXiv:2609.25049v1 Announce Type: new Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that…

  7. [7] 2026-09-23 Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations

    arXiv:2609.03426v2 Announce Type: replace Abstract: Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to high parameter and activation costs. We propose Lngram v2, which decouples the number…

  8. [8] 2026-09-23 Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

    arXiv:2609.25048v1 Announce Type: new Abstract: How many prompts does on-policy distillation (OPD) need, and how does the answer depend on the student policies that generate its training responses? We study these two controls jointly: prompt breadth and rollout refresh. A 3×3 mathematical-reasoning experiment fixes 14,080 trajectories and 110 optimizer updates while varying the prompt bank and the number of response-generating…

  9. [9] 2026-09-23 S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

    arXiv:2608.00528v2 Announce Type: replace Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memory costs due to the Key-Value (KV) cache. Although low-rank compression of KV cache is a promising remedy, existing methods face a dilemma: offline approaches depend on external calibration data, whereas online approaches incur substantial compute…

  10. [10] 2026-09-23 Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

    arXiv:2609.25009v1 Announce Type: new Abstract: Numerically equivalent word problems should yield the same canonical answer whether a quantity is written as a decimal, fraction, percentage, number word, scientific notation, or an exactly converted unit. We generate 3,600 exact-rational problems and 8,600 prompts spanning five identity-preserving transformation families, and evaluate five open-weight systems. After a fixed syntax audit that normalizes common…

  11. [11] 2026-09-23 A Computational Approach to Measuring Semantic Change in Sanskrit Literature

    arXiv:2609.25012v1 Announce Type: new Abstract: Diachronic word embeddings have become the modern standard for tracking semantic change, yet they have been largely validated on modern, high-resource, and well-segmented languages. This paper tests whether the paradigm transfers to Sanskrit, an ancient, low-resource language whose phonological fusion (sandhi), morphological inflection, compounding, and polysemy pose a unique challenge. I assemble a 2.7M-token corpus…

  12. [12] 2026-09-23 Peerify: Benchmarking Peer-Review Claim Verification

    arXiv:2609.25046v1 Announce Type: new Abstract: Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely manual and time-consuming process. We present Peerify, a pipeline for manuscript-grounded verification of peer-review claims. Given a manuscript and a review comment, the Peerify pipeline decomposes reviews into atomic claims, retrieves relevant manuscript…

  13. [13] 2026-09-23 Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

    arXiv:2609.25028v1 Announce Type: new Abstract: QMSum provides no scorer, making query-focused meeting summarization results difficult to compare. We rescore or generate 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1 when moved from capped long input to 2,000-word retrieved spans. Fine-tuning it on this span regime recovers the loss. On test…

  14. [14] 2026-09-23 From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication

    arXiv:2609.25034v1 Announce Type: new Abstract: Central bank press conferences are not merely information releases — they are structured narratives. We study whether the shape of sentiment within a statement, not just its average tone, carries policy-relevant signals. Constructing sentiment arcs for ECB and Fed press conferences along three dimensions — monetary stance, economic outlook, and uncertainty — we assess their…

  15. [15] 2026-09-23 MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

    arXiv:2609.20850v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios,…

  16. [16] 2026-09-23 AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search

    arXiv:2609.25047v1 Announce Type: new Abstract: Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering. A popular line of such agents frames model building as a code search problem and solves it by tree search, in which each node is a candidate program and the tree grows by generating a child program…

  17. [17] 2026-09-23 Efficient Iterative Retrieval with Heterogeneous Batching

    arXiv:2609.25405v1 Announce Type: cross Abstract: Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these models in isolation. Coarse-grained partitioning, such as dedicating GPUs to specific tasks, fails to adapt to dynamic workloads and creates computational "bubbles". To address these,…

  18. [18] 2026-09-23 Low-Rank Attention Residuals

    arXiv:2607.09694v2 Announce Type: replace-cross Abstract: Attention Residuals (AttnRes) replace the fixed residual sum with depth-wise attention over previous sub-layer outputs in Large Language Models (LLMs), but use each output as both a full-dimensional key and value. This couples routing with representation and makes the cost of computing depth-routing scores scale with hidden width $d$. We propose Low-Rank Attention Residuals (LR-AttnRes),…

  19. [19] 2026-09-23 A Semantic Approach to the Academic Publishing Network: Document Vector Representations and Hybrid Structural-Semantic Fusion over OpenAlex Data

    arXiv:2609.26218v1 Announce Type: cross Abstract: Structural graph analysis of the academic publishing network captures the topological relationships between entities but does not see the content of works. Building on our structural approach, this work complements it with a semantic layer and a parameterized structural-semantic fusion. We represent scientific documents by citation-informed vector embeddings (SPECTER2) and store them in an embedded…

  20. [20] 2026-09-23 Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

    arXiv:2609.25938v1 Announce Type: cross Abstract: A conformal abstention certificate for text-to-SQL is only as truthful as the correctness labels it is calibrated on. The uncertainty pipelines that read confidence off execution consistency take those labels from the single database a benchmark ships, an oracle known to be lenient. We run a preregistered intervention on Spider-Realistic, swapping that database for the…

  21. [21] 2026-09-23 Faithful Autoformalization via Roundtrip Verification and Repair

    arXiv:2604.25031v3 Announce Type: replace Abstract: When an LLM formalizes natural language, how do we know the output is faithful? We propose a roundtrip verification approach which does not require ground-truth annotations: formalize a statement, translate the result back to natural language, re-formalize, and use a formal tool to check logical equivalence. When the two formalizations agree, this provides evidence of…

  22. [22] 2026-09-23 SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking

    arXiv:2508.15526v2 Announce Type: replace Abstract: The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabilities. To this end, numerous LLM safety evaluation benchmarks are proposed. However, existing benchmarks generally rely on labor-intensive manual curation, which causes excessive time and resource consumption. They also exhibit significant redundancy and limited difficulty. To…

  23. [23] 2026-09-23 Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

    arXiv:2609.25845v1 Announce Type: cross Abstract: Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-training raises equal-weight macro accuracy from 70.6% to 76.1%, with the gain concentrated on…

  24. [24] 2026-09-23 Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction

    arXiv:2609.25179v1 Announce Type: new Abstract: Predicting the remaining useful life (RUL) is essential for effective predictive maintenance. Spatio-Temporal Graph Neural Networks (ST-GNNs), which can model both temporal and spatial relationships by representing time series data as a sequence of graphs, have shown exceptional performance in RUL prediction. However, current ST-GNNs face several drawbacks. First, they require domain expertise or significant…

  25. [25] 2026-09-23 Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising

    arXiv:2605.10953v2 Announce Type: replace-cross Abstract: The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dense geophone deployments, distributed acoustic sensing (DAS) arrays, and large-scale 2D and 3D surveys. This expansion makes complex noise suppression increasingly challenging, especially when signal fidelity must be preserved. Conventional supervised deep learning methods…

  26. [26] 2026-09-23 Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning

    arXiv:2609.25166v1 Announce Type: new Abstract: Diffusion data-point unlearning is typically evaluated immediately after each deletion, even though subsequent requests may repeatedly update the same model. We identify sequential reappearance, a failure mode in which an instance that is initially judged to be forgotten later returns to the memorized regime without reuse of the deleted data or adversarial fine-tuning. To capture…

  27. [27] 2026-09-23 Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies

    arXiv:2609.20906v2 Announce Type: replace Abstract: Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to…

  28. [28] 2026-09-23 Learning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian

    arXiv:2609.25163v1 Announce Type: new Abstract: The paper proposes a novel data-driven framework for designing and training a feedback linearizing controller by explicitly incorporating relative degree based conditions into the learning process. This enables the conventional feedback controller components to be replaced by neural Lie derivatives, thereby facilitating a fully data-driven feedback linearization framework. Furthermore, practical closed-loop stability is established by…

  29. [29] 2026-09-23 Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion

    arXiv:2607.20742v2 Announce Type: replace Abstract: Multimodal learning is a robust approach to improve predictive performance in applications such as medical prognosis. However, the clinical applicability of models that use multimodal learning is hampered by their poor performance under noisy or uninformative data streams. Present fusion approaches often lack robust mechanisms for the dynamic assessment of data quality and for the…

  30. [30] 2026-09-23 Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

    arXiv:2609.25131v1 Announce Type: new Abstract: Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku puzzles shows that 9.4% of generated cells are correct at an intermediate step but incorrect in the final output. We introduce generation order into…

  31. [31] 2026-09-23 The Probabilistic Structure of Large Language Models

    arXiv:2609.25134v1 Announce Type: new Abstract: This paper presents a probabilistic perspective on large language models (LLMs), developed with the aim of bringing together, in a single self-contained account, tools that are usually treated separately across the literature. LLMs are described through probability measures on the set of sequences of tokens, specified via their autoregressive conditional distributions. Training is formulated as…

  32. [32] 2026-09-23 Stable Unsupervised Continual Chunking with Sheaf SyncMap

    arXiv:2609.25143v1 Announce Type: new Abstract: Unsupervised Continual chunking is a fundamental problem in machine learning and neuroscience, where the goal is to identify groups of states that frequently co-occur in temporal sequences. A key challenge is to form accurate chunks while maintaining their stability over time. In this work, we propose sheaf regularization to reduce local inconsistencies in Decentralized SyncMap,…

  33. [33] 2026-09-23 Brain-Inspired Hierarchical Modularity for General Continual Learning

    arXiv:2609.25146v1 Announce Type: new Abstract: Continual learning, the ability to learn from sequential experience while retaining and adapting prior knowledge, is central to intelligent systems operating in changing environments. However, conventional continual learning is typically studied with offline task-wise training and clear task boundaries, leaving a substantial gap from general continual learning under online, uncertain, and evolving data streams. In…

  34. [34] 2026-09-23 Dual-GNN Multilevel Coarsening for Maximum Independent Set

    arXiv:2609.25149v1 Announce Type: new Abstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive. Researchers often employ graph sparsification methods to improve computational efficiency. Traditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information. In this paper, we propose Graph Edge Sparsification (GES), a learning-based sparsification approach for Euclidean TSP.…

  35. [35] 2026-09-23 Exposing Blind Spots in Deep Imbalanced Regression Evaluation

    arXiv:2609.25152v1 Announce Type: new Abstract: Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to perform best in densely populated target regions even when reliable performance is required across the full target range. Despite rapid methodological progress, DIR evaluation remains constrained by three blind spots: it is dominated by image-based…

  36. [36] 2026-09-23 Glucose-ML: A collection of longitudinal diabetes datasets for development of robust AI solutions

    arXiv:2507.14077v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) algorithms are a critical part of state-of-the-art digital health technology for diabetes management. Yet, access to large high-quality datasets is creating barriers that impede development of robust AI solutions. To accelerate development of transparent, reproducible, and robust AI solutions, we present Glucose-ML, a collection of 10 publicly available diabetes datasets, released within…

  37. [37] 2026-09-23 Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules

    arXiv:2609.08333v3 Announce Type: replace-cross Abstract: Molecular size is coupled to composition, structure, and function, yet most 3D molecular generators require a predefined atom count. We introduce Equivariant-Free Transformer-Autoencoded Latent Flow Matching, a two-stage framework that samples a fixed-dimensional latent vector using flow matching and uses an autoregressive Transformer to determine molecular size, atom types, coordinates, and chemical attributes. Canonical atom…

  38. [38] 2026-09-23 RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty

    arXiv:2609.26069v1 Announce Type: cross Abstract: Simulation-based tutor selection can be unstable when predictively adequate learner models imply different policy rankings. RankCert certifies one of eight equal-budget tutoring policies only when model-averaged utility, probability-best, posterior regret, cross-domain rank, family coverage, and leave-one-domain-out and leave-one-visible-family-out averages support the same candidate; otherwise it abstains. We evaluated RankCert in 1,280 frozen held-out settings spanning…

  39. [39] 2026-09-23 Discovery-Driven Integration of Disjoint Tables via Text

    arXiv:2609.26658v1 Announce Type: cross Abstract: Integrating heterogeneous datasets within data lakes is a critical challenge, particularly for semantically related tables that lack the explicit attributes needed to be joined. We study Discovery-Driven Integration, where the relevant sources and their missing relational structure must be discovered before integration. In this setting, unstructured text provides the evidence that connects otherwise disjoint tables.…

  40. [40] 2026-09-23 GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement

    arXiv:2609.26361v1 Announce Type: cross Abstract: With the rapid pace of AI research and the hundreds of daily new publications, staying up-to-date with the latest developments has become increasingly difficult. For researchers, quickly identifying impactful work is essential, yet manually reviewing each new publication is impractical. Automated impact prediction methods help address this challenge, usually by combining various information sources available,…

  41. [41] 2026-09-23 Event-Based Early Warning of Vineyard Disease Risk from Environmental Time Series

    arXiv:2605.04548v2 Announce Type: replace Abstract: Accurate early warning of vineyard disease risk from environmental observations is essential for timely intervention and more sustainable crop protection. However, many existing studies formulate disease prediction as daily presence classification, which can favor persistence-driven predictions and provide only limited support for actionable short-horizon warning. In this paper, we present an event-based approach for early…

  42. [42] 2026-09-23 Transport-Coupled Bayesian Flows for Molecular Graph Generation

    arXiv:2510.10211v4 Announce Type: replace Abstract: Molecular graph generation (MGG) is essentially a multi-class generative task, aimed at predicting categories of atoms and bonds under strict chemical and structural constraints. However, many prevailing diffusion paradigms learn to regress numerical embeddings and rely on a hard discretization rule during sampling to recover discrete labels. This introduces a fundamental discrepancy between training and…

  43. [43] 2026-09-23 AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

    arXiv:2609.22592v1 Announce Type: new Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success from failure. Constructing such gyms remains manual, expensive, and static. Task sets saturate as models improve and are increasingly exposed to contamination. Synthetic generation offers scale, but…

  44. [44] 2026-09-23 Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective

    arXiv:2609.24103v1 Announce Type: cross Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first step toward safer algorithms in partially observable settings. Specifically, we extend Distributional Reinforcement Learning (DistRL)-which models the entire return distribution for fully observable domains-to…

  45. [45] 2026-09-23 EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability

    arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence. However, retrieval augmented generation (RAG) over heterogeneous enterprise data can suffer from citation drift, unsupported content, and weak source traceability. We present EvidenT (T = Trust + Transparency + Traceability), a lightweight pipeline that verifies extracted evidence against retrieved documents before…

  46. [46] 2026-09-23 HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

    arXiv:2607.04265v2 Announce Type: replace-cross Abstract: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion. We propose HALO-WA, a hybrid-attention latent-guided online reinforcement learning (RL) framework for WA models, which leverages latent features…

  47. [47] 2026-09-23 IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

    arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of…

  48. [48] 2026-09-23 Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization

    arXiv:2604.24952v2 Announce Type: replace-cross Abstract: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holistic annotations, resulting in severe label noise: images that excel in some dimensions but are deficient in others are simply marked as winner or loser. We theoretically demonstrate that compressing multi-dimensional preferences into binary labels generates…

  49. [49] 2026-09-23 PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

    arXiv:2609.22353v1 Announce Type: new Abstract: Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual predictions into timely and inspectable guidance is a distinct challenge. An object label or steering command does not explain which evidence supports a decision or when that evidence is unreliable. We present…

  50. [50] 2026-09-23 Social Influence and the Allocation of Scientific Attention in AI Populations

    arXiv:2609.22408v1 Announce Type: new Abstract: AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the…

  51. [51] 2026-09-23 Learning 3D biophysical cell properties from 2D images and cell-population statistics

    arXiv:2609.22410v1 Announce Type: new Abstract: Inferring 3D cellular properties from 2D microscopy is difficult when a reference instrument reports only population statistics rather than labels for individual cells. Here we develop a population-supervised framework that maps single 2D red-cell images to latent biophysical quantities and aggregates them to mean corpuscular volume, red-cell distribution width and mean corpuscular haemoglobin. The model…

  52. [52] 2026-09-23 Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

    arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and…

  53. [53] 2026-09-23 Goal-driven Variant Categorization

    arXiv:2609.22475v1 Announce Type: new Abstract: Process discovery rarely yields a single coherent process structure. For analysis, a common step is to cluster process variants based on structural similarity and then assign business meaning to the resulting groups. Since these partitions are not derived from the organization's goals, analysts must manually interpret and consolidate variants into business-meaningful categories. This judgment-intensive step…

  54. [54] 2026-09-23 Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

    arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial…

  55. [55] 2026-09-23 The Wisdom of Artificial Deliberative Crowds

    arXiv:2609.22497v1 Announce Type: new Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also…

  56. [56] 2026-09-23 KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

    arXiv:2609.24298v1 Announce Type: cross Abstract: What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width…

  57. [57] 2026-09-23 Organizational Principles Enable Collective Intelligence in Embodied AI

    arXiv:2609.11737v2 Announce Type: replace-cross Abstract: Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the physical tasks they perform impose fundamentally different coordination requirements. Here we show that principles from human organization theory can be operationalized to…

  58. [58] 2026-09-23 Subgoal Search For Complex Reasoning Tasks

    arXiv:2108.11204v4 Announce Type: replace Abstract: Humans excel in solving complex reasoning tasks through a mental process of moving from one idea to a related one. Inspired by this, we propose Subgoal Search (kSubS) method. Its key component is a learned subgoal generator that produces a diversity of subgoals that are both achievable and closer to the solution. Using subgoals reduces…

  59. [59] 2026-09-23 Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

    arXiv:2609.24651v1 Announce Type: cross Abstract: Diffusion and flow models, as promising generative paradigms for speech enhancement, face a training–inference mismatch: training uses analytical path states, whereas inference recursively evaluates models on self-generated rollout states along discretized sampling trajectories. This mismatch causes prediction and discretization errors to accumulate. To address it, we introduce Corrective Forcing (CoF), a post-training paradigm that forces…

  60. [60] 2026-09-23 Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition

    arXiv:2609.16752v2 Announce Type: replace Abstract: Cognitive Field Theory (CFT) proposes that cognition arises from memory-dressed collective dynamics that generate a persistent macroscopic cognitive field. Here we develop a Cognitive Field Network (CFN), a recurrent Transformer in which the organized hidden field re-enters subsequent inference through [ Phi_{n+1}=F_{theta}(X_{n+1},Phi_n). ] Rather than prescribing an explicit memory operation, the CFN allows new information…

  61. [61] 2026-09-23 How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

    arXiv:2606.05256v2 Announce Type: replace Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView. The intervention, conducted by unknown, external researchers and halted following ethical backlash, involved undisclosed AI-generated accounts engaging users in live debate. After public disclosure, Reddit authorized moderators to release an archive of the AI-generated comments, creating a rare opportunity to…

  62. [62] 2026-09-23 Toward an Unbiased Collective Memory for Efficient LLM-Based Agentic 6G Cross-Domain Management

    arXiv:2509.26200v2 Announce Type: replace-cross Abstract: Agentic artificial intelligence is a candidate enabler of Level-4 autonomy in sixth-generation (6G) networks, but agents reasoning over a shared memory inherit its distortions. We study cross-domain radio access network (RAN)–edge orchestration in which a RAN agent minimizing energy and an edge agent minimizing latency negotiate, validate proposals against a digital twin (DT), and share…

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Consulting & Strategy practice, or Request an AI Roadmap.

Sources

  1. [1] Offloaded inference for real-world physical AI robotics
  2. [2] MIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow
  3. [3] PAGE: Partition-Aware Gated KV-Cache Eviction
  4. [4] FrontierMath ErdH{o}s
  5. [5] FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research
  6. [6] Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione
  7. [7] Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations
  8. [8] Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation
  9. [9] S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching
  10. [10] Same Quantity, Different Answer: Numerical Representation Invariance in Language Models
  11. [11] A Computational Approach to Measuring Semantic Change in Sanskrit Literature
  12. [12] Peerify: Benchmarking Peer-Review Claim Verification
  13. [13] Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum
  14. [14] From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication
  15. [15] MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs
  16. [16] AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search
  17. [17] Efficient Iterative Retrieval with Heterogeneous Batching
  18. [18] Low-Rank Attention Residuals
  19. [19] A Semantic Approach to the Academic Publishing Network: Document Vector Representations and Hybrid Structural-Semantic Fusion over OpenAlex Data
  20. [20] Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL
  21. [21] Faithful Autoformalization via Roundtrip Verification and Repair
  22. [22] SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
  23. [23] Visual Jev: Accurate and Efficient Decisions from Shared Visual Context
  24. [24] Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction
  25. [25] Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising
  26. [26] Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning
  27. [27] Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies
  28. [28] Learning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian
  29. [29] Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion
  30. [30] Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow
  31. [31] The Probabilistic Structure of Large Language Models
  32. [32] Stable Unsupervised Continual Chunking with Sheaf SyncMap
  33. [33] Brain-Inspired Hierarchical Modularity for General Continual Learning
  34. [34] Dual-GNN Multilevel Coarsening for Maximum Independent Set
  35. [35] Exposing Blind Spots in Deep Imbalanced Regression Evaluation
  36. [36] Glucose-ML: A collection of longitudinal diabetes datasets for development of robust AI solutions
  37. [37] Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules
  38. [38] RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty
  39. [39] Discovery-Driven Integration of Disjoint Tables via Text
  40. [40] GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement
  41. [41] Event-Based Early Warning of Vineyard Disease Risk from Environmental Time Series
  42. [42] Transport-Coupled Bayesian Flows for Molecular Graph Generation
  43. [43] AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
  44. [44] Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective
  45. [45] EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
  46. [46] HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
  47. [47] IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
  48. [48] Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
  49. [49] PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation
  50. [50] Social Influence and the Allocation of Scientific Attention in AI Populations
  51. [51] Learning 3D biophysical cell properties from 2D images and cell-population statistics
  52. [52] Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
  53. [53] Goal-driven Variant Categorization
  54. [54] Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
  55. [55] The Wisdom of Artificial Deliberative Crowds
  56. [56] KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
  57. [57] Organizational Principles Enable Collective Intelligence in Embodied AI
  58. [58] Subgoal Search For Complex Reasoning Tasks
  59. [59] Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement
  60. [60] Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition
  61. [61] How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment
  62. [62] Toward an Unbiased Collective Memory for Efficient LLM-Based Agentic 6G Cross-Domain Management

Leave a comment

0.0/5