Skip to content Skip to footer

AI Research & Papers — September 25, 2026

Findings

  1. [1] 2026-09-25 MIT students gain a humanist lens on technical innovation in Tulsa, Oklahoma

    When MIT mechanical engineering student Daphne Wang arrived for her internship at the Muscogee Creek Nation Department of Health in Tulsa, Oklahoma last summer, she expected to be writing code. To her surprise, she found herself analyzing tribal history and… also delivered remarks at TASC, encouraging students as future leaders to take a public interest in technology. “Technology can be used for good, or it can be used for harm. Your generation has the opportunity to bend that arc toward something…

  2. [2] 2026-09-25 SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions

    arXiv:2512.14277v2 Announce Type: replace-cross Abstract: The advent of large language models is contributing to the emergence of novel approaches that promise to better tackle the challenge of generating structured queries, such as SPARQL queries, from natural language. However, these new approaches mostly focus on response accuracy while ignoring other evaluation criteria, such as runtime and cost to generate SPARQL queries.…

  3. [3] 2026-09-25 Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

    arXiv:2609.29001v1 Announce Type: new Abstract: Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is…

  4. [4] 2026-09-25 DuplexDrama: A Synthesized Dialogue Dataset with Scenarios, Full-Duplex Behaviors, Expressive Speech, and Sound Events

    arXiv:2609.12872v2 Announce Type: replace Abstract: We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); (iii) expressive speech with persona-aligned emotion labels; and (iv) script-aware sound events. DuplexDrama is built via a 4-stage pipeline; quality validation on both scripts and synthesized audio confirms…

  5. [5] 2026-09-25 Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

    arXiv:2609.28854v1 Announce Type: new Abstract: Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism – here the sales representative, a witness recorded in the CRM -…

  6. [6] 2026-09-25 Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction

    arXiv:2601.18395v3 Announce Type: replace Abstract: Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard practices include prompting decoder-only LLMs using greedy decoding to avoid output variability. Rather than treating this variability as a limitation, we show that sampling can produce substantially better solutions than greedy…

  7. [7] 2026-09-25 COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages

    arXiv:2609.28826v1 Announce Type: new Abstract: Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from English-pivot content and often fail to capture the linguistic diversity, cultural complexity, and domain-specific characteristics of Indian languages. We present COILD, an Indic-centric parallel corpus comprising over 1.16…

  8. [8] 2026-09-25 MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis

    arXiv:2410.14948v3 Announce Type: replace Abstract: Medical multimodal large language models (MLLMs) can perform well on existing medical visual question answering (MedVQA) benchmarks, but their training data often does not match clinical diagnosis. Most supervision is organized around isolated images or short QA pairs, leaving two structures weakly specified: how evidence leads to a decision, and how views, series, modalities, and…

  9. [9] 2026-09-25 Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks

    arXiv:2609.28673v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically,…

  10. [10] 2026-09-25 An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection

    arXiv:2609.28703v1 Announce Type: new Abstract: Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary classification, evaluated their frameworks on a single dataset, and provide limited insight into how decisions are made, which limits their real-world applicability.…

  11. [11] 2026-09-25 PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

    arXiv:2609.28727v1 Announce Type: new Abstract: Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals.…

  12. [12] 2026-09-25 Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus

    arXiv:2609.28747v1 Announce Type: new Abstract: A language model's confidence in an answer is often read as a proxy for how well it knows the corresponding fact. This manual documents an open toolkit built to test that reading directly: a small causal language model is fine-tuned on a corpus that consistently asserts one fabricated arithmetic answer for each of the 81…

  13. [13] 2026-09-25 Temporal Taxation Compounds Under Post-Training Compression of Whisper Models

    arXiv:2609.28739v1 Announce Type: new Abstract: Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which alters model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups. Across the Whisper family on…

  14. [14] 2026-09-25 Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

    arXiv:2605.29064v3 Announce Type: replace Abstract: This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional layers: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations from each of…

  15. [15] 2026-09-25 Script Choice in LLMs: Evidence for Late-Layer Commitment

    arXiv:2609.28784v1 Announce Type: new Abstract: In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a clear asymmetry: both the input script and the instructed output script are encoded in the earliest layers of the network, while, in contrast, commitment…

  16. [16] 2026-09-25 Recovering the Zipfian Distribution in Unsupervised Term Discovery

    arXiv:2606.10781v4 Announce Type: replace-cross Abstract: Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Zipfian distribution, yet the dominant centre-based clustering approach — K-means — produces a more uniform distribution due to an inductive bias toward spherical clusters. In this paper we revisit graph-based…

  17. [17] 2026-09-25 Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research

    arXiv:2609.27690v2 Announce Type: replace Abstract: Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong…

  18. [18] 2026-09-25 CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding

    arXiv:2609.29474v1 Announce Type: cross Abstract: Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge —i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve— remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for…

  19. [19] 2026-09-25 Reward Hacking Challenges Oversight of Autonomous Research Agents

    arXiv:2609.28614v1 Announce Type: new Abstract: Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack without instructions to do so, (2)…

  20. [20] 2026-09-25 PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

    arXiv:2609.30094v1 Announce Type: cross Abstract: Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce textbf{PrivDrift}, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic…

  21. [21] 2026-09-25 iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

    arXiv:2609.29626v1 Announce Type: cross Abstract: Recursive AI, the prospect of AI taking an increasingly complete role in building and improving AI, is a crown jewel of AI for AI. Although recursive self-development has become practical for small models, bounded tasks, and fixed time budgets, a more consequential realization of this ambition, i.e., developing a release-ready, frontier-competitive model, remains far more…

  22. [22] 2026-09-25 DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models

    arXiv:2609.29092v1 Announce Type: cross Abstract: Vision-based legged locomotion methods assume clean depth at training time and rely on hand-tuned post-processing filters at deployment. However, filter parameters are rarely disclosed, hindering reproducibility, and performance degrades substantially when depth noise is left unaddressed. Building noise robustness directly into the learning pipeline would eliminate this dependency. While such robustness has been explored for…

  23. [23] 2026-09-25 SGA: Uncertainty Quantification for Multi-Step Forecasting in Time Series Foundation Models

    arXiv:2609.28582v1 Announce Type: new Abstract: The recent emergence of Time Series Foundation Models (TSFMs) has significantly advanced multi-step forecasting performance, enabling accurate predictions over extended future horizons. However, existing TSFMs often suffer from significantly inherent uncertainty, which typically manifests as derived forecast branches emerging at each time step and spreading to subsequent steps; different forecast branches often exhibit varying forecasting…

  24. [24] 2026-09-25 IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

    arXiv:2604.07709v5 Announce Type: replace-cross Abstract: A strongly safety-trained model will provide a doctor with a benzodiazepine taper schedule, but not a patient who asks for one. The model knows the information, but how much it shares depends on the framing. We introduce IatroBench, a benchmark that evaluates models on two axes of harm (commission and omission) across 60 pre-registered clinical…

  25. [25] 2026-09-25 Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning

    arXiv:2609.28581v1 Announce Type: new Abstract: Reinforcement-learning (RL) policies are often distributed as opaque neural checkpoints, while training logs show that a run occurred without explaining what the policy learned. We study whether independently trained policies can be represented and composed through auditable discrete behavioral rules. We define auditability as six separately testable predicates: trace integrity, lossless coding, rule coverage, behavioral…

  26. [26] 2026-09-25 CyFM: Cylindrical Optimal Transport for Few-Step Complex-Valued Flow Matching

    arXiv:2609.14171v2 Announce Type: replace Abstract: Complex-valued signals like MRI and audio spectrograms are typically modelled as flat two-channel Euclidean data. The inherited Euclidean metric $dA^2 + A^2 dtheta^2$ vanishes at the origin, leaving phase unpenalised exactly where the signal is weakest. We replace it with the decoupled product metric $dA^2 + dtheta^2$ on the cylindrical closure $[0, infty) times S^1$,…

  27. [27] 2026-09-25 Uncovering Residential PV-EV Co-Adoption from Smart-Meter Data: Load Archetypes and Detection for Demand-Side Planning

    arXiv:2609.28578v1 Announce Type: new Abstract: The increasing adoption of electric vehicles (EVs) and rooftop photovoltaic (PV) systems is reshaping residential electricity demand and creating new challenges for demand-side management (DSM), tariff design, and low-voltage network planning. Much of the existing literature examines EV charging or PV generation in isolation, leaving the behavioral dynamics of household co-adoption less understood. We develop…

  28. [28] 2026-09-25 Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

    arXiv:2608.02829v4 Announce Type: replace Abstract: Model families are trained size by size. Can a pretrained large model instead be converted into a smaller sibling? We study the 1.4B->410M conversion in Pythia end to end. Representations align strongly across sizes (ridge R^2=0.84); parameters align weakly. Dense weight projection is destructive; a bit-exact control places the fault in basis mixing, which breaks…

  29. [29] 2026-09-25 CFD Correction of Open Tip Clearance Flow in a Compressor Cascade Using VAE Latent Space Adaptation

    arXiv:2609.28558v1 Announce Type: new Abstract: CFD predictions of open tip clearance flow in compressor cascades are subject to discrepancies relative to experiments, while experimental observations are sparse and high-resolution experimental ground truth is unavailable. This study proposes a non-intrusive correction method based on a variational autoencoder (VAE) and latent-space adaptation. A VAE is first trained using a dataset of 166…

  30. [30] 2026-09-25 CARE: Condition-Aware Representation Regularization for Diffusion Models

    arXiv:2609.28561v1 Announce Type: new Abstract: Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE…

  31. [31] 2026-09-25 SpaFactor: Lightweight Spatial Context-Aware Gene Program Modeling for Histology-to-Transcriptomics Inference

    arXiv:2609.28563v1 Announce Type: new Abstract: Spatial transcriptomics (ST) profiles gene expression within tissue architecture, but its cost and experimental complexity limit routine use. Predicting spatial expression from routinely available hematoxylin and eosin (HE) images therefore offers a scalable alternative. However, conventional methods often fit high-dimensional gene outputs as independent targets, overlooking the biological coordination among genes while remaining vulnerable to…

  32. [32] 2026-09-25 Leakage-Safe Machine Learning for Hydrogen Embrittlement Detection in 316L Stainless Steel: A Region-Held-Out Evaluation of Texture and Deep Features in SEM Micrographs

    arXiv:2609.28567v1 Announce Type: new Abstract: Scanning electron microscopy (SEM) is routinely used to characterize the microstructural changes caused by hydrogen embrittlement (HE) in structural steels. Machine learning can automate this characterization, but models are often evaluated using image-level splits. When several images come from the same specimen region, such splits leak information between the training and test sets. Here, we…

  33. [33] 2026-09-25 When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages

    arXiv:2609.28565v1 Announce Type: new Abstract: Post hoc explanation methods such as SHAP and LIME are widely used to interpret text classifiers, but their visualizations are mainly designed for left-to-right languages. When applied to right-to-left (RTL) languages such as Urdu, Arabic, Persian, and Hebrew, the attribution values remain mathematically valid, while their visual presentation fails. Tokens appear out of sequence, connected…

  34. [34] 2026-09-25 Foundations of Large Language Models

    arXiv:2501.09223v3 Announce Type: replace-cross Abstract: This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals,…

  35. [35] 2026-09-25 Time-Series Foundation Models That Understand Data Revisions

    arXiv:2609.28576v1 Announce Type: new Abstract: Historical observations are not always fixed: statistical agencies revise previously published values as new evidence arrives. Forecasting from a contemporary download can therefore expose a model to information unavailable at the date it purportedly made a prediction. We propose VINTAGE-TS, a revision-aware adaptation of a time-series foundation model that distinguishes observation time from information-availability time.…

  36. [36] 2026-09-25 AFT Neural Function Approximators for 1D Nonlinear Force Laws

    arXiv:2609.29242v1 Announce Type: cross Abstract: Nonlinear contacts and friction strongly influence the vibration response of assembled structures, but their accurate numerical treatment is computationally demanding. The harmonic balance method is widely used to compute periodic steady-state responses, yet the required alternating frequency-time scheme becomes costly for nonsmooth and hysteretic nonlinearities and must be repeated throughout the nonlinear solution process. Here…

  37. [37] 2026-09-25 An explicit solution of the five-expert prediction PDE and the exact optimality set of COMB

    arXiv:2609.14892v2 Announce Type: replace-cross Abstract: In this paper, we derive an explicit solution of the stationary prediction with expert advice PDE for five experts. The formula is given in three regions. In the first two regions, it is the four-expert solution plus a single integral with an elementary positive density. In the third region, it is a finite sum of…

  38. [38] 2026-09-25 Cost-Sensitive Online Window Size Selection for Portfolio Management

    arXiv:2609.29887v1 Announce Type: cross Abstract: This paper investigates cost-sensitive online window size selection for portfolio management under changing market conditions. Specifically, we propose a two-level framework that constructs portfolios using candidate window sizes and dynamically aggregates them through online learning. By treating candidate window sizes as “experts,'' we dynamically update their aggregation weights using turnover-inclusive losses. Moreover, we derive finite-horizon…

  39. [39] 2026-09-25 An Exploratory Ablation of a Small MLA–SSM Hybrid Language Model

    arXiv:2609.29618v1 Announce Type: cross Abstract: We report an exploratory, single-seed ablation of TALH (Adaptive Latent Hybrid), a decoder-only language model with parallel Multi-head Latent Attention (MLA) and a custom recurrent state-space (SSM) branch. Five variants, spanning 117–217M estimated active parameters per token, are trained from scratch on a FineWeb sample for the same number of optimisation steps and tokens. In…

  40. [40] 2026-09-25 Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

    arXiv:2602.08329v2 Announce Type: replace Abstract: A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve the quality of dense attention while sharply reducing computation and bandwidth, existing sparse methods generally rely on posterior heuristics, i.e., selectors conditioned on observed attention or proxy scores.…

  41. [41] 2026-09-25 Learning and interpreting policies for simultaneous entanglement requests in quantum networks

    arXiv:2609.30157v1 Announce Type: cross Abstract: Future quantum networks will make use of entanglement to perform numerous tasks, such as sending quantum information over long distances, distributed quantum computing, and quantum sensing. In general, these tasks will need to be performed simultaneously in various regions of a network, while minimizing resources and latency. We will thus require policies for scheduling link-level…

  42. [42] 2026-09-25 Aftab: A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning

    arXiv:2608.07335v3 Announce Type: replace-cross Abstract: Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this question through a progressive three-phase study within the Parallelized Q-Network (PQN) framework. First, we compare eight convolutional encoder topologies on Atari-57 under…

  43. [43] 2026-09-25 Driving Epidemic Models with AI Agents: the Epydemix Agent Framework

    arXiv:2609.28692v1 Announce Type: new Abstract: Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an…

  44. [44] 2026-09-25 HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

    arXiv:2602.00993v3 Announce Type: replace-cross Abstract: End-to-end autonomous driving models increasingly benefit from large vision-language models for semantic understanding, yet safe and reliable planning under long-tail conditions remains challenging, particularly in mixed-traffic environments involving heterogeneous road users and rare safety-critical interactions. This paper proposes HERMES, a holistic risk-aware end-to-end multimodal driving framework that explicitly incorporates long-tail semantic knowledge into trajectory planning.…

  45. [45] 2026-09-25 Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

    arXiv:2609.28690v1 Announce Type: new Abstract: Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real…

  46. [46] 2026-09-25 Answering Path Queries under Linear and Guarded Existential Rules

    arXiv:2607.22636v2 Announce Type: replace Abstract: Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance and an ontology. While most work in the area focuses on conjunctive queries (CQs), navigational queries have gained increasing attention. In this paper, we investigate the complexity of answering two-way (conjunctive) regular path queries ((C)RPQs) over…

  47. [47] 2026-09-25 Training Object Permanence in World Models

    arXiv:2609.28654v1 Announce Type: new Abstract: Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them…

  48. [48] 2026-09-25 To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

    arXiv:2609.30227v1 Announce Type: cross Abstract: Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can verify claims directly from speech. We introduce VeriSpeak, a probe benchmark for studying speech-based fact verification in Large Audio Language Models (LALMs). VeriSpeak contains 3,879 spoken claims spanning…

  49. [49] 2026-09-25 PAWS: Policy-driven Agentic World Simulation

    arXiv:2609.28547v1 Announce Type: new Abstract: Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action…

  50. [50] 2026-09-25 Pistis Technical Report

    arXiv:2609.28554v1 Announce Type: new Abstract: We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL),…

  51. [51] 2026-09-25 BaseCamp — An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

    arXiv:2609.28557v1 Announce Type: new Abstract: DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings…

  52. [52] 2026-09-25 TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

    arXiv:2609.28575v1 Announce Type: new Abstract: Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality — whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted…

  53. [53] 2026-09-25 DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

    arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the emph{correction chain} from reward to parameter update. At the rollout level, hard queries—those with high semantic entropy—frequently produce unanimously wrong sample groups, collapsing the…

  54. [54] 2026-09-25 SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design

    arXiv:2609.27621v2 Announce Type: replace Abstract: Physical modeling and inverse design require computation that can continue from reusable state. We introduce SHRAV, an architecture-independent computational framework organized around State, Hypothesis, Reason, Action, and Verify. Its central mechanism is a state-continuation core with declared reuse boundaries and explicit roles for learned evolution and numerical quantities. Forward configurations evolve predictive state and read…

  55. [55] 2026-09-25 Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents

    arXiv:2609.28609v1 Announce Type: new Abstract: Role-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where it performs poorly also change, while the training distribution…

  56. [56] 2026-09-25 Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

    arXiv:2609.26176v2 Announce Type: replace-cross Abstract: Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. A high refusal rate there is reported as safety, and it is equally consistent with a model that has stopped telling the request apart from anything else in the same format. We…

  57. [57] 2026-09-25 When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills

    arXiv:2605.25832v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator results shape the next population but are not preserved as reusable design knowledge. We present Auto-Robotist, a self-evolving LLM agent that distills morphology-search traces into an explicit natural-language skill library. Each skill stores a structural…

  58. [58] 2026-09-25 Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning

    arXiv:2609.29711v1 Announce Type: cross Abstract: Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output control concerns whether a trained model selectively reduces the likelihood of sensitive content in its outputs. In this work, we investigate the latter together with continual-learning utility under…

  59. [59] 2026-09-25 When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI

    arXiv:2609.29547v1 Announce Type: cross Abstract: Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the runtime infrastructure that defines authority, records action, interrupts execution, checks outcomes, and supports repair. We call this the reduced-supervision paradox.…

  60. [60] 2026-09-25 Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits

    arXiv:2609.30017v1 Announce Type: cross Abstract: Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it. Internal nodes of the tree provide cheap but biased estimates of a region's…

  61. [61] 2026-09-25 MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots

    arXiv:2609.29908v1 Announce Type: cross Abstract: Neural models can learn to generate various solutions to the inverse kinematics problem from data, but are usually limited to a single robot. We present MorphIK, a flow-matching model that solves inverse kinematics for revolute-joint-based kinematic chains it has never seen during training. The model uses a transformer architecture to encode the robot's morphology along…

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Consulting & Strategy practice, or Request an AI Roadmap.

Sources

  1. [1] MIT students gain a humanist lens on technical innovation in Tulsa, Oklahoma
  2. [2] SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions
  3. [3] Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms
  4. [4] DuplexDrama: A Synthesized Dialogue Dataset with Scenarios, Full-Duplex Behaviors, Expressive Speech, and Sound Events
  5. [5] Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding
  6. [6] Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
  7. [7] COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages
  8. [8] MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis
  9. [9] Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks
  10. [10] An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection
  11. [11] PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs
  12. [12] Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus
  13. [13] Temporal Taxation Compounds Under Post-Training Compression of Whisper Models
  14. [14] Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation
  15. [15] Script Choice in LLMs: Evidence for Late-Layer Commitment
  16. [16] Recovering the Zipfian Distribution in Unsupervised Term Discovery
  17. [17] Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research
  18. [18] CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
  19. [19] Reward Hacking Challenges Oversight of Autonomous Research Agents
  20. [20] PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
  21. [21] iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model
  22. [22] DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models
  23. [23] SGA: Uncertainty Quantification for Multi-Step Forecasting in Time Series Foundation Models
  24. [24] IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models
  25. [25] Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning
  26. [26] CyFM: Cylindrical Optimal Transport for Few-Step Complex-Valued Flow Matching
  27. [27] Uncovering Residential PV-EV Co-Adoption from Smart-Meter Data: Load Archetypes and Detection for Demand-Side Planning
  28. [28] Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling
  29. [29] CFD Correction of Open Tip Clearance Flow in a Compressor Cascade Using VAE Latent Space Adaptation
  30. [30] CARE: Condition-Aware Representation Regularization for Diffusion Models
  31. [31] SpaFactor: Lightweight Spatial Context-Aware Gene Program Modeling for Histology-to-Transcriptomics Inference
  32. [32] Leakage-Safe Machine Learning for Hydrogen Embrittlement Detection in 316L Stainless Steel: A Region-Held-Out Evaluation of Texture and Deep Features in SEM Micrographs
  33. [33] When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages
  34. [34] Foundations of Large Language Models
  35. [35] Time-Series Foundation Models That Understand Data Revisions
  36. [36] AFT Neural Function Approximators for 1D Nonlinear Force Laws
  37. [37] An explicit solution of the five-expert prediction PDE and the exact optimality set of COMB
  38. [38] Cost-Sensitive Online Window Size Selection for Portfolio Management
  39. [39] An Exploratory Ablation of a Small MLA–SSM Hybrid Language Model
  40. [40] Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
  41. [41] Learning and interpreting policies for simultaneous entanglement requests in quantum networks
  42. [42] Aftab: A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning
  43. [43] Driving Epidemic Models with AI Agents: the Epydemix Agent Framework
  44. [44] HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving
  45. [45] Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency
  46. [46] Answering Path Queries under Linear and Guarded Existential Rules
  47. [47] Training Object Permanence in World Models
  48. [48] To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech
  49. [49] PAWS: Policy-driven Agentic World Simulation
  50. [50] Pistis Technical Report
  51. [51] BaseCamp — An Agentic AI Framework for Automating DNA Sequencing Data Pipelines
  52. [52] TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment
  53. [53] DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs
  54. [54] SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design
  55. [55] Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents
  56. [56] Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation
  57. [57] When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills
  58. [58] Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning
  59. [59] When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI
  60. [60] Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits
  61. [61] MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots

Leave a comment

0.0/5