What Happened
Two recent research publications address different gaps in AI planning. The Lincoln AI Computing Survey now tracks more than 120 commercial accelerators, up from 57 in its first survey. It compares publicly reported peak performance and power across CPUs, GPUs, ASICs, FPGAs and dataflow systems, while examining how architecture affects performance. Earlier work…
What Happened
Recent AI papers point less to a single model breakthrough than to improvements in the systems around models: culturally appropriate data, evidence retrieval, tool controls, evaluation and selective use of compute. Most results are research benchmarks, not production guarantees.
Data provenance matters. WAON’s roughly 155 million native Japanese image–text examples outperformed translated English…
What Happened
Recent AI papers point to a common problem: strong benchmark results do not necessarily translate into reliable decisions. An audit of the IBM Telco churn benchmark found that applying SMOTE before the train/test split inflated churn-class F1 by 13.1 percentage points. Under the authors’ retention-cost assumptions, the cost-optimal decision threshold was roughly 5–10…
What Happened
Several recent papers point to the same engineering lesson: improving an AI model is not the same as improving the decisions or actions of an application.
Agent failures can become useful training data. The Agent Error Dataset contains 50,228 error–diagnosis pairs with execution traces. In matched replays, proposed corrections raised verifier pass rates…
What Happened
Recent papers point to a practical shift: agent performance depends not only on the base model, but also on how teams train for failures, construct environments, manage execution, and verify outputs.
Failure data can improve recovery. The Agent Error Dataset contains 50,228 error–diagnosis pairs with execution traces. In matched replays, first-proposal corrections raised…
What Happened
A large set of recent papers and lab releases advance evaluation, robustness, retrieval hygiene, agent steering, domain adaptation, benchmarks and applied systems across production-relevant areas. Key highlights:
Operational forecasting: Microsoft Research built an end‑to‑end ML pipeline that produces 30–60 minute, location‑specific geomagnetic risk forecasts for ~67k U.S. substations with sub‑second inference…
Findings [1] 2026-09-29 Introducing Quine: An AI research system designed for the complexity of biology At a glance Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interactive harness connecting models, scientific tools, literature, and researchers. In collaboration with researchers at the Broad Institute…
Findings [1] 2026-09-28 One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact On July 24, 2025, Microsoft Research Asia – Singapore (MSRA – Singapore) opened its doors as Microsoft’s first research lab in Southeast Asia. The launch built on more than two decades of collaboration…
Findings [1] 2026-09-25 MIT students gain a humanist lens on technical innovation in Tulsa, Oklahoma When MIT mechanical engineering student Daphne Wang arrived for her internship at the Muscogee Creek Nation Department of Health in Tulsa, Oklahoma last summer, she expected to be writing code. To her surprise, she found herself analyzing tribal history…
Findings [1] 2026-09-24 Estimating suicide risk from text When people reach out during a mental health crisis, a top priority for counselors is identifying those with a high risk of suicide. The distressed person’s language holds critical clues, and a new tool developed by scientists at MIT’s McGovern… However, they say they often use…
Findings [1] 2026-09-23 Offloaded inference for real-world physical AI robotics At a glance Challenges a core assumption in robotics AI: Our research shows that running physical AI inference exclusively on onboard GPUs can limit robot performance, battery life, and scalability, and that offloading inference to edge or cloud GPUs can… AI Testing and Evaluation:…
Findings [1] 2026-09-22 Poitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research Patricia and James Poitras ’63, longtime MIT supporters, have launched a fellowship program for graduate students and postdocs studying major mental illness, expanding their MIT philanthropy to directly support early-career scientists. The commitment establishes 50 two-year…