Frontier AI Watch

AI Research

AI research papers and research-source updates across AI-for-science, scientific discovery, benchmarks, evaluation, safety and alignment, agents, model behavior, research infrastructure, and applied research signals.

Research papers, evaluation, and AI-for-science

AI research illustration
Google Research BlogPublished

GlucoFM: Foundation model for continuous glucose monitoring

AI research illustration
arXiv cs.AI Atom FeedPublished

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled…

AI research illustration
arXiv cs.AI Atom FeedPublished

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that…

AI research illustration
arXiv cs.LG Atom FeedPublished

Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role

Designing machine learning algorithms for wireless resource management is labour-intensive: the architecture, the loss function and the training recipe are all specified by hand. We demonstrate that this design layer can be surrendered to an autonomous agent in its entirety. We adopt the autoresearch protocol, in which an AI coding agent edits a training…

AI research illustration
arXiv cs.CL Atom FeedPublished

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

Civil infrastructure compliance checking has long relied on engineers manually reading legacy 2D plans; however, OCR-based automation strips away the geometry and layout essential for interpreting these plans. We present a Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG. It indexes and reasons directly over plan…

AI research illustration
arXiv cs.CL Atom FeedPublished

SwarmWorld: Stigmergic technological evolution in societies of language-model agents

Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions to accumulate into durable social organization. Language-model agents offer a new substrate for this process, yet most multi-agent systems rely on direct conversation, predefined roles, or centralized workflows. It remains unclear whether…

AI research illustration
Google DeepMind NewsPublished

Intelligent transcription with Gemini 3.5 Transcribe

Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

AI research illustration
Google Research BlogPublished

AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

AI research illustration
arXiv cs.LG Atom FeedPublished

What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable…

AI research illustration
Google DeepMind NewsPublished

From Atari to EVE Online: Building on 15 Years of AI Research in Games

Google DeepMind partners with game studios to prototype breakthrough AI gameplay.

AI research illustration
Microsoft Research AIPublished

Broadening access to Skala creates a faster path to predictive DFT

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft…

AI research illustration
Microsoft Research AIPublished

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .