Frontier AI Watch

AI Research

AI research papers and research-source updates across AI-for-science, scientific discovery, benchmarks, evaluation, safety and alignment, agents, model behavior, research infrastructure, and applied research signals.

Research papers, evaluation, and AI-for-science

AI research illustration
arXiv cs.CL Atom FeedPublished

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency.…

AI research illustration
arXiv cs.AI Atom FeedPublished

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain scattered across optimization histories,…

AI research illustration
arXiv cs.CL Atom FeedPublished

TTPO: Test-Time Policy Optimization

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet…

AI research illustration
arXiv cs.AI Atom FeedPublished

SWE-Prime: Fewer Trajectories, Better Performance

To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective,…

AI research illustration
arXiv cs.LG Atom FeedPublished

Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study

Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflects contemporary critical care. No alternative learned directly from patient trajectories is in routine use. We conducted a retrospective two-cohort study on a total of 29,116 and…

AI research illustration
arXiv cs.LG Atom FeedPublished

Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling

Friend recommendation is inherently graph-structured: the relevance of a potential connection depends on multi-hop social context rather than user attributes alone. However, deploying message-passing GNNs on a production-scale social graph with hundreds of millions of users and tens of billions of edges requires addressing numerous modeling and systems…

AI research illustration
Google Research BlogPublished

Planetary prediction engine: Automating global models via Earth AI

AI research illustration
Google DeepMind NewsPublished

Gemini Omni 1.1 Flash lets you build with more control

Gemini Omni 1.1 Flash lets you build with more control

AI research illustration
Google DeepMind NewsPublished

Piloting the world's first double-blind AI evaluations

Piloting the world's first double-blind AI evaluations

AI research illustration
Google Research BlogPublished

GlucoFM: Foundation model for continuous glucose monitoring

AI research illustration
Microsoft Research AIPublished

Broadening access to Skala creates a faster path to predictive DFT

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft…

AI research illustration
Microsoft Research AIPublished

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .