Researchers from IBM Research, EPFL, and ETH Zurich have fine-tuned a Granite-based large language model on nearly 3,000 single-atom catalyst publications to generate synthesis protocols from user-defined prompts. The study illustrates a potential role for LLMs in materials science R&D pipelines and raises questions about scientific AI platforms, laboratory automation, and regulatory context. Commercial deployment pathways and quantitative performance metrics are not verifiable from the available source alone.
At AWS Summit New York 2026, Amazon Web Services announced a coordinated set of agent-focused products—including AWS Continuum, AWS Context, and expanded Bedrock AgentCore capabilities—highlighting a platform direction that places AI agents at the center of enterprise software delivery.
A joint whitepaper from Microsoft Advertising and Publicis Groupe, published June 18, 2025, finds that 75% of users report equivalent or better satisfaction with conversational AI search versus traditional search, and that 46% of consumers who notice ads in that environment report an improved experience. The findings are a reference point for discussions about search advertising and Microsoft's ad product strategy.
A Public First survey of more than 18,000 respondents across 15 countries suggests that people in key U.S.-allied markets increasingly view China as the world’s AI leader, while American confidence in AI is weakening over resource use, labor displacement, and information reliability. The result matters as a signal that perception can influence procurement, regulation, and go-to-market strategy.
Google DeepMind says a randomized controlled trial across 12 schools in Sierra Leone and 1,763 junior secondary students found that guided AI learning lifted mathematics scores by 0.258 standard deviations. The result reinforces a broader shift in edtech: AI tools will increasingly be judged by learning outcomes, not by novelty or usage alone.
Stanford University's Center for Artificial Intelligence in Medicine & Imaging is conducting prospective real-time clinical validation studies of AI models for medical imaging. This is a systematic approach to evaluating the safety and effectiveness of AI tools in actual clinical settings, helping build evidence that can inform regulatory review and healthcare deployment.
Nature has introduced a benchmark of expert-level academic questions designed to assess the scholarly capabilities of AI systems. The benchmark aims to move beyond existing evaluation tools by testing advanced reasoning abilities required in real research environments. The research community anticipates this will enable more accurate measurement of AI models' scientific problem-solving capacity.
Anthropic has published a research blog post proposing that biological data infrastructure become more agent-friendly. The company outlines deterministic execution layers, reliable access to biological databases, and agent-accessible context engines to support scientific discovery.
OpenAI has released PaperBench, a new benchmark designed to measure AI agents' ability to replicate state-of-the-art research. The benchmark evaluates how accurately AI systems can reproduce empirical contributions from published papers, establishing a new standard for automated scientific research capabilities.