Research
It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
Researchers find that it's alarmingly simple to bypass safety measures in some advanced AI models. This raises concerns about the reliability and security of these systems in real-world applications.
The State of Simulation for Physical AI: An Overview
Hugging Face is providing an overview of the current state of simulation for physical AI. This resource helps developers understand how to better train AI models in simulated environments before real-world deployment.
🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Xaira launched its X-Cell model designed for drug discovery using causal data. This approach aims to enhance the accuracy of predictions in pharmaceutical research.
Google Deepmind argues video generators already contain the world models computer vision has been missing
Google DeepMind claims that video generators are providing the world models that computer vision lacks. This could enhance how AI understands and interacts with visual data, leading to better applications in various fields.

AI text detectors struggle when language models mimic an author's style
AI text detectors are having a tough time when language models imitate an author's writing style. This means users might find it harder to distinguish between human and AI-generated content.

AI chatbots reading X-rays can be dangerously confident even when they're wrong
AI chatbots are showing overconfidence in diagnosing X-rays, often making incorrect assessments. This raises concerns about the reliability of AI in medical settings and the potential risks for patient care.

The risk of weather data sabotage is rising
Researchers are warning about the increasing risk of weather data sabotage, which could disrupt climate models and forecasts. This threat could lead to inaccurate weather predictions, impacting everything from agriculture to disaster preparedness.
Kimi K3, and what we can still learn from the pelican benchmark
Kimi K3 just introduced insights from the pelican benchmark to enhance AI performance. This update helps developers understand how to optimize their models based on real-world tasks.
AI Isn’t Smarter Than a Baby—Yet
Researchers find that AI still lacks the cognitive abilities of a human baby. This highlights the limitations of current AI systems in understanding and interacting with the world like humans do.
GPT-5.6 Sol reportedly disproves a 30-year-old statistics conjecture in 90 minutes after humans couldn't crack it
GPT-5.6 Sol just disproved a 30-year-old statistics conjecture in 90 minutes, a task humans struggled with for decades. This breakthrough showcases the model's advanced problem-solving capabilities and potential for tackling complex mathematical challenges.

What Anthropic’s latest AI discovery does—and doesn’t—show
Anthropic just revealed new insights about AI alignment and safety. Their findings could lead to better understanding and control of AI behaviors in future models.
Scientists’ Side Hustle? Using AI and Quantum Computing to Generate New Peptides
Scientists are using AI and quantum computing to generate new peptides. This approach could accelerate drug discovery and lead to more effective treatments.
The Download: Claude’s inner workings and OpenAI’s “super app”
Anthropic is revealing the inner workings of Claude, detailing its architecture and capabilities. This transparency aims to enhance user trust and understanding of how Claude operates in various applications.
Can AI answer the $3 trillion question?
TechCrunch explores how AI could tackle the $3 trillion question of global economic challenges. By leveraging advanced models, AI aims to provide insights that could reshape economic strategies and decision-making.
Why this CEO thinks video games make better training data than the internet
A CEO argues that video games provide superior training data for AI compared to the internet. This approach could enhance AI's learning by leveraging structured environments and rich interactions found in games.
[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI
Lilian Weng summarizes 35 papers on Harness Engineering for Reinforcement Learning. This compilation provides insights into improving the efficiency and effectiveness of reinforcement learning systems.
How Imperial College London is accelerating dementia research with a modern data platform
Imperial College London is using a modern data platform to speed up dementia research. This approach allows researchers to analyze vast amounts of data more efficiently, potentially leading to faster breakthroughs in treatment.
Building a World Map with only 500 bytes
Simon Willison just created a world map using only 500 bytes of data. This compact representation showcases how much can be achieved with minimal resources in data visualization.
Better Models: Worse Tools
Simon Willison critiques the trend of improving AI models while neglecting the tools that utilize them. He argues that better models without effective tools lead to less practical applications for users.
A 26,000-student study shows AI's hidden learning cost takes two full years to surface
A study with 26,000 students reveals that the hidden learning costs of AI tools take two years to become apparent. This means educators need to be aware of long-term impacts when integrating AI into learning environments.

A device that revives eyeballs from dead donors could make eye transplants possible
Researchers developed a device that can revive eyeballs from deceased donors, potentially making eye transplants feasible. This breakthrough could significantly improve the availability of donor organs for those in need of vision restoration.
UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do
UK's AI Security Institute reveals that standard benchmarks fail to accurately assess AI agents' capabilities. This means existing evaluations might not reflect the true potential and performance of these systems.

GPT and Claude failed Bridgewater's finance tests because the right answers were never public
GPT and Claude struggled with Bridgewater's finance tests since the correct answers weren't publicly available. This highlights the limitations of AI models when faced with proprietary knowledge and specific industry standards.

AIEWF Daily Dispatch: The great loops debate and the state of AI engineering
AI engineers are debating the effectiveness of different loop structures in AI programming. This discussion could lead to more efficient coding practices and improved AI performance.
Best practices for multi-turn reinforcement learning in Amazon SageMaker AI
AWS just shared best practices for multi-turn reinforcement learning in Amazon SageMaker AI. This guidance helps developers optimize their models for better performance in interactive environments.
Skill engineering and the case against one-shot AI design
Researchers are arguing against one-shot AI design, advocating for skill engineering instead. This approach aims to create more adaptable and capable AI systems by focusing on developing specific skills over a single, generalized model.
Teaching AI to run with the turbines
MIT researchers are training AI to optimize the operation of wind turbines. This could lead to more efficient energy production and reduced operational costs in wind farms.
Roundtables: Longevity’s Next Frontier: “Reprogramming” Your Body
Researchers are exploring ways to 'reprogram' the human body to extend longevity. This could lead to breakthroughs in health and aging, potentially changing how we approach life extension.
Inside Genebench-Pro
OpenAI just launched Genebench-Pro, a new framework for evaluating AI models on genetic data. This tool helps researchers assess model performance in genomics, making it easier to develop AI applications in healthcare.
Only three AI models finished above starting capital in a 500-day startup survival test
Only three AI models managed to exceed their starting capital in a 500-day startup survival test. This shows that while many AI models are promising, only a few can sustain long-term success in real-world applications.
