AI benchmarks have a trust problem and Google wants to fix it

Google is addressing trust issues in AI benchmarks by developing new evaluation methods. This shift aims to improve the reliability of AI performance assessments for developers and users alike.
More in Research
The Download: a secretive antiaging drug and joining virtual power plants
Researchers are developing a secretive anti-aging drug that shows promise in extending lifespan. This could revolutionize how we approach aging and longevity treatments.
The Humanoids at China’s Robot Games Were Faster Than Usain Bolt—but I’m More Impressed by Their Tweezer Mastery
Humanoid robots at China's Robot Games outpaced Usain Bolt in speed tests. Their impressive dexterity with tweezers showcases advancements in robotic manipulation skills.
How to evaluate LLMs before production
GitHub just released a guide on evaluating large language models (LLMs) before deploying them. This helps developers ensure their models meet production standards and perform reliably in real-world applications.
Addressing a sticking point in sustainable adhesives
Researchers at MIT are developing a new sustainable adhesive that uses plant-based materials. This innovation could reduce reliance on harmful chemicals in manufacturing processes.