Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
Lukas Petersson and Axel Backlund from Andon Labs are launching a new evaluation framework called Reality. This tool aims to enhance the assessment of AI models by providing more accurate performance metrics.
More in Models
Alibaba's Qwen takes on Kimi K3 with open-weight Qwen 3.8, says model is "second only to Fable 5"
Alibaba just launched the open-weight Qwen 3.8 model, claiming it's 'second only to Fable 5'. This upgrade enhances their competitive edge in the AI model landscape.

Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math
Moonshot's Kimi K3 outperforms Fable 5 in generating frontend code but struggles with complex math tasks. Users can expect better performance in web development, but may face limitations in advanced calculations.

Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost
Open-weight models are now achieving frontier cyber performance that was only available a few months ago, and they do it at a much lower cost. This shift makes advanced capabilities more accessible for various applications in cybersecurity.

[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
Kimi just released the K3 2.8T-A50B, the largest open model to date, priced competitively at Sonnet 5 levels. This gives developers access to powerful capabilities without breaking the bank.