There are no lossless transformations of natural-language text
Simon Willison reveals that lossless transformations of natural-language text are impossible. This means that some information will always be lost when converting text formats, impacting how we handle text data.
More in Research
Stealing Reasoning Traces from Proprietary LLM APIs
Researchers are extracting reasoning traces from proprietary LLM APIs to analyze their decision-making processes. This could lead to better understanding and improvements in AI model transparency and accountability.
Old OCR text cripples language model training, and FineBooks wants to fix that at scale
FineBooks is tackling the issue of outdated OCR text that hampers language model training. Their solution aims to enhance the quality of training data, which could lead to better-performing AI models.

Readers rate AI-generated short stories higher than human ones until they learn a machine wrote them
Researchers found that readers initially rate AI-generated short stories higher than those written by humans. However, ratings drop significantly once readers discover the stories were machine-generated.

The Download: a censorship conspiracy theory and the first virus created by AI
Researchers just revealed the first virus created by AI, showcasing its potential to generate malicious code. This raises concerns about cybersecurity and the need for better defenses against AI-driven threats.