The Download: how your data is being used to train AI, and why chatbots aren’t doctors

Millions of images of passports, credit cards, birth certificates, and other documents containing personally identifiable information are likely included in one of the biggest open-source AI training sets, new research has found.

Thousands of images—including identifiable faces—were found in a small subset of DataComp CommonPool, a major AI training set for image generation scraped from the web. Because the researchers audited just 0.1% of CommonPool’s data, they estimate that the real number of images containing personally identifiable information, including faces and identity documents, is in the hundreds of millions.

The bottom line? Anything you put online can be and probably has been scraped. Read the full story.

—Eileen Guo

AI companies have stopped warning you that their chatbots aren’t doctors

AI companies have now mostly abandoned the once-standard practice of including medical disclaimers and warnings in response to health questions, new research has found. In fact, many leading AI models will now not only answer health questions but even ask follow-ups and attempt a diagnosis.

Such disclaimers serve an important reminder to people asking AI about everything from eating disorders to cancer diagnoses, the authors say, and their absence means that users of AI are more likely to trust unsafe medical advice. Read the full story.

—James O’Donnell

Source link

Judge rejects SEC and Ripple’s proposed settlement deal, upholds $125M penalty

FTI Consulting: A Hidden Gem or Just Fine?

Quick Access to Ruby Documentation

CDC Decides America’s Children Could Do With More Lead In Their Blood

The Great Unracking: Saying goodbye to the servers at our physical datacenter

The Download: how your data is being used to train AI, and why chatbots aren’t doctors

Leave a Reply Cancel reply

admin

Leave a Reply Cancel reply

Related Posts