Focus on Faculty: Professor Hutson in ACM Symposium on Computer Science and the Law
A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset, co-authored by Professor Jevan Hutson and several colleagues in the proceedings of the ACM Symposium on Computer Science and Law, examines the privacy implications of large-scale web-scraped datasets used to train AI systems.
Through an empirical audit of DataComp CommonPool — one of the most widely downloaded AI training datasets, with over 2.2 million downloads — the authors identify significant quantities of personally identifiable information. These include credit card numbers, passport numbers, and an estimated 136,000 résumé images, despite the dataset’s sanitization efforts. Drawing on these findings, the paper analyzes existing privacy and data protection laws, highlights legal risks in current data curation practices, and argues for a reorientation of how “publicly available” information is defined within privacy and AI governance frameworks.