A company asked me to build a model using data scraped from social media profiles without users'…
knowledge.
I said no.
Not because it was technically difficult. Because it was wrong.
The data was public (people had posted it online). The use case wasn't malicious (predicting consumer preferences). But the users had no idea their data would be used this way, and they'd never consented to it.
"But the data is publicly available" is the tech industry's favorite justification for ethical shortcuts. Public availability doesn't equal consent for any use.
Where I draw the line: if users would be surprised or uncomfortable knowing their data was used this way, it's not ethical to use it. Full stop.
Ethical alternatives that work: first-party data collected with clear consent. Aggregated, anonymized datasets. Synthetic data generated from statistical distributions. Licensed datasets from providers who handle consent.
Yes, these are more expensive and more work. Yes, they limit what you can build. But the alternative is building AI systems on a foundation of broken trust.
Every AI engineer makes choices about data that have real consequences for real people. Those choices matter more than the model architecture or the accuracy score.
Build AI you'd be comfortable explaining to the people whose data you used.