All posts
// / Blog

Every time I see a LinkedIn post about some fancy new model architecture, I think about the…

three weeks I once spent cleaning a client's customer data.

Three weeks. Not building models. Not fine-tuning. Just figuring out why 30% of the date fields were in five different formats and why someone had entered "yes" in a numerical column 847 times.

Nobody posts about this. But it's 90% of the job.

The #1 bottleneck in AI has never been models. It's data. And the engineers who understand Apache Spark, dbt, data quality frameworks like Great Expectations, and proper orchestration with Airflow — they're the ones who actually get ML projects across the finish line.

I've started telling junior engineers: if you want to be an ML engineer, spend your first year becoming a really good data engineer. Learn to build pipelines that are reliable, tested, and documented.

The model is the easy part. Getting clean, reliable data to that model? That's where the real skill lives.

#DataEngineering#BigData#ApacheSpark#MachineLearning#DataQuality#DataPipelines