Data Pipelines and Storage
Training datasets routinely run into terabytes, and the cost of storing, versioning, and moving that data between environments compounds quietly. Feature stores in particular tend to retain far more historical data than any active model actually queries.