ML Process:
- Sourcing
- Cleaning
- Feature Engineering
- Machine Learning
- Prediction (some label)
Datasets for:
- books
- banking data
- titanic passengers (credit for: https://github.com/ranvirm/scala-spark-titanic-example-project/blob/master/src/main/scala/DataCleaner.scala)
Part for getting bussiness knowledge about particular scheme and production data. Business and scale understanding.
- How much user?
- Which social groups?
- For what they pay?
- Which countries?