GitHub - rahulbakshee/ml: Machine Learning | Exploratory Data Analysis

train_test_split: We will try to classify the iris dataset using K-nearest neighbours algorithm from sklearn.neighbours.KNeighborsClassifier. For the train and test split we use sklearn.cross_validation.train_test_split.
Bias vs Variance source: https://www.coursera.org/learn/machine-learning/supplement/VTe37/the-problem-of-overfitting

Underfitting, or high bias, is when the form of our hypothesis function h maps poorly to the trend of the data. It is usually caused by a function that is too simple or uses too few features. At the other extreme, overfitting, or high variance, is caused by a hypothesis function that fits the available data but does not generalize well to predict new data. It is usually caused by a complicated function that creates a lot of unnecessary curves and angles unrelated to the data.

This terminology is applied to both linear and logistic regression. There are two main options to address the issue of overfitting:

Reduce the number of features:

Manually select which features to keep. Use a model selection algorithm (studied later in the course). 2) Regularization

Keep all the features, but reduce the magnitude of parameters θj. Regularization works well when we have a lot of slightly useful features.

Precison vs Recall
classification-optimization code
regression-optimizatio code

Name		Name	Last commit message	Last commit date
Latest commit History 47 Commits
AnalyticsVidhya		AnalyticsVidhya
Kggle-AI-Ethics		Kggle-AI-Ethics
Bias vs Variance		Bias vs Variance
README.md		README.md
combine-rare-categories-into-one.py		combine-rare-categories-into-one.py
cross-entropy.py		cross-entropy.py
how-to-import-kaggle-dataset-into-colab		how-to-import-kaggle-dataset-into-colab
hyperparameter-optimization-classification.py		hyperparameter-optimization-classification.py
train_test_split.ipynb		train_test_split.ipynb
xcode__f1-mcc.py		xcode__f1-mcc.py
xcode__xgb.py		xcode__xgb.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

AnalyticsVidhya

AnalyticsVidhya

Kggle-AI-Ethics

Kggle-AI-Ethics

Bias vs Variance

Bias vs Variance

README.md

README.md

combine-rare-categories-into-one.py

combine-rare-categories-into-one.py

cross-entropy.py

cross-entropy.py

how-to-import-kaggle-dataset-into-colab

how-to-import-kaggle-dataset-into-colab

hyperparameter-optimization-classification.py

hyperparameter-optimization-classification.py

train_test_split.ipynb

train_test_split.ipynb

xcode__f1-mcc.py

xcode__f1-mcc.py

xcode__xgb.py

xcode__xgb.py

Repository files navigation

Contents

About

Releases

Packages

Languages

rahulbakshee/ml

Folders and files

Latest commit

History

Repository files navigation

Contents

About

Topics

Resources

Stars

Watchers

Forks

Languages