<img src="images/python.png">

 - **Python is an interpreted high-level programming language for general-purpose programming**
    
 - **Less is More - making programs work in the least lines of code**
 
 - **Python’s Compatibility with Hadoop - The PyDoop package offers access to the HDFS API for Hadoop and hence allows to write Hadoop MapReduce programs and applications**
 
 - **Ease of Learning - ample learning resources, readable code and large community**
 
 - **Powerful Packages -  Has a powerful set of packages for a wide range of data science and analytical needs**
 
 - **Data Visualization - Though Python toughest competitor R is better when it comes to data visualization**

### <font color=blue>Basic data types - Numbers</font>

** Integers and floats work as you would expect from other languages **

In [None]:
import warnings
warnings.filterwarnings("ignore", category=DeprecationWarning) 

In [None]:
x = 5
print(type(x)) # Prints "<class 'int'>"
print(x)       # Prints "3"
print(x + 5)   # Addition; prints "4"
print(x - 2)   # Subtraction; prints "2"
print(x * 3)   # Multiplication; prints "6"
print(x ** 2)  # Exponentiation; prints "9"
x = x  + 1
x += 1
print(x)  # Prints "4"
x *= 2
x = x * 2
print(x)  # Prints "8"
y = 3.543
print(type(y)) # Prints "<class 'float'>"
print(y, y + 1, y * 2, y ** 2) # Prints "2.5 3.5 5.0 6.25"

### <font color=blue>Basic data types - Booleans</font>

** Python implements all of the usual operators for Boolean logic, but uses English words rather than symbols **

In [None]:
t = True
f = False
print(type(t)) # Prints "<class 'bool'>"
print(t and f) # Logical AND; prints "False"
print(t or f)  # Logical OR; prints "True"
print(not t)   # Logical NOT; prints "False"
print(t != f)  # Logical XOR; prints "True"

### <font color=blue>Basic data types - Strings</font>

** Python has great support for strings and String objects have a bunch of useful methods **

In [None]:
hello = 'hello'    # String literals can use single quotes
world = "world"    # or double quotes; it does not matter.
print(hello)       # Prints "hello"
print(len(hello))  # String length; prints "5"
hw = hello + '          ' + world  # String concatenation
print(len(hw))
print(hw)  # prints "hello world"
hw12 = '%s %s %d' % (hello, world, 12)  # sprintf style string formatting
print(hw12)  # prints "hello world 12"

In [None]:
s = "hello"
print(s.capitalize())  # Capitalize a string; prints "Hello"
print(s.upper())       # Convert a string to uppercase; prints "HELLO"
print(s.rjust(7))      # Right-justify a string, padding with spaces; prints "  hello"
print(s.center(7))     # Center a string, padding with spaces; prints " hello "
print(s.replace('l', '(ell)'))  # Replace all instances of one substring with another;
                                # prints "he(ell)(ell)o"
print('  world '.strip())  # Strip leading and trailing whitespace; prints "world"
len('  world '.strip())

### <font color=blue>Basic data types - Lists</font>

** A list is the Python equivalent of an array, but is resizeable and can contain elements of different types **

** In addition to accessing list elements one at a time, Python provides concise syntax to access sublists; this is known as slicing **

** You can loop over the elements **

** You can make code simpler using a list comprehension **

In [None]:
xs = [3, 1, 2]    # Create a list
print(xs, xs[0])  # Prints "[3, 1, 2] 2"
print(xs[-1])     # Negative indices count from the end of the list; prints "2"
xs[2] = 'foo'     # Lists can contain elements of different types
print(xs)         # Prints "[3, 1, 'foo']"
xs.append('bar')
xs.append(1.2) # Add a new element to the end of the list
print(xs)         # Prints "[3, 1, 'foo', 'bar']"
x = xs.pop()      # Remove and return the last element of the list
print(x, xs)      # Prints "bar [3, 1, 'foo']"

In [None]:
nums = list(range(5))     # range is a built-in function that creates a list of integers
print(nums)               # Prints "[0, 1, 2, 3, 4]"
print(nums[2:3])          # Get a slice from index 2 to 4 (exclusive); prints "[2, 3]"
print(nums[2:])           # Get a slice from index 2 to the end; prints "[2, 3, 4]"
print(nums[:2])           # Get a slice from the start to index 2 (exclusive); prints "[0, 1]"
print(nums[:])            # Get a slice of the whole list; prints "[0, 1, 2, 3, 4]"
print(nums[:-1])          # Slice indices can be negative; prints "[0, 1, 2, 3]"
nums[2:4] = [8, 9]        # Assign a new sublist to a slice
print(nums)               # Prints "[0, 1, 8, 9, 4]"

In [None]:
animals = ['cat', 'dog', 'monkey']
for a in animals:
    print(a)
# Prints "cat", "dog", "monkey", each on its own line.

In [None]:
animals = ['cat', 'dog', 'monkey']
for idx, animal in enumerate(animals):
    print(idx)
    print('#%d: %s' % (idx + 1, animal))
# Prints "#1: cat", "#2: dog", "#3: monkey", each on its own line

In [None]:
nums = [0, 1, 2, 3, 4]
squares = []
squares = list()
for x in nums:
    if x % 2 == 0:
        squares.append(x ** 2)
print(squares)   # Prints [0, 1, 4, 9, 16]

In [None]:
nums = [0, 1, 2, 3, 4]
squares = [x ** 2 for x in nums]
print(squares)   # Prints [0, 1, 4, 9, 16]

In [None]:
nums = [0, 1, 2, 3, 4]
even_squares = [x ** 2 for x in nums if x % 2 == 0]
print(even_squares)  # Prints "[0, 4, 16]"

### <font color=blue>Basic data types - Dictionaries</font>

** A dictionary stores (key, value) pairs, similar to a Map in Java or an object in Javascript. **

** It is easy to iterate over the keys in a dictionary **

** Access to keys and their corresponding values through items **

** Dictionary comprehensions  are similar to list comprehensions, but allow you to easily construct dictionaries. **

In [None]:
d = {'cat': 'cute', 'dog': 'cute'}  # Create a new dictionary with some data
print(d['cat'])       # Get an entry from a dictionary; prints "cute"
print('horse' in d)     # Check if a dictionary has a given key; prints "True"
d['fish'] = 'wet'     # Set an entry in a dictionary
print(d['fish'])      # Prints "wet"
#print(d['monkey'])  # KeyError: 'monkey' not a key of d
print(d.get('monkey', 'this is not here'))  # Get an element with a default; prints "N/A"
print(d.get('fish', 'N/A'))    # Get an element with a default; prints "wet"
del d['fish']         # Remove an element from a dictionary
print(d.get('fish', 'N/A')) # "fish" is no longer a key; prints "N/A"

In [None]:
d = {'person': 2, 'cat': 4, 'spider': 8}
for animal in d:
    legs = d[animal]
    print('A %s has %d legs' % (animal, legs))
# Prints "A person has 2 legs", "A cat has 4 legs", "A spider has 8 legs"

In [None]:
nums = [0, 1, 2, 3, 4]
even_num_to_square = {x: x ** 2 for x in nums if x % 2 == 0}
print(even_num_to_square)  # Prints "{0: 0, 2: 4, 4: 16}"

### <font color=blue>Basic data types - Sets</font>

** A set is an unordered collection of distinct elements. **

** Iterating over a set has the same syntax as iterating over a list **

** However since sets are unordered, you cannot make assumptions about the order in which you visit the elements of the set **

** Like lists and dictionaries, we can easily construct sets using set comprehensions **

In [None]:
animals = {'cat', 'dog', 'horse'}
print('cat' in animals)   # Check if an element is in a set; prints "True"
print('fish' in animals)  # prints "False"
animals.add('fish')       # Add an element to a set
print('fish' in animals)  # Prints "True"
print(len(animals))       # Number of elements in a set; prints "3"
animals.add('cat')        # Adding an element that is already in the set does nothing
print(len(animals))       # Prints "3"
animals.remove('cat')     # Remove an element from a set
print(len(animals))       # Prints "2"

In [None]:
animals = {'cat', 'dog', 'fish'}
for idx, animal in enumerate(animals):
    print('#%d: %s' % (idx + 1, animal))
# Prints "#1: fish", "#2: dog", "#3: cat"

In [None]:
from math import sqrt
nums = {int(sqrt(x)) for x in range(30)}
print(range(30))
print(nums)  # Prints "{0, 1, 2, 3, 4, 5}"

### <font color=blue>Basic data types - Tuples</font>

** A tuple is an (immutable) ordered list of values.  **

** A tuple is in many ways similar to a list; one of the most important differences is that tuples can be used as keys in dictionaries and as elements of sets, while lists cannot. **

In [None]:
d = {(x, x + 1): x for x in range(10)}  # Create a dictionary with tuple keys
print(d)
t = (5, 6)        # Create a tuple
print(type(t))    # Prints "<class 'tuple'>"
print(d[t])       # Prints "5"
print(d[(1, 2)])  # Prints "1"

### <font color=blue>Functions</font>

** Python functions are defined using the def keyword **

** We can define functions to take optional keyword arguments **

In [None]:
def sign(x):
    if x > 0:
        return 'positive'
    elif x < 0:
        return 'negative'
    else:
        return 'zero'

for x in [-1, 0, 1, -2, 5, -90]:
    print(sign(x))
# Prints "negative", "zero", "positive"

In [None]:
def hello(name, loud=False):
    if loud:
        print('HELLO, %s!' % name.upper())
    else:
        print('Hello, %s' % name)

hello('Bob') # Prints "Hello, Bob"
hello('Fred', loud=True)  # Prints "HELLO, FRED!"

<img src="images/pandas.png">

- ** Pandas is an open-source Python Library **
- ** Providing high-performance data manipulation and analysis. **
- ** The name Pandas is derived from the word Panel Data. **
- ** Fast and efficient DataFrame object **
- ** Default and customized indexing. **
- ** Tools for loading data into in-memory data objects from  different file formats. **
- ** Handling of missing data. **
- ** Label-based slicing of large data sets. **
- ** Series: 1D labeled homogeneous array, size immutable. **
- ** DataFrame: General 2D labeled, size-mutable tabular  structure with potentially heterogeneously typed columns. **
- ** Panel: General 3D labeled, size-mutable array. **

In [None]:
import pandas as pd
# create simple dataset of people
data = {'Name': ['John', 'Anna', 'Peter', 'Linda'],
       'Location': ['New York', 'Paris', 'Berlin', 'London'],
       'Age': [24, 13, 53, 33]
       }
data_pandas = pd.DataFrame(data)
data_pandas

In [None]:
data_pandas.columns

In [None]:
data_pandas['Name']

In [None]:
list(data_pandas['Name'])

In [None]:
data_pandas['Name'] == 'Anna'

In [None]:
data_pandas[data_pandas['Name'] == 'Anna']

In [None]:
import pandas as pd
# create simple dataset of people
data = {'Name': ['John', 'Anna', 'Peter', 'Linda'],
       'Location': ['New York', 'Paris', 'Berlin', 'London'],
       'Age': [24, 13, None , 33]
       }
data_pandas = pd.DataFrame(data)
data_pandas

In [None]:
data_pandas.fillna(0, inplace=True)
data_pandas

In [None]:
data_pandas.shape

In [None]:
data_pandas.to_excel('data/pd_example.xlsx')

In [None]:
data_pandas.to_csv('data/pd_example.csv')

In [None]:
excel_file = 'data/movies.xls'
movies = pd.read_excel('data/movies.xls')

In [None]:
movies.shape

In [None]:
movies.head()

In [None]:
movies.fillna(3, inplace=True)
movies.head()

In [None]:
sorted_by_gross = movies.sort_values(['Gross Earnings'], ascending=False)

In [None]:
sorted_by_gross["Gross Earnings"].head(10)

In [None]:
movies.describe()

In [None]:
movies["Year"].median()

In [None]:
movies["Net Earnings"] = movies["Gross Earnings"] - movies["Budget"]

In [None]:
sorted_movies = movies[['Net Earnings']].sort_values(['Net Earnings'], ascending=[False])

In [None]:
sorted_movies.head()

In [None]:
movies_subset = movies[['Year', 'Gross Earnings']]
movies_subset.head()

In [None]:
movies.to_csv('data/movies.csv')

In [None]:
excel_file = 'data/movies.csv'
movies = pd.read_csv(excel_file)

In [None]:
movies.shape

In [None]:
movies.head()

<img src="images/numpy.jpeg">

- ** NumPy is a package for linear algebra and advanced  mathematics in Python. **

- ** It provides a fast implementation of multidimensional  numerical arrays (C / FORTRAN), vectors, matrices,  tensors and operations on them. **

- ** NumPy is the core library for scientific computing in Python. **

- ** It provides a high-performance multidimensional array object, and tools for working with these arrays. **

- ** Essential for Machine Learning tasks. **

### <font color=blue>Arrays</font>

** A numpy array is a grid of values, all of the same type, and is indexed by a tuple of nonnegative integers. **

** The number of dimensions is the rank of the array. **

** The shape of an array is a tuple of integers giving the size of the array along each dimension. **

** We can initialize numpy arrays from nested Python lists, and access elements using square brackets. **

In [None]:
import numpy as np

a = np.array([1, 2, 3])   # Create a rank 1 array
print(type(a))            # Prints "<class 'numpy.ndarray'>"
print(a.shape)            # Prints "(3,)"
print(a[0], a[1], a[2])   # Prints "1 2 3"
a[0] = 5                  # Change an element of the array
print(a)                  # Prints "[5, 2, 3]"

b = np.array([[1,2,3],[4,5,6]])    # Create a rank 2 array
print(b)
print(b.shape)                     # Prints "(2, 3)"
print(b[0, 0], b[0, 1], b[1, 0])   # Prints "1 2 4"

### <font color=blue>Built-in Arrays</font>

** Numpy also provides many functions to create arrays. **

In [None]:
import numpy as np

a = np.zeros((2,2))   # Create an array of all zeros
print(a)              # Prints "[[ 0.  0.]
                      #          [ 0.  0.]]"
    
b = np.ones((3,2))    # Create an array of all ones
print(b)              # Prints "[[ 1.  1.]]"

c = np.full((2,2),9)  # Create a constant array
print(c)               # Prints "[[ 7.  7.]
                       #          [ 7.  7.]]"

d = np.eye(2)         # Create a 2x2 identity matrix
print(d)              # Prints "[[ 1.  0.]
                      #          [ 0.  1.]]"

e = np.random.random((2,3))  # Create an array filled with random values
print(e)                     # Might print "[[ 0.91940167  0.08143941]
                             #               [ 0.68744134  0.87236687]]"

### <font color=blue>Array indexing</font>

** Numpy offers several ways to index into arrays. **

** Similar to Python lists, numpy arrays can be sliced. **

** Since arrays may be multidimensional, you must specify a slice for each dimension of the array. **

In [6]:
import numpy as np

# Create the following rank 2 array with shape (3, 4)
# [[ 1  2  3  4]
#  [ 5  6  7  8]
#  [ 9 10 11 12]]
a = np.array([[1,2,3,4], [5,6,7,8], [9,10,11,12]])
print(a)

# Use slicing to pull out the subarray consisting of the first 2 rows
# and columns 1 and 2; b is the following array of shape (2, 2):
# [[2 3]
#  [6 7]]
#b = np.array([[]])
b = a[:2, 1:3]
b = np.array(a[:2, 1:3])
print(b)
# A slice of an array is a view into the same data, so modifying it
# will modify the original array.
print(a[0, 1])   # Prints "2"
b[0, 0] = 77     # b[0, 0] is the same piece of data as a[0, 1]
print(b)
print(a[0, 1])   # Prints "77"
print(a)

[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
[[2 3]
 [6 7]]
2
[[77  3]
 [ 6  7]]
2
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]


In [7]:
import numpy as np
np.sin?

In [8]:
a = 5
help(a)

Help on int object:

class int(object)
 |  int(x=0) -> integer
 |  int(x, base=10) -> integer
 |  
 |  Convert a number or string to an integer, or return 0 if no arguments
 |  are given.  If x is a number, return x.__int__().  For floating point
 |  numbers, this truncates towards zero.
 |  
 |  If x is not a number or if base is given, then x must be a string,
 |  bytes, or bytearray instance representing an integer literal in the
 |  given base.  The literal can be preceded by '+' or '-' and be surrounded
 |  by whitespace.  The base defaults to 10.  Valid bases are 0 and 2-36.
 |  Base 0 means to interpret the base from the string as an integer literal.
 |  >>> int('0b100', base=0)
 |  4
 |  
 |  Methods defined here:
 |  
 |  __abs__(self, /)
 |      abs(self)
 |  
 |  __add__(self, value, /)
 |      Return self+value.
 |  
 |  __and__(self, value, /)
 |      Return self&value.
 |  
 |  __bool__(self, /)
 |      self != 0
 |  
 |  __ceil__(...)
 |      Ceiling of an Integral retur

In [None]:
import numpy as np
# Create the following rank 2 array with shape (3, 4)
# [[ 1  2  3  4]
#  [ 5  6  7  8]
#  [ 9 10 11 12]]
a = np.array([[1,2,3,4], [5,6,7,8], [9,10,11,12]])

# Two ways of accessing the data in the middle row of the array.
# Mixing integer indexing with slices yields an array of lower rank,
# while using only slices yields an array of the same rank as the
# original array:
row_r1 = a[1, :]    # Rank 1 view of the second row of a
row_r2 = a[1:2, :]  # Rank 2 view of the second row of a
print(row_r1, row_r1.shape)  # Prints "[5 6 7 8] (4,)"
print(row_r2, row_r2.shape)  # Prints "[[5 6 7 8]] (1, 4)"

# We can make the same distinction when accessing columns of an array:
col_r1 = a[:, 1]
col_r2 = a[:, 1:2]
print(col_r1, col_r1.shape)  # Prints "[ 2  6 10] (3,)"
print(col_r2, col_r2.shape)  # Prints "[[ 2]
                             #          [ 6]
                             #          [10]] (3, 1)"

In [None]:
import numpy as np

a = np.array([[1,2], [3, 4], [5, 6]])
print(a)
bool_idx = (a > 2)   # Find the elements of a that are bigger than 2;
                     # this returns a numpy array of Booleans of the same
                     # shape as a, where each slot of bool_idx tells
                     # whether that element of a is > 2.

print(bool_idx)      # Prints "[[False False]
                     #          [ True  True]
                     #          [ True  True]]"

# We use boolean array indexing to construct a rank 1 array
# consisting of the elements of a corresponding to the True values
# of bool_idx

print(a[bool_idx])  # Prints "[3 4 5 6]"

# We can do all of the above in a single concise statement:
print(a[a > 2])     # Prints "[3 4 5 6]"

### <font color=blue>Datatypes</font>

** Every numpy array is a grid of elements of the same type. **

** Numpy provides a large set of numeric datatypes that you can use to construct arrays. **

** Numpy tries to guess a datatype when you create an array, but functions that construct arrays usually also include an optional argument to explicitly specify the datatype. **

In [None]:
import numpy as np

x = np.array([1, 2])   # Let numpy choose the datatype
print(x.dtype)         # Prints "int64"

x = np.array([1.0, 2.0])   # Let numpy choose the datatype
print(x.dtype)             # Prints "float64"

x = np.array([1, 2], dtype=np.float64)   # Force a particular datatype
print(x)
print(x.dtype)                         # Prints "int64"

### <font color=blue>Array Math</font>

** Basic mathematical functions operate elementwise on arrays, and are available both as operator overloads and as functions in the numpy module **

In [None]:
import numpy as np
x = np.array([[1,2],[3,4]], dtype=np.float64)
y = np.array([[5,6],[7,8]], dtype=np.float64)

# Elementwise sum; both produce the array
# [[ 6.0  8.0]
#  [10.0 12.0]]
print(x + y)
print(np.add(x, y))

# Elementwise difference; both produce the array
# [[-4.0 -4.0]
#  [-4.0 -4.0]]
print(x - y)
print(np.subtract(x, y))

# Elementwise product; both produce the array
# [[ 5.0 12.0]
#  [21.0 32.0]]
print(x * y)
print(np.multiply(x, y))

# Elementwise division; both produce the array
# [[ 0.2         0.33333333]
#  [ 0.42857143  0.5       ]]
print(x / y)
print(np.divide(x, y))

# Elementwise square root; produces the array
# [[ 1.          1.41421356]
#  [ 1.73205081  2.        ]]
print(np.sqrt(x))

** We use the dot function to compute inner products of vectors, to multiply a vector by a matrix, and to multiply matrices. **

** dot is available both as a function in the numpy module and as an instance method of array objects. **

In [None]:
import numpy as np

x = np.array([[1,2],[3,4]])
y = np.array([[5,6],[7,8]])

v = np.array([9,10])
w = np.array([11, 12])

# Inner product of vectors; both produce 219
print(v.dot(w))
print(np.dot(v, w))

# Matrix / vector product; both produce the rank 1 array [29 67]
print(x.dot(v))
print(np.dot(x, v))

# Matrix / matrix product; both produce the rank 2 array
# [[19 22]
#  [43 50]]
print(x.dot(y))
print(np.dot(x, y))

In [None]:
import numpy as np

x = np.array([[1,2],[3,4]])

print(np.sum(x))  # Compute sum of all elements; prints "10"
print(np.sum(x, axis=0))  # Compute sum of each column; prints "[4 6]"
print(np.sum(x, axis=1))  # Compute sum of each row; prints "[3 7]"

import numpy as np

x = np.array([[1,2], [3,4]])
print(x)    # Prints "[[1 2]
            #          [3 4]]"
print(x.T)  # Prints "[[1 3]
            #          [2 4]]"

# Note that taking the transpose of a rank 1 array does nothing:
v = np.array([1,2,3])
print(v)    # Prints "[1 2 3]"
print(v.T)  # Prints "[1 2 3]"

<img src="images/scipy.jpg">

- ** Python-based ecosystem of open-source software for mathematics, science, and engineering. **

- ** The SciPy library provides many user-friendly and efficient numerical routines such as routines  for numerical integration and optimization. **

- ** One of its main aim is to provide a reimplementation of the  MATLAB toolboxes. **

- ** SciPy builds on this, and provides a large number of functions that operate on numpy arrays and are useful for different types of scientific and engineering applications. **


In [None]:
from scipy.misc import imread, imsave, imresize

# Read an JPEG image into a numpy array
img = imread('assets/cat.jpg')
print(img.dtype, img.shape)  # Prints "uint8 (400, 248, 3)"

# We can tint the image by scaling each of the color channels
# by a different scalar constant. The image has shape (400, 248, 3);
# we multiply it by the array [1, 0.95, 0.9] of shape (3,);
# numpy broadcasting means that this leaves the red channel unchanged,
# and multiplies the green and blue channels by 0.95 and 0.9
# respectively.
img_tinted = img * [1, 0.95, 0.9]

# Resize the tinted image to be 300 by 300 pixels.
img_tinted = imresize(img_tinted, (300, 300))

# Write the tinted image back to disk
imsave('assets/cat_tinted.jpg', img_tinted)

# <img src="assets/cat.jpg"> <img src="assets/cat_tinted.jpg">

In [None]:
import numpy as np
from scipy.spatial.distance import pdist, squareform

# Create the following array where each row is a point in 2D space:
# [[0 1]
#  [1 0]
#  [2 0]]
x = np.array([[0, 1], [1, 0], [2, 0]])
print(x)

# Compute the Euclidean distance between all rows of x.
# d[i, j] is the Euclidean distance between x[i, :] and x[j, :],
# and d is the following array:
# [[ 0.          1.41421356  2.23606798]
#  [ 1.41421356  0.          1.        ]
#  [ 2.23606798  1.          0.        ]]
d = squareform(pdist(x, 'euclidean'))
print(d)

<img src="images/matplotlib.jpg">

- ** Matplotlib is a Python 2D plotting library **

- ** Produces quality figures in a variety of hardcopy formats and interactive environments across platforms **

- ** Matplotlib tries to make easy things easy and hard things possible **

- ** Can generate plots, histograms, power spectra, bar charts, error charts, scatterplots, etc. **


In [None]:
import numpy as np
import matplotlib.pyplot as plt

# Compute the x and y coordinates for points on a sine curve
x = np.arange(0, 3 * np.pi, 0.1)
print(x)
y = np.sin(x)

# Plot the points using matplotlib
plt.plot(x, y)
plt.show()  # You must call plt.show() to make graphics appear.

In [None]:
import numpy as np
import matplotlib.pyplot as plt

# Compute the x and y coordinates for points on sine and cosine curves
x = np.arange(0, 3 * np.pi, 0.1)
y_sin = np.sin(x)
y_cos = np.cos(x)

# Plot the points using matplotlib
plt.plot(x, y_sin)
plt.plot(x, y_cos)
plt.xlabel('x axis label')
plt.ylabel('y axis label')
plt.title('Sine and Cosine')
plt.legend(['Sine', 'Cosine'])
plt.show()

In [None]:
import numpy as np
import matplotlib.pyplot as plt

# Compute the x and y coordinates for points on sine and cosine curves
x = np.arange(0, 3 * np.pi, 0.1)
y_sin = np.sin(x)
y_cos = np.cos(x)

# Set up a subplot grid that has height 2 and width 1,
# and set the first such subplot as active.
plt.subplot(2, 1, 1)

# Make the first plot
plt.plot(x, y_sin)
plt.title('Sine')

# Set the second subplot as active, and make the second plot.
plt.subplot(2, 1, 2)
plt.plot(x, y_cos)
plt.title('Cosine')

# Show the figure.
plt.show()

In [None]:
import numpy as np
from scipy.misc import imread, imresize
import matplotlib.pyplot as plt

img = imread('assets/cat.jpg')
img_tinted = img * [1, 0.95, 0.9]

# Show the original image
plt.subplot(1, 2, 1)
plt.imshow(img)

# Show the tinted image
plt.subplot(1, 2, 2)

# A slight gotcha with imshow is that it might give strange results
# if presented with data that is not uint8. To work around this, we
# explicitly cast the image to uint8 before displaying it.
plt.imshow(np.uint8(img_tinted))
plt.show()

In [None]:
import pandas as pd
excel_file = 'data/movies.xls'
movies = pd.read_excel(excel_file)
sorted_by_gross = movies.sort_values(['Gross Earnings'], ascending=False)

import matplotlib.pyplot as plt
%matplotlib inline

sorted_by_gross['Gross Earnings'].head(10).plot(kind="barh")
plt.show()

In [None]:
movies['IMDB Score'].plot(kind="hist")
plt.show()

In [None]:
movies["Net Earnings"] = movies["Gross Earnings"] - movies["Budget"]
sorted_movies = movies[['Net Earnings']].sort_values(['Net Earnings'], ascending=[False])
sorted_movies.head(10)['Net Earnings'].plot.barh()
plt.show()

In [None]:
earnings_by_year = movies_subset.pivot_table(index=['Year'])
earnings_by_year.plot()
plt.show()

In [None]:
movies_subset = movies[['Country', 'Language', 'Gross Earnings']]
earnings_by_co_lang = movies_subset.pivot_table(index=['Country', 'Language'])
earnings_by_co_lang.head(20).plot(kind='bar', figsize=(20,8))
plt.show()

<img src="images/scikit.png">

- ** Scikit – Learn python based statistical Machine Learning library **

- ** Simple and efficient tools for data mining and data analysis **

- ** Accessible to everybody, and reusable in various contexts **

- ** Built on NumPy, SciPy, and matplotlib **

- ** Open source, commercially usable - BSD license **

- ** Performs Preprocessing, Dimensionality reduction, Classification, Clustering and Regression **


### <font color=blue>Data-Set</font>

In [None]:
import numpy as np
X = np.random.random((10,5))
#shape of X
print(X.shape)
print(X)

In [None]:
y = np.array(['M','F','M','M','M','F','M','F','M','F'])
#shape of y
print(y.shape)
y

### <font color=blue>Split Data-Set into Train Set And Test Set</font>

In [None]:
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

### <font color=blue>Initialize Models</font>

In [None]:
#support vector machines(svm)
from sklearn.svm import SVC
svc = SVC(kernel='linear')

#KNN
from sklearn import neighbors
knn = neighbors.KNeighborsClassifier(n_neighbors=4)

#Naive Bayes 
from sklearn.naive_bayes import GaussianNB
gnb = GaussianNB()

# Principal Component Analysis(PCA) 
from sklearn.decomposition import PCA
pca = PCA(n_components=0.95)

# K-Means 
from sklearn.cluster import KMeans
k_means = KMeans(n_clusters=4, random_state=0)

### <font color=blue>Fit or Train Data Into Models</font>

In [None]:
svc.fit(X_train, y_train)

In [None]:
knn.fit(X_train, y_train)

In [None]:
gnb.fit(X_train, y_train)

In [None]:
k_means.fit(X_train)

In [None]:
pca_model = pca.fit_transform(X_train)

### <font color=blue>Prediction or Testing</font>

In [None]:
pred = svc.predict(X_test)
print(pred)

In [None]:
pred = knn.predict(X_test)
print(pred)

In [None]:
pred = gnb.predict(X_test)
print(pred)

In [None]:
pred = k_means.predict(X_test)
pred