<font color='darkred'> Unless otherwise noted, **this notebook will not be reviewed or autograded.**</font> You are welcome to use it for scratchwork, but **only the files listed in the exercises will be checked.**

---

# Exercises

For these exercises, add your functions to the *apputil\.py* file. If you like, you're welcome to adjust the *app\.py* file, but it is not required.

## Notes on Recursion

A [recursive function](https://www.w3schools.com/python/gloss_python_function_recursion.asp) is one which calls itself.

1. When the function is called, your CPU runs through each line of code until the function needs to be called again.
2. At that point, all variables are saved in memory, and the function runs through each line of code again until the function is called (again, but with a different passed argument), and so on.
3. Eventually, this process will stop at the "bottom of the **stack**", where the function doesn't get a chance to call itself again (likely because of some condition un/met by the latest passed argument).
4. Then, your CPU will work its way back up the stack to the final result. For example, take a look at [this visual example](https://realpython.com/python-recursion/#calculate-factorial) of calculating 4!.

When you write these functions, keep two things in mind:

- You will need a built-in stopping point (i.e., the "bottom"), where your function returns some result before it calls itself.
- **Don't think too hard about this.** Recursion can be perplexing to conceptualize when writing the code. So, when you call the function inside the function, think about it as a magical "hidden" function that has already done what you want it to do.
- [Python Tutor](https://pythontutor.com/) ([editor](https://pythontutor.com/visualize.html#mode=edit)) can be a helpful resource for this exercise!

## Exercise 1

The Fibonacci Series starts with 0 and 1. Each of the following numbers are the sum of the previous two numbers in the series:

`0 1 1 2 3 5 8 13 21 34 ...`

So, `fib(9) = 34`.

Write a recursive function (`fib`) that, given `n`, will return the `n`th number of the Fibonacci Series.

*Test your function using Google or any other tool that can calculate the Fibonacci Series.*

In [None]:
"""
Write a recursive function to compute the nth Fibonacci number (sum of previous two numbers within the series starting with 0 and 1).
The fib function is defined as follows:
fib(0) = 0
fib(1) = 1
fib(n) = fib(n-1) + fib(n-2) for n > 1
The function should raise a ValueError if the input is not a non-negative integer, which will account for any invalid inputs.
"""

def fib(n):
    if type(n) is int and n >= 0:
        if n == 0:
            return 0
        elif n == 1:
            return 1
        else:
            return fib(n-1) + fib(n-2)
    else:
        raise ValueError("Input must be a non-negative integer")

In [6]:
fib(-1)

ValueError: Input must be a non-negative integer


## Exercise 2

Write a (single) recursive function, `to_binary()`, that [converts](https://en.wikipedia.org/wiki/Binary_number#Conversion_to_and_from_other_numeral_systems) an integer into its [binary](https://en.wikipedia.org/wiki/Binary_number) representation. So, for example:

```python
to_binary(2)   -->  10
to_binary(12)  -->  1100
```

*Note: you can test your function with the built in `bin()` function.*

In [12]:
"""
Write a function to convert a non-negative integer to its binary representation as a string.
If the number is 0, the function will automatically return 0.
While the current number is greater than 0, add the remainer of the number when divided by 2 to the binary string and find the new number (which is the current number divided by 2 using the floor division operator (aka the divided number down)).
Once the number is 0, return the found converted binary string.
The function should raise a ValueError if the input is not a non-negative integer, which will account for any invalid inputs.
"""
def to_binary(num):
    binary_string = ""
    if type(num) is int and num >= 0:
        if num == 0:
            return "0"
        while num > 0:
            binary_string = str(num % 2) + binary_string
            num = num // 2
        return binary_string
    else:
        raise ValueError("Input must be a non-negative integer")

In [10]:
to_binary(2)

'10'

## Exercise 3 

Use the raw Bellevue Almshouse Dataset (`df_bellevue`) extracted at the top of the lab (i.e., with `pd.read_csv ...`).

**Write a function for each of the following tasks. Name these functions `task_i()`** (i.e., without any input arguments).

1. Return a list of all column names, *sorted* such that the first column has the *least* missing values, and the last column has the *most* missing values (use the raw column names).
   - *Note: there is an issue with the `gender` column you'll need to remedy first ...*
2. Return a **data frame** with two columns:
   - the year (for each year in the data), `year`
   - the total number of entries (immigrant admissions) for each year, `total_admissions`
3. Return a **series** with:
   - Index: gender (for each gender in the data)
   - Values: the average age for the indexed gender.
4. Return a list of the 5 most common professions *in order of prevalence* (so, the most common is first).

For each of these, if there are messy data issues, use the `print` statement to explain.


In [None]:
#use print statements to explain messy data issues
import pandas as pd
import seaborn as sns
import numpy as np
url = 'https://github.com/melaniewalsh/Intro-Cultural-Analytics/raw/master/book/data/bellevue_almshouse_modified.csv'

df_bellevue = pd.read_csv(url)

#return list of all col names, sorted, so that first col has LEAST missing vals and last col has MOST (use raw col names)
#remedy gender col first (making it so that there are only M, F, and NaN gender values)
df_bellevue['gender'] = df_bellevue['gender'].replace(['h', 'g', '?'], np.nan)

def task_1():
    col_list = df_bellevue.columns.tolist()
    return df_bellevue[col_list].isna().sum().sort_values()
    
#return df with 2 cols, year for each year and total num of entries (immigrant admins) for each year    
def task_2():
    return 0
#return a series with index: gender (for each gender) and values: avg age of indexed gender
def task_3():
    return 0
#return a list of the top 5 professions in order of prevalence (most common first)
def task_4():
    return 0


date_in          0
last_name        0
first_name       4
gender           5
age             50
profession    1019
disease       3087
children      9547
dtype: int64

In [None]:
task_1()