# Housing Rental Analysis for San Francisco

In this challenge, your job is to use your data visualization skills, including aggregation, interactive visualizations, and geospatial analysis, to find properties in the San Francisco market that are viable investment opportunities.

## Instructions

Use the `san_francisco_housing.ipynb` notebook to visualize and analyze the real-estate data.

Note that this assignment requires you to create a visualization by using hvPlot and GeoViews. Additionally, you need to read the `sfo_neighborhoods_census_data.csv` file from the `Resources` folder into the notebook and create the DataFrame that you’ll use in the analysis.

The main task in this Challenge is to visualize and analyze the real-estate data in your Jupyter notebook. Use the `san_francisco_housing.ipynb` notebook to complete the following tasks:

* Calculate and plot the housing units per year.

* Calculate and plot the average prices per square foot.

* Compare the average prices by neighborhood.

* Build an interactive neighborhood map.

* Compose your data story.

### Calculate and Plot the Housing Units per Year

For this part of the assignment, use numerical and visual aggregation to calculate the number of housing units per year, and then visualize the results as a bar chart. To do so, complete the following steps:

1. Use the `groupby` function to group the data by year. Aggregate the results by the `mean` of the groups.

2. Use the `hvplot` function to plot the `housing_units_by_year` DataFrame as a bar chart. Make the x-axis represent the `year` and the y-axis represent the `housing_units`.

3. Style and format the line plot to ensure a professionally styled visualization.

4. Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of the resulting bar chart.](Images/zoomed-housing-units-by-year.png)

5. Answer the following question:

    * What’s the overall trend in housing units over the period that you’re analyzing?

### Calculate and Plot the Average Sale Prices per Square Foot

For this part of the assignment, use numerical and visual aggregation to calculate the average prices per square foot, and then visualize the results as a bar chart. To do so, complete the following steps:

1. Group the data by year, and then average the results. What’s the lowest gross rent that’s reported for the years that the DataFrame includes?

2. Create a new DataFrame named `prices_square_foot_by_year` by filtering out the “housing_units” column. The new DataFrame should include the averages per year for only the sale price per square foot and the gross rent.

3. Use hvPlot to plot the `prices_square_foot_by_year` DataFrame as a line plot.

    > **Hint** This single plot will include lines for both `sale_price_sqr_foot` and `gross_rent`.

4. Style and format the line plot to ensure a professionally styled visualization.

5. Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of the resulting plot.](Images/avg-sale-px-sq-foot-gross-rent.png)

6. Use both the `prices_square_foot_by_year` DataFrame and interactive plots to answer the following questions:

    * Did any year experience a drop in the average sale price per square foot compared to the previous year?

    * If so, did the gross rent increase or decrease during that year?

### Compare the Average Sale Prices by Neighborhood

For this part of the assignment, use interactive visualizations and widgets to explore the average sale price per square foot by neighborhood. To do so, complete the following steps:

1. Create a new DataFrame that groups the original DataFrame by year and neighborhood. Aggregate the results by the `mean` of the groups.

2. Filter out the “housing_units” column to create a DataFrame that includes only the `sale_price_sqr_foot` and `gross_rent` averages per year.

3. Create an interactive line plot with hvPlot that visualizes both `sale_price_sqr_foot` and `gross_rent`. Set the x-axis parameter to the year (`x="year"`). Use the `groupby` parameter to create an interactive widget for `neighborhood`.

4. Style and format the line plot to ensure a professionally styled visualization.

5. Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of the resulting plot.](Images/pricing-info-by-neighborhood.png)

6. Use the interactive visualization to answer the following question:

    * For the Anza Vista neighborhood, is the average sale price per square foot for 2016 more or less than the price that’s listed for 2012? 

### Build an Interactive Neighborhood Map

For this part of the assignment, explore the geospatial relationships in the data by using interactive visualizations with hvPlot and GeoViews. To build your map, use the `sfo_data_df` DataFrame (created during the initial import), which includes the neighborhood location data with the average prices. To do all this, complete the following steps:

1. Read the `neighborhood_coordinates.csv` file from the `Resources` folder into the notebook, and create a DataFrame named `neighborhood_locations_df`. Be sure to set the `index_col` of the DataFrame as “Neighborhood”.

2. Using the original `sfo_data_df` Dataframe, create a DataFrame named `all_neighborhood_info_df` that groups the data by neighborhood. Aggregate the results by the `mean` of the group.

3. Review the two code cells that concatenate the `neighborhood_locations_df` DataFrame with the `all_neighborhood_info_df` DataFrame. Note that the first cell uses the [Pandas concat function](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.concat.html) to create a DataFrame named `all_neighborhoods_df`. The second cell cleans the data and sets the “Neighborhood” column. Be sure to run these cells to create the `all_neighborhoods_df` DataFrame, which you’ll need to create the geospatial visualization.

4. Using hvPlot with GeoViews enabled, create a `points` plot for the `all_neighborhoods_df` DataFrame. Be sure to do the following:

    * Set the `geo` parameter to True.
    * Set the `size` parameter to “sale_price_sqr_foot”.
    * Set the `color` parameter to “gross_rent”.
    * Set the `frame_width` parameter to 700.
    * Set the `frame_height` parameter to 500.
    * Include a descriptive title.

Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of a scatter plot created with hvPlot and GeoViews.](Images/6-4-geoviews-plot.png)

5. Use the interactive map to answer the following question:

    * Which neighborhood has the highest gross rent, and which has the highest sale price per square foot?

### Compose Your Data Story

Based on the visualizations that you created, answer the following questions:

* How does the trend in rental income growth compare to the trend in sales prices? Does this same trend hold true for all the neighborhoods across San Francisco?

* What insights can you share with your company about the potential one-click, buy-and-rent strategy that they're pursuing? Do neighborhoods exist that you would suggest for investment, and why?

In [17]:
# Import the required libraries and dependencies
import pandas as pd
import hvplot.pandas
from pathlib import Path
from bokeh.models.widgets import Dropdown
import warnings

## Import the data 

In [2]:
# Using the read_csv function and Path module, create a DataFrame 
# by importing the sfo_neighborhoods_census_data.csv file from the Resources folder
sfo_data_df = pd.read_csv(Path("Resources/sfo_neighborhoods_census_data.csv"))

# Review the first and last five rows of the DataFrame
sfo_data_df

Unnamed: 0,year,neighborhood,sale_price_sqr_foot,housing_units,gross_rent
0,2010,Alamo Square,291.182945,372560,1239
1,2010,Anza Vista,267.932583,372560,1239
2,2010,Bayview,170.098665,372560,1239
3,2010,Buena Vista Park,347.394919,372560,1239
4,2010,Central Richmond,319.027623,372560,1239
...,...,...,...,...,...
392,2016,Telegraph Hill,903.049771,384242,4390
393,2016,Twin Peaks,970.085470,384242,4390
394,2016,Van Ness/ Civic Center,552.602567,384242,4390
395,2016,Visitacion Valley,328.319007,384242,4390


---

## Calculate and Plot the Housing Units per Year

For this part of the assignment, use numerical and visual aggregation to calculate the number of housing units per year, and then visualize the results as a bar chart. To do so, complete the following steps:

1. Use the `groupby` function to group the data by year. Aggregate the results by the `mean` of the groups.

2. Use the `hvplot` function to plot the `housing_units_by_year` DataFrame as a bar chart. Make the x-axis represent the `year` and the y-axis represent the `housing_units`.

3. Style and format the line plot to ensure a professionally styled visualization.

4. Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of the resulting bar chart.](Images/zoomed-housing-units-by-year.png)

5. Answer the following question:

    * What’s the overall trend in housing units over the period that you’re analyzing?



### Step 1: Use the `groupby` function to group the data by year. Aggregate the results by the `mean` of the groups.

In [3]:
# Create a numerical aggregation that groups the data by the year and then averages the results.
housing_units_by_year = sfo_data_df.groupby('year').mean().loc[:, 'housing_units']

# Review the DataFrame
housing_units_by_year= housing_units_by_year.apply(lambda x: "{:.0f}K".format(x/1000))
housing_units_by_year

year
2010    373K
2011    375K
2012    376K
2013    378K
2014    380K
2015    382K
2016    384K
Name: housing_units, dtype: object

### Step 2: Use the `hvplot` function to plot the `housing_units_by_year` DataFrame as a bar chart. Make the x-axis represent the `year` and the y-axis represent the `housing_units`.

### Step 3: Style and format the line plot to ensure a professionally styled visualization.

In [4]:
# Create a visual aggregation explore the housing units by year
housing_units_by_year.hvplot(
    xlabel="Year",
    ylabel="Housing Units (in thousands)",
    title="Housing Units by Year: 2010 - 2016",
    rot=45
)

### Step 5: Answer the following question:

**Question:** What is the overall trend in housing_units over the period being analyzed?

**Answer:** San Franscisco is experience stable (albeit slow) growth over the years 2010 to 2016.

---

## Calculate and Plot the Average Sale Prices per Square Foot

For this part of the assignment, use numerical and visual aggregation to calculate the average prices per square foot, and then visualize the results as a bar chart. To do so, complete the following steps:

1. Group the data by year, and then average the results. What’s the lowest gross rent that’s reported for the years that the DataFrame includes?

2. Create a new DataFrame named `prices_square_foot_by_year` by filtering out the “housing_units” column. The new DataFrame should include the averages per year for only the sale price per square foot and the gross rent.

3. Use hvPlot to plot the `prices_square_foot_by_year` DataFrame as a line plot.

    > **Hint** This single plot will include lines for both `sale_price_sqr_foot` and `gross_rent`.

4. Style and format the line plot to ensure a professionally styled visualization.

5. Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of the resulting plot.](Images/avg-sale-px-sq-foot-gross-rent.png)

6. Use both the `prices_square_foot_by_year` DataFrame and interactive plots to answer the following questions:

    * Did any year experience a drop in the average sale price per square foot compared to the previous year?

    * If so, did the gross rent increase or decrease during that year?



### Step 1: Group the data by year, and then average the results.

In [5]:
# Create a numerical aggregation by grouping the data by year and averaging the results
sfo_avg_by_year = sfo_data_df.groupby('year').mean()

# Review the resulting DataFrame
sfo_avg_by_year.sort_values('gross_rent')

Unnamed: 0_level_0,sale_price_sqr_foot,housing_units,gross_rent
year,Unnamed: 1_level_1,Unnamed: 2_level_1,Unnamed: 3_level_1
2010,369.344353,372560.0,1239.0
2011,341.903429,374507.0,1530.0
2012,399.389968,376454.0,2324.0
2013,483.600304,378401.0,2971.0
2014,556.277273,380348.0,3528.0
2015,632.540352,382295.0,3739.0
2016,697.643709,384242.0,4390.0


**Question:** What is the lowest gross rent reported for the years included in the DataFrame?

**Answer:** The lowest average gross rent reported in the dataframe is $1,239 which was reported in 2010.

### Step 2: Create a new DataFrame named `prices_square_foot_by_year` by filtering out the “housing_units” column. The new DataFrame should include the averages per year for only the sale price per square foot and the gross rent.

In [6]:
# Filter out the housing_units column, creating a new DataFrame 
# Keep only sale_price_sqr_foot and gross_rent averages per year
prices_square_foot_by_year = sfo_avg_by_year.drop('housing_units', axis=1)

# Review the resulting DataFrame
#prices_square_foot_by_year= prices_square_foot_by_year.round(2).apply(lambda x: '${:,.2f}'.format(x))
#prices_square_foot_by_year['sale_price_sqr_foot']= prices_square_foot_by_year['sale_price_sqr_foot'].round(2).apply(lambda x: '${:,.2f}'.format(x))
#prices_square_foot_by_year['gross_rent']= prices_square_foot_by_year['gross_rent'].round(2).apply(lambda x: '${:,.2f}'.format(x))

prices_square_foot_by_year

Unnamed: 0_level_0,sale_price_sqr_foot,gross_rent
year,Unnamed: 1_level_1,Unnamed: 2_level_1
2010,369.344353,1239.0
2011,341.903429,1530.0
2012,399.389968,2324.0
2013,483.600304,2971.0
2014,556.277273,3528.0
2015,632.540352,3739.0
2016,697.643709,4390.0


### Step 3: Use hvPlot to plot the `prices_square_foot_by_year` DataFrame as a line plot.

> **Hint** This single plot will include lines for both `sale_price_sqr_foot` and `gross_rent`

### Step 4: Style and format the line plot to ensure a professionally styled visualization.


In [7]:
# Plot prices_square_foot_by_year. 
# Inclued labels for the x- and y-axes, and a title.
prices_square_foot_by_year.hvplot(
    xlabel="Year",
    ylabel="Price",
    title="Housing Costs: 2010 - 2016",
    rot=45
).opts(
    yformatter='%.2f$'
)  

### Step 6: Use both the `prices_square_foot_by_year` DataFrame and interactive plots to answer the following questions:

**Question:** Did any year experience a drop in the average sale price per square foot compared to the previous year?

**Answer:** Yes - in 2011, SFO experienced a drop in price per square foot from $369.34 to $341.90.

**Question:** If so, did the gross rent increase or decrease during that year?

**Answer:** Interestingly, during the year that prices per square dropped, there was a continued increase in gross rent. There could be several reasons for this including growth and movement to parts of the city that have a lower price point and/or movement to and building of larger apartments and homes.

---

## Compare the Average Sale Prices by Neighborhood

For this part of the assignment, use interactive visualizations and widgets to explore the average sale price per square foot by neighborhood. To do so, complete the following steps:

1. Create a new DataFrame that groups the original DataFrame by year and neighborhood. Aggregate the results by the `mean` of the groups.

2. Filter out the “housing_units” column to create a DataFrame that includes only the `sale_price_sqr_foot` and `gross_rent` averages per year.

3. Create an interactive line plot with hvPlot that visualizes both `sale_price_sqr_foot` and `gross_rent`. Set the x-axis parameter to the year (`x="year"`). Use the `groupby` parameter to create an interactive widget for `neighborhood`.

4. Style and format the line plot to ensure a professionally styled visualization.

5. Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of the resulting plot.](Images/pricing-info-by-neighborhood.png)

6. Use the interactive visualization to answer the following question:

    * For the Anza Vista neighborhood, is the average sale price per square foot for 2016 more or less than the price that’s listed for 2012? 


### Step 1: Create a new DataFrame that groups the original DataFrame by year and neighborhood. Aggregate the results by the `mean` of the groups.

In [8]:
# Group by year and neighborhood and then create a new dataframe of the mean values
prices_by_year_by_neighborhood = sfo_data_df.groupby(['year', 'neighborhood']).mean()

# Review the DataFrame
prices_by_year_by_neighborhood

Unnamed: 0_level_0,Unnamed: 1_level_0,sale_price_sqr_foot,housing_units,gross_rent
year,neighborhood,Unnamed: 2_level_1,Unnamed: 3_level_1,Unnamed: 4_level_1
2010,Alamo Square,291.182945,372560.0,1239.0
2010,Anza Vista,267.932583,372560.0,1239.0
2010,Bayview,170.098665,372560.0,1239.0
2010,Buena Vista Park,347.394919,372560.0,1239.0
2010,Central Richmond,319.027623,372560.0,1239.0
...,...,...,...,...
2016,Telegraph Hill,903.049771,384242.0,4390.0
2016,Twin Peaks,970.085470,384242.0,4390.0
2016,Van Ness/ Civic Center,552.602567,384242.0,4390.0
2016,Visitacion Valley,328.319007,384242.0,4390.0


### Step 2: Filter out the “housing_units” column to create a DataFrame that includes only the `sale_price_sqr_foot` and `gross_rent` averages per year.

In [9]:
# Filter out the housing_units
prices_by_year_by_neighborhood = prices_by_year_by_neighborhood.drop('housing_units', axis=1)

# Review the first and last five rows of the DataFrame
display(prices_by_year_by_neighborhood.head(5))
print("-----------------------------")
display(prices_by_year_by_neighborhood.tail(5))


Unnamed: 0_level_0,Unnamed: 1_level_0,sale_price_sqr_foot,gross_rent
year,neighborhood,Unnamed: 2_level_1,Unnamed: 3_level_1
2010,Alamo Square,291.182945,1239.0
2010,Anza Vista,267.932583,1239.0
2010,Bayview,170.098665,1239.0
2010,Buena Vista Park,347.394919,1239.0
2010,Central Richmond,319.027623,1239.0


-----------------------------


Unnamed: 0_level_0,Unnamed: 1_level_0,sale_price_sqr_foot,gross_rent
year,neighborhood,Unnamed: 2_level_1,Unnamed: 3_level_1
2016,Telegraph Hill,903.049771,4390.0
2016,Twin Peaks,970.08547,4390.0
2016,Van Ness/ Civic Center,552.602567,4390.0
2016,Visitacion Valley,328.319007,4390.0
2016,Westwood Park,631.195426,4390.0


### Step 3: Create an interactive line plot with hvPlot that visualizes both `sale_price_sqr_foot` and `gross_rent`. Set the x-axis parameter to the year (`x="year"`). Use the `groupby` parameter to create an interactive widget for `neighborhood`.

### Step 4: Style and format the line plot to ensure a professionally styled visualization.

In [19]:
# Use hvplot to create an interactive line plot of the average price per square foot
# The plot should have a dropdown selector for the neighborhood

neighborhood_price_plot = prices_by_year_by_neighborhood.hvplot(
    x='year',
    xlabel="Year",
    ylabel="Cost (in dollars)",
    groupby="neighborhood",
    widget="dropdown",
    title="Housing costs by neighborhood: 2010 - 2016"
).opts(
    yformatter='%2f$'
)

neighborhood_price_plot



### Step 6: Use the interactive visualization to answer the following question:

**Question:** For the Anza Vista neighborhood, is the average sale price per square foot for 2016 more or less than the price that’s listed for 2012? 

**Answer:** The average sale price per square foot for the Anza Vista neighborhood in 2016 is less than the price listed in 2012. There was a huge decline from $344.5 to $88. It would be worthwhile to look into causes of this rapid decline.

---

## Build an Interactive Neighborhood Map

For this part of the assignment, explore the geospatial relationships in the data by using interactive visualizations with hvPlot and GeoViews. To build your map, use the `sfo_data_df` DataFrame (created during the initial import), which includes the neighborhood location data with the average prices. To do all this, complete the following steps:

1. Read the `neighborhood_coordinates.csv` file from the `Resources` folder into the notebook, and create a DataFrame named `neighborhood_locations_df`. Be sure to set the `index_col` of the DataFrame as “Neighborhood”.

2. Using the original `sfo_data_df` Dataframe, create a DataFrame named `all_neighborhood_info_df` that groups the data by neighborhood. Aggregate the results by the `mean` of the group.

3. Review the two code cells that concatenate the `neighborhood_locations_df` DataFrame with the `all_neighborhood_info_df` DataFrame. Note that the first cell uses the [Pandas concat function](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.concat.html) to create a DataFrame named `all_neighborhoods_df`. The second cell cleans the data and sets the “Neighborhood” column. Be sure to run these cells to create the `all_neighborhoods_df` DataFrame, which you’ll need to create the geospatial visualization.

4. Using hvPlot with GeoViews enabled, create a `points` plot for the `all_neighborhoods_df` DataFrame. Be sure to do the following:

    * Set the `size` parameter to “sale_price_sqr_foot”.

    * Set the `color` parameter to “gross_rent”.

    * Set the `size_max` parameter to “25”.

    * Set the `zoom` parameter to “11”.

Note that your resulting plot should appear similar to the following image:

![A screenshot depicts an example of a scatter plot created with hvPlot and GeoViews.](Images/6-4-geoviews-plot.png)

5. Use the interactive map to answer the following question:

    * Which neighborhood has the highest gross rent, and which has the highest sale price per square foot?


### Step 1: Read the `neighborhood_coordinates.csv` file from the `Resources` folder into the notebook, and create a DataFrame named `neighborhood_locations_df`. Be sure to set the `index_col` of the DataFrame as “Neighborhood”.

In [11]:
# Load neighborhoods coordinates data
neighborhood_locations_df = pd.read_csv(Path("Resources/neighborhoods_coordinates.csv"),index_col="Neighborhood")

# Review the DataFrame
neighborhood_locations_df

Unnamed: 0_level_0,Lat,Lon
Neighborhood,Unnamed: 1_level_1,Unnamed: 2_level_1
Alamo Square,37.791012,-122.402100
Anza Vista,37.779598,-122.443451
Bayview,37.734670,-122.401060
Bayview Heights,37.728740,-122.410980
Bernal Heights,37.728630,-122.443050
...,...,...
West Portal,37.740260,-122.463880
Western Addition,37.792980,-122.435790
Westwood Highlands,37.734700,-122.456854
Westwood Park,37.734150,-122.457000


### Step 2: Using the original `sfo_data_df` Dataframe, create a DataFrame named `all_neighborhood_info_df` that groups the data by neighborhood. Aggregate the results by the `mean` of the group.

In [59]:
# Calculate the mean values for each neighborhood
all_neighborhood_info_df = sfo_data_df.groupby("neighborhood").mean()

# Review the resulting DataFrame
all_neighborhood_info_df

Unnamed: 0_level_0,year,sale_price_sqr_foot,housing_units,gross_rent
neighborhood,Unnamed: 1_level_1,Unnamed: 2_level_1,Unnamed: 3_level_1,Unnamed: 4_level_1
Alamo Square,2013.000000,366.020712,378401.00,2817.285714
Anza Vista,2013.333333,373.382198,379050.00,3031.833333
Bayview,2012.000000,204.588623,376454.00,2318.400000
Bayview Heights,2015.000000,590.792839,382295.00,3739.000000
Bernal Heights,2013.500000,576.746488,379374.50,3080.333333
...,...,...,...,...
West Portal,2012.250000,498.488485,376940.75,2515.500000
Western Addition,2012.500000,307.562201,377427.50,2555.166667
Westwood Highlands,2012.000000,533.703935,376454.00,2250.500000
Westwood Park,2015.000000,687.087575,382295.00,3959.000000


### Step 3: Review the two code cells that concatenate the `neighborhood_locations_df` DataFrame with the `all_neighborhood_info_df` DataFrame. 

Note that the first cell uses the [Pandas concat function](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.concat.html) to create a DataFrame named `all_neighborhoods_df`. 

The second cell cleans the data and sets the “Neighborhood” column. 

Be sure to run these cells to create the `all_neighborhoods_df` DataFrame, which you’ll need to create the geospatial visualization.

In [13]:
# Using the Pandas `concat` function, join the 
# neighborhood_locations_df and the all_neighborhood_info_df DataFrame
# The axis of the concatenation is "columns".
# The concat function will automatially combine columns with
# identical information, while keeping the additional columns.
all_neighborhoods_df = pd.concat(
    [neighborhood_locations_df, all_neighborhood_info_df], 
    axis="columns",
    sort=False
)

# Review the resulting DataFrame
display(all_neighborhoods_df.head())
display(all_neighborhoods_df.tail())


Unnamed: 0,Lat,Lon,year,sale_price_sqr_foot,housing_units,gross_rent
Alamo Square,37.791012,-122.4021,2013.0,366.020712,378401.0,2817.285714
Anza Vista,37.779598,-122.443451,2013.333333,373.382198,379050.0,3031.833333
Bayview,37.73467,-122.40106,2012.0,204.588623,376454.0,2318.4
Bayview Heights,37.72874,-122.41098,2015.0,590.792839,382295.0,3739.0
Bernal Heights,37.72863,-122.44305,,,,


Unnamed: 0,Lat,Lon,year,sale_price_sqr_foot,housing_units,gross_rent
Yerba Buena,37.79298,-122.39636,2012.5,576.709848,377427.5,2555.166667
Bernal Heights,,,2013.5,576.746488,379374.5,3080.333333
Downtown,,,2013.0,391.434378,378401.0,2817.285714
Ingleside,,,2012.5,367.895144,377427.5,2509.0
Outer Richmond,,,2013.0,473.900773,378401.0,2817.285714


In [44]:
# Call the dropna function to remove any neighborhoods that do not have data
all_neighborhoods_df = all_neighborhoods_df.reset_index().dropna()

# Rename the "index" column as "Neighborhood" for use in the Visualization
all_neighborhoods_df = all_neighborhoods_df.rename(columns={"index": "Neighborhood"})

# Review the resulting DataFrame
display(all_neighborhoods_df.sort_values('gross_rent').head())
display(all_neighborhoods_df.sort_values('gross_rent').tail())
print("---------------------------------------------------------------------")
display(all_neighborhoods_df.sort_values('sale_price_sqr_foot').head())
display(all_neighborhoods_df.sort_values('sale_price_sqr_foot').tail())

Unnamed: 0,Neighborhood,Neighborhood.1,Neighborhood.2,Neighborhood.3,Neighborhood.4,Neighborhood.5,Lat,Lon,year,sale_price_sqr_foot,housing_units,gross_rent
17,17,17,17,17,19,Forest Knolls,37.75706,-122.45509,2011.0,321.840837,374507.0,1781.5
11,11,11,11,11,12,Diamond Heights,37.72863,-122.44305,2011.5,434.991739,375480.5,2016.0
56,56,56,56,56,60,South Beach,37.78323,-122.40065,2011.666667,650.124479,375805.0,2099.0
35,35,35,35,35,38,Miraloma Park,37.73415,-122.457,2011.75,779.810842,375967.25,2155.25
66,66,66,66,66,70,Westwood Highlands,37.7347,-122.456854,2012.0,533.703935,376454.0,2250.5


Unnamed: 0,Neighborhood,Neighborhood.1,Neighborhood.2,Neighborhood.3,Neighborhood.4,Neighborhood.5,Lat,Lon,year,sale_price_sqr_foot,housing_units,gross_rent
33,33,33,33,33,36,Merced Heights,37.71993,-122.46595,2014.0,788.844818,380348.0,3414.0
55,55,55,55,55,59,Silver Terrace,37.73467,-122.40106,2014.0,170.292549,380348.0,3528.0
63,63,63,63,63,67,Visitacion Valley,37.72874,-122.41098,2014.5,301.46618,381321.5,3657.0
3,3,3,3,3,3,Bayview Heights,37.72874,-122.41098,2015.0,590.792839,382295.0,3739.0
67,67,67,67,67,71,Westwood Park,37.73415,-122.457,2015.0,687.087575,382295.0,3959.0


---------------------------------------------------------------------


Unnamed: 0,Neighborhood,Neighborhood.1,Neighborhood.2,Neighborhood.3,Neighborhood.4,Neighborhood.5,Lat,Lon,year,sale_price_sqr_foot,housing_units,gross_rent
55,55,55,55,55,59,Silver Terrace,37.73467,-122.40106,2014.0,170.292549,380348.0,3528.0
22,22,22,22,22,24,Hunters Point,37.72551,-122.37178,2012.5,170.62492,377427.5,2489.0
2,2,2,2,2,2,Bayview,37.73467,-122.40106,2012.0,204.588623,376454.0,2318.4
44,44,44,44,44,47,Outer Mission,37.7228,-122.43869,2013.5,242.370952,379374.5,2995.75
63,63,63,63,63,67,Visitacion Valley,37.72874,-122.41098,2014.5,301.46618,381321.5,3657.0


Unnamed: 0,Neighborhood,Neighborhood.1,Neighborhood.2,Neighborhood.3,Neighborhood.4,Neighborhood.5,Lat,Lon,year,sale_price_sqr_foot,housing_units,gross_rent
67,67,67,67,67,71,Westwood Park,37.73415,-122.457,2015.0,687.087575,382295.0,3959.0
47,47,47,47,47,51,Pacific Heights,37.79298,-122.43579,2013.0,689.555817,378401.0,2817.285714
35,35,35,35,35,38,Miraloma Park,37.73415,-122.457,2011.75,779.810842,375967.25,2155.25
33,33,33,33,33,36,Merced Heights,37.71993,-122.46595,2014.0,788.844818,380348.0,3414.0
61,61,61,61,61,65,Union Square District,37.79101,-122.4021,2012.5,903.993258,377427.5,2555.166667


### Step 4: Using hvPlot with GeoViews enabled, create a `points` plot for the `all_neighborhoods_df` DataFrame. Be sure to do the following:

* Set the `geo` parameter to True.
* Set the `size` parameter to “sale_price_sqr_foot”.
* Set the `color` parameter to “gross_rent”.
* Set the `frame_width` parameter to 700.
* Set the `frame_height` parameter to 500.
* Include a descriptive title.

In [38]:
# Create a plot to analyze neighborhood info
# Set size = 300 to make the points easier to see on the map
SFO_geoplot = all_neighborhoods_df.hvplot.points(
    'Lon', 
    'Lat', 
    geo=True, 
    color="gross_rent",
    size = "sale_price_sqr_foot",
    tiles='OSM',
    frame_width = 700,
    frame_height = 500,
    title = "Gross Rent and Sale Price (by sqft) by SFO Neighborhood: 2010-2016",
    hover_cols=['Neighborhood', 'gross_rent', 'sale_price_spr_foot'],
    legend = 'bottom_right'
    )

warnings.filterwarnings('ignore')
SFO_geoplot

### Step 5: Use the interactive map to answer the following question:

**Question:** Which neighborhood has the highest gross rent, and which has the highest sale price per square foot?

**Answer:** The neighborhood with the highest gross rent is Westwood Park and the neighborhood with the highest sale price per square foot is the Union Square District.

## Compose Your Data Story

Based on the visualizations that you have created, compose a data story that synthesizes your analysis by answering the following questions:

**Question:**  How does the trend in rental income growth compare to the trend in sales prices? Does this same trend hold true for all the neighborhoods across San Francisco?

**Answer:** As you can see in Exhibits 2 and 3 below, the growth trend over time is much higher for rental income growth versus sales prices. Over the 6- year period from 2010-2016, sales prices increase almost 90% and gross rents increase by 254%. 

Not all the neighborhoods follow suit. Anza Vista and Hayes Valley showed a decline in sales prices over the 6 years.

**Question:** What insights can you share with your company about the potential one-click, buy-and-rent strategy that they're pursuing? Do neighborhoods exist that you would suggest for investment, and why?

**Answer:** There's a lot of factors that go into real estate purchases. However given the data we have for the years 2010-2016, here are my observations:
1. In general, SFO is a growing city with increasing prices to purchase and rent.
2. Gross rental rates are growing at a faster rate then price per square foot - but those are two different metrics.
3. As you can see in Exhibit 5 below, the cities with the fast growing prices are:
    Russian Hill                     1.826949
    Twin Peaks                       1.925857
    Downtown                         1.972553
    Outer Richmond                   2.539050
    Van Ness/ Civic Center           4.116078
4. One would have to compare the total cost verus the gross rent to understand the ROI on investments or the rent per square foot versus the cost per square foot. 

### *Additional Exibits Below*

In [48]:
#Data Story Exhibit 1
prices_square_foot_by_year

Unnamed: 0_level_0,sale_price_sqr_foot,gross_rent
year,Unnamed: 1_level_1,Unnamed: 2_level_1
2010,369.344353,1239.0
2011,341.903429,1530.0
2012,399.389968,2324.0
2013,483.600304,2971.0
2014,556.277273,3528.0
2015,632.540352,3739.0
2016,697.643709,4390.0


In [52]:
#Data Story Exhibit 2
pct_change = prices_square_foot_by_year.pct_change()
pct_change

Unnamed: 0_level_0,sale_price_sqr_foot,gross_rent
year,Unnamed: 1_level_1,Unnamed: 2_level_1
2010,,
2011,-0.074296,0.234867
2012,0.168137,0.518954
2013,0.210847,0.278399
2014,0.150283,0.187479
2015,0.137095,0.059807
2016,0.102924,0.174111


In [56]:
#Data Story Exhibit 3
pct_change_6 = prices_square_foot_by_year.pct_change(periods=6)
pct_change_6

Unnamed: 0_level_0,sale_price_sqr_foot,gross_rent
year,Unnamed: 1_level_1,Unnamed: 2_level_1
2010,,
2011,,
2012,,
2013,,
2014,,
2015,,
2016,0.888871,2.54318


In [184]:
#Data Story Exhibit 4
all_neighborhood_sales = sfo_data_df[['year','neighborhood','sale_price_sqr_foot']]
all_neighborhood_sales_pivoted = all_neighborhood_sales.pivot_table(
    index='neighborhood', columns='year', values='sale_price_sqr_foot')
all_neighborhood_sales_pivoted = all_neighborhood_sales_pivoted.dropna()
all_neighborhood_sales_pivoted.head()

year,2010,2011,2012,2013,2014,2015,2016
neighborhood,Unnamed: 1_level_1,Unnamed: 2_level_1,Unnamed: 3_level_1,Unnamed: 4_level_1,Unnamed: 5_level_1,Unnamed: 6_level_1,Unnamed: 7_level_1
Alamo Square,291.182945,272.52731,183.099317,387.794144,484.443552,602.115877,340.981835
Central Richmond,319.027623,261.6638,279.03014,343.054261,423.952727,473.163815,661.064426
Central Sunset,418.172493,391.42645,334.487873,330.736431,456.241996,453.379377,581.370875
Cow Hollow,569.379968,390.595653,644.818307,707.402809,691.865411,598.620541,1059.065602
Downtown,241.804552,234.281624,269.109529,361.454692,428.726533,485.886886,718.77683


In [183]:
#Data Story Exhibit 5
sales_pct_change_neighborhood = all_neighborhood_sales_pivoted.pct_change(axis="columns", periods=6)
display=sales_pct_change_neighborhood.iloc[:,6].sort_values()
display

neighborhood
Hayes Valley                    -0.357048
Park North                      -0.351870
Marina                           0.115004
Alamo Square                     0.171023
South of Market                  0.323553
Outer Sunset                     0.386264
Central Sunset                   0.390266
Inner Richmond                   0.411055
Potrero Hill                     0.439926
Inner Sunset                     0.617015
Jordan Park/Laurel Heights       0.677987
Outer Parkside                   0.695346
Lower Pacific Heights            0.701424
Telegraph Hill                   0.720772
Cow Hollow                       0.860033
Pacific Heights                  1.006762
Central Richmond                 1.072123
Haight Ashbury                   1.089940
Noe Valley                       1.097652
Nob Hill                         1.525399
Eureka Valley/Dolores Heights    1.547121
Inner Mission                    1.717728
Parnassus/Ashbury Heights        1.781770
Russian Hill         