# Guided Project: Exploring Hacker News Posts

![hacker news logo](https://s3.amazonaws.com/dq-content/354/hacker_news.jpg)

In this project, we'll work with a data set of submissions to popular technology site [Hacker News](https://news.ycombinator.com/). Hacker News is a site started by the startup incubator [Y Combinator](https://www.ycombinator.com/), where user-submitted stories (known as "posts") are voted and commented upon, in a similar way as as reddit. Hacker News is highly popular in technology and startup circles, and posts that make it to the top of Hacker News' listings can get hundreds of thousands of visitors as a result.

You can find the data set [here](https://www.kaggle.com/hacker-news/hacker-news-posts), but note that it has been reduced from almost 300,000 rows to approximately 20,000 rows by removing all submissions that did not receive any comments, and then randomly sampling from the remaining submissions. Below are descriptions of the columns:

* `id`: The unique identifier from Hacker News for the post
* `title`: The title of the post
* `url`: The URL that the posts links to, if it the post has a URL
* `num_points`: The number of points the post acquired, calculated as the total number of upvotes minus the total number of downvotes
* `num_comments`: The number of comments that were made on the post
* `author`: The username of the person who submitted the post
* `created_at`: The date and time at which the post was submitted


We're specifically interested in posts whose titles begin with either `Ask HN` or `Show HN`. Users submit `Ask HN` posts to ask the Hacker News community a specific question. Below are a couple examples:

* Ask HN: How to improve my personal website?
* Ask HN: Am I the only one outraged by Twitter shutting down share counts?
* Ask HN: Any recent changes to CSS that broke mobile?


Likewise, users submit `Show HN` posts to show the Hacker News community a project, product, or just generally something interesting. Below are a couple of examples:

* Show HN: Wio Link  ESP8266 Based Web of Things Hardware Development Platform'
* Show HN: Something pointless I made
* Show HN: Shanhu.io, a programming playground powered by e8vm

We'll compare these two types of posts to determine the following:

* Do `Ask HN` or `Show HN` receive more comments on average?
* Do posts created at a certain time receive more comments on average?

Let's start by importing the libraries we need and reading the data set into a list of lists.

## Open and explore the data set

First, let's open the data set. As mentioned before, it can be found on Kaggle using [this link](https://www.kaggle.com/hacker-news/hacker-news-posts). For ease of use, I have renamed the file to "hacker_news.csv".

In [15]:
from csv import reader
opened_file = open("hacker_news.csv")
read_file = reader(opened_file)
hn = list(read_file) # a list of lists of the full data set
headers = hn[0]      # to display the headers separately
hn_rows = hn[1:]     # to access the data without the headers

Let's see how many entries there are in this data set.

In [16]:
print("There are {:,.0f} entries in the data set.".format(len(hn_rows)))

There are 293,119 entries in the data set.


Let's also explore the colums for this data set:

In [17]:
print("There are {} columns:".format(len(headers)))
print(headers)

There are 7 columns:
['id', 'title', 'url', 'num_points', 'num_comments', 'author', 'created_at']


The first five rows look as follows:

In [18]:
for row in hn_rows[:5]:
    print('\n')
    print(row)



['12579008', 'You have two days to comment if you want stem cells to be classified as your own', 'http://www.regulations.gov/document?D=FDA-2015-D-3719-0018', '1', '0', 'altstar', '9/26/2016 3:26']


['12579005', 'SQLAR  the SQLite Archiver', 'https://www.sqlite.org/sqlar/doc/trunk/README.md', '1', '0', 'blacksqr', '9/26/2016 3:24']


['12578997', 'What if we just printed a flatscreen television on the side of our boxes?', 'https://medium.com/vanmoof/our-secrets-out-f21c1f03fdc8#.ietxmez43', '1', '0', 'pavel_lishin', '9/26/2016 3:19']


['12578989', 'algorithmic music', 'http://cacm.acm.org/magazines/2011/7/109891-algorithmic-composition/fulltext', '1', '0', 'poindontcare', '9/26/2016 3:16']


['12578979', 'How the Data Vault Enables the Next-Gen Data Warehouse and Data Lake', 'https://www.talend.com/blog/2016/05/12/talend-and-Â\x93the-data-vaultÂ\x94', '1', '0', 'markgainor1', '9/26/2016 3:14']


## Extracting 'Ask HN' and 'Show HN' posts
Now that we've removed the headers from `hn` (the list of lists of the full data set), we're ready to filter our data further. Since we're mostly concerned with post titles beginning with `Ask HN` or `Show HN`, we'll create new lists of lists containing just the data for those titles, as well as a third list for other posts.

In [19]:
ask_posts = []
show_posts = []
other_posts = []

for row in hn:
    title = row[1]
    if title.lower().startswith('ask hn'):
        ask_posts.append(row)
    elif title.lower().startswith('show hn'):
        show_posts.append(row)
    else:
        other_posts.append(row)
        
print("There are {:,.0f} number of 'Ask HN' posts.".format(len(ask_posts)))
print("There are {:,.0f} number of 'Show HN' posts.".format(len(show_posts)))
print("There are {:,.0f} number of posts in the other category.".format(len(other_posts)))

There are 9,139 number of 'Ask HN' posts.
There are 10,158 number of 'Show HN' posts.
There are 273,823 number of posts in the other category.


## Calculating the average number of comments for 'Ask HN' and 'Show HN' posts

Now that we have separated the `Ask HN` and the `Show HN` posts into two list of lists, we will determine which of these receive more comments on average.

In [27]:
total_ask_comments = 0
total_show_comments = 0

for row in ask_posts:
    num_ask_comments = int(row[4])
    total_ask_comments += num_ask_comments 
    
for row in show_posts:
    num_show_comments = int(row[4])
    total_show_comments += num_show_comments 
    
avg_ask_comments = round(total_ask_comments / len(ask_posts), 2)
avg_show_comments = round(total_show_comments / len(show_posts), 2)

print("'Ask HN' posts have an average of {} comments per post.".format(avg_ask_comments))
print("'Show HN' posts have an average of {} comments per post.".format(avg_show_comments))

'Ask HN' posts have an average of 10.39 comments per post.
'Show HN' posts have an average of 4.89 comments per post.


Because the average number of comments is markedly higher in the `Ask HN` posts than in the `Show HN` posts, we can conclude that posts containing questions receive more engagement from the Hacker News community. 

Since `Ask HN` posts are more likely to receive comments, we'll focus our remaining analysis just on these posts. Next, we'll determine if posts created at a certain time are more likely to attract comments.

## Finding the amount of 'Ask HN' posts and comments by hour created

First, we will calculate the amount of `Ask HN` posts created in each hour of the day, along with the number of comments they receive.

In [28]:
results_list = []

for row in ask_posts:
    created_at = row[6]
    n_comments = int(row[4])
    results_list.append([created_at, n_comments])

Next, we will make two dictionaries to calculate the amount of posts and their number of comments. We will use the `datetime` module to work with the data in the `created_at` column.

In [29]:
import datetime as dt

counts_by_hour = {}
comments_by_hour = {}
date_format = "%m/%d/%Y %H:%M"

for row in results_list:
    date = dt.datetime.strptime(row[0], date_format)
    hour = date.strftime("%H")
    n_comments = row[1]
    if hour in counts_by_hour:
        counts_by_hour[hour] += 1
        comments_by_hour[hour] += n_comments
    else:
        counts_by_hour[hour] = 1
        comments_by_hour[hour] = n_comments

comment= 'Counts by hour:\n{}\nComments by hour:\n{}\n' 
print(comment.format(counts_by_hour, comments_by_hour))

Counts by hour:
{'02': 269, '01': 282, '22': 383, '21': 518, '19': 552, '17': 587, '15': 646, '14': 513, '13': 444, '11': 312, '10': 282, '09': 222, '07': 226, '03': 271, '23': 343, '20': 510, '16': 579, '08': 257, '00': 301, '18': 614, '12': 342, '04': 243, '06': 234, '05': 209}
Comments by hour:
{'02': 2996, '01': 2089, '22': 3372, '21': 4500, '19': 3954, '17': 5547, '15': 18525, '14': 4972, '13': 7245, '11': 2797, '10': 3013, '09': 1477, '07': 1585, '03': 2154, '23': 2297, '20': 4462, '16': 4466, '08': 2362, '00': 2277, '18': 4877, '12': 4234, '04': 2360, '06': 1587, '05': 1838}



Next, we will calculate the average number of comments `Ask HN` posts receive by hour created.

## Calculating the average number of comments for 'Ask HN' posts by hour

Using the two dictionaries created above, we will now calculate the average number of comments for `Ask HN` posts created during each hour of the day. 

In [23]:
avg_by_hour = []

for hour in counts_by_hour:
    avg = round(comments_by_hour[hour]/counts_by_hour[hour], 2)
    avg_by_hour.append([hour, avg])

sorted(avg_by_hour)

[['00', 7.56],
 ['01', 7.41],
 ['02', 11.14],
 ['03', 7.95],
 ['04', 9.71],
 ['05', 8.79],
 ['06', 6.78],
 ['07', 7.01],
 ['08', 9.19],
 ['09', 6.65],
 ['10', 10.68],
 ['11', 8.96],
 ['12', 12.38],
 ['13', 16.32],
 ['14', 9.69],
 ['15', 28.68],
 ['16', 7.71],
 ['17', 9.45],
 ['18', 7.94],
 ['19', 7.16],
 ['20', 8.75],
 ['21', 8.69],
 ['22', 8.8],
 ['23', 6.7]]

Although we now have the results we need, this format makes it hard to identify the hours with the highest values. 
The easiest way to solve this would be to change the order of items in the list and then sort on the `avg` item.

In [31]:
avg_by_hour_sorted = []

for hour in counts_by_hour:
    avg = round(comments_by_hour[hour]/counts_by_hour[hour], 2)
    avg_by_hour_sorted.append([avg, hour])

avg_by_hour_s = sorted(avg_by_hour_sorted, reverse = True)
avg_by_hour_s

[[28.68, '15'],
 [16.32, '13'],
 [12.38, '12'],
 [11.14, '02'],
 [10.68, '10'],
 [9.71, '04'],
 [9.69, '14'],
 [9.45, '17'],
 [9.19, '08'],
 [8.96, '11'],
 [8.8, '22'],
 [8.79, '05'],
 [8.75, '20'],
 [8.69, '21'],
 [7.95, '03'],
 [7.94, '18'],
 [7.71, '16'],
 [7.56, '00'],
 [7.41, '01'],
 [7.16, '19'],
 [7.01, '07'],
 [6.78, '06'],
 [6.7, '23'],
 [6.65, '09']]

Let's finish by sorting the list of lists and printing the five highest values in a format that's easier to read.

In [32]:
print("Top 5 Hours for 'Ask HN' Posts Comments:")

for row in avg_by_hour_s[:5]:
    date = dt.datetime.strptime(row[1], "%H")
    time = date.strftime("%H:%M")
    print("{}: {:.2f} average comments per post".format(time, row[0]))

Top 5 Hours for 'Ask HN' Posts Comments:
15:00: 28.68 average comments per post
13:00: 16.32 average comments per post
12:00: 12.38 average comments per post
02:00: 11.14 average comments per post
10:00: 10.68 average comments per post


## Conclusion
From the analysis above, we can conclude that `Ask HN` posts have received the most comments on average between 15:00-16:00 GMT. Another time slot on which many comments have been made was between 12:00-14:00 GMT. 

Fewest comments were made on average on `Ask HN` posts that were placed between 6:00-8:00 GMT and between 00:00-02:00 GMT. It is advised to avoid these time slots when trying to obtain the maximum amount of comments, based on averages from the dataset. A remarkable outlier is 2:00, which is the timeslot that receives the 4th most comments. Perhaps, this is because the comments are posted by users from different time zones.

Please note that these times are expressed in GMT and need to be converted to one's local time zone. This [site](http://www.worldtimeserver.com) can help you with this.