In this project, we'll work with a data set of submissions to popular technology site Hacker News.

We are wroking with submissions that received any comments. Below are descriptions of the columns:

* id: The unique identifier from Hacker News for the post
* title: The title of the post
* url: The URL that the posts links to, if it the post has a URL
* num_points: The number of points the post acquired, calculated as the total number of upvotes minus the total number of downvotes
* num_comments: The number of comments that were made on the post
* author: The username of the person who submitted the post
* created_at: The date and time at which the post was submitted

We're specifically interested in posts whose titles begin with either Ask HN or Show HN. Users submit Ask HN posts to ask the Hacker News community a specific question. Below are a couple examples:

1. Ask HN: How to improve my personal website?
2. Ask HN: Am I the only one outraged by Twitter shutting down share counts?
3. Ask HN: Aby recent changes to CSS that broke mobile?

Likewise, users submit Show HN posts to show the Hacker News community a project, product, or just generally something interesting. Below are a couple of examples:

1. Show HN: Wio Link ESP8266 Based Web of Things Hardware Development Platform'
2. Show HN: Something pointless I made
3. Show HN: Shanhu.io, a programming playground powered by e8vm

We'll compare these two types of posts to determine the following:

1. Do Ask HN or Show HN receive more comments on average?
2. Do posts created at a certain time receive more comments on average?

In [21]:
from csv import reader

hn = list(reader(open("hacker_news.csv", encoding = "utf-8")))
print(len(hn))

20101


In [22]:
headers = hn[0]
hn = hn[1:]

In [23]:
# converting created_at column to datetime data type
import datetime as dt

date_format = "%m/%d/%Y %H:%M"

for i in hn:
    i[-1] = dt.datetime.strptime(i[-1],date_format)

In [24]:
# Identify posts that begin with either `Ask HN` or `Show HN` 
# and separate the data into different lists.

ask_posts = []
show_posts = [] 
other_posts = []

for i in hn:
    title = i[1]
    if title.lower().startswith("ask hn"):
        ask_posts.append(i)
    elif title.lower().startswith("show hn"):
        show_posts.append(i)
    else:
        other_posts.append(i)

In [25]:
print(len(ask_posts))
print(len(show_posts))
print(len(other_posts))

1744
1162
17194


In [26]:
# Calculating the Average Number of Comments for Ask HN

total_ask_comments = 0
no_of_comments = 0

for i in ask_posts:
    no_of_comments += 1
    num_comments = int(i[4])
    total_ask_comments += num_comments

avg_ask_comments = total_ask_comments/no_of_comments

print(avg_ask_comments)

14.038417431192661


In [27]:
# Calculating the Average Number of Comments for Show HN

total_show_comments = 0
no_of_comments = 0

for i in show_posts:
    no_of_comments += 1
    num_comments = int(i[4])
    total_show_comments += num_comments
    
avg_show_comments = total_show_comments/no_of_comments
print(avg_show_comments)

10.31669535283993


In [36]:
# Calculate the amount of ask posts created during each hour of day
# and the number of comments received.

counts_by_hour = {}   
comments_by_hour = {}

for i in ask_posts:
    hour = i[-1].hour
    comment = int(i[4])
    
    if hour not in (counts_by_hour and comments_by_hour):
        counts_by_hour[hour] = 1
        comments_by_hour[hour] = comment
    else:
        counts_by_hour[hour] += 1
        comments_by_hour[hour] += comment



In [37]:
print(counts_by_hour)

{9: 45, 13: 85, 10: 59, 14: 107, 16: 108, 23: 68, 12: 73, 17: 100, 15: 116, 21: 109, 20: 80, 2: 58, 18: 109, 3: 54, 5: 46, 19: 110, 1: 60, 22: 71, 8: 48, 4: 47, 0: 55, 6: 44, 7: 34, 11: 58}


In [38]:
print(comments_by_hour)

{9: 251, 13: 1253, 10: 793, 14: 1416, 16: 1814, 23: 543, 12: 687, 17: 1146, 15: 4477, 21: 1745, 20: 1722, 2: 1381, 18: 1439, 3: 421, 5: 464, 19: 1188, 1: 683, 22: 479, 8: 492, 4: 337, 0: 447, 6: 397, 7: 267, 11: 641}


In [46]:
# Calculating the Average Number of Comments for Ask HN Posts by Hour
avg_by_hour = []

for hour in comments_by_hour:
    total_comments_per_hour = comments_by_hour[hour]
    posts_per_hour = counts_by_hour[hour]
    avg_by_hour.append([hour,round((total_comments_per_hour/posts_per_hour),2)])

In [47]:
print(avg_by_hour)

[[9, 5.58], [13, 14.74], [10, 13.44], [14, 13.23], [16, 16.8], [23, 7.99], [12, 9.41], [17, 11.46], [15, 38.59], [21, 16.01], [20, 21.52], [2, 23.81], [18, 13.2], [3, 7.8], [5, 10.09], [19, 10.8], [1, 11.38], [22, 6.75], [8, 10.25], [4, 7.17], [0, 8.13], [6, 9.02], [7, 7.85], [11, 11.05]]


In [53]:
hour_sorted_by_comments = sorted(avg_by_hour, key = lambda lst: lst[1], reverse = True)

In [57]:
print("Top 5 Hours for Ask Posts Comments:")

for hours, avg in hour_sorted_by_comments[:5]:
    hour = dt.datetime.strptime(str(hours), "%H").strftime("%H:%M")
    print(f'{avg} average comments per post at {hour} hours')

Top 5 Hours for Ask Posts Comments:
38.59 average comments per post at 15:00 hours
23.81 average comments per post at 02:00 hours
21.52 average comments per post at 20:00 hours
16.8 average comments per post at 16:00 hours
16.01 average comments per post at 21:00 hours


The hour that receives the most comments per post on average is 15:00, with an average of 38.59 comments per post. There's about a 60% increase in the number of comments between the hours with the highest and second highest average number of comments.

# Conclusion

In this project, we analyzed ask posts and show posts to determine which type of post and time receive the most comments on average. Based on our analysis, to maximize the amount of comments a post receives, we'd recommend the post be categorized as ask post and created between 15:00 and 16:00 (3:00 pm est - 4:00 pm est).