# Exploring Hacker News Posts

Hacker News is a site started by the startup incubator Y Combinator, where user-submitted stories (known as "posts") are voted and commented upon, similar to reddit. Hacker News is extremely popular in technology and startup circles, and posts that make it to the top of Hacker News' listings can get hundreds of thousands of visitors as a result.

You can find the data set [here](https://www.kaggle.com/hacker-news/hacker-news-posts), but note that it has been reduced from almost 300,000 rows to approximately 20,000 rows by removing all submissions that did not receive any comments, and then randomly sampling from the remaining submissions. Below are descriptions of the columns:

- id: The unique identifier from Hacker News for the post
title: The title of the post
- url: The URL that the posts links to, if it the post has a URL
- num_points: The number of points the post acquired, calculated as the total number of upvotes minus the total number of downvotes
- num_comments: The number of comments that were made on the post
- author: The username of the person who submitted the post
- created_at: The date and time at which the post was submitted

We're specifically interested in posts whose titles begin with either ***Ask HN*** or ***Show HN***. Users submit ***Ask HN*** posts to ask the Hacker News community a specific question. Below are a couple examples:

```
Ask HN: How to improve my personal website?
Ask HN: Am I the only one outraged by Twitter shutting down share counts?
Ask HN: Aby recent changes to CSS that broke mobile?
```

Likewise, users submit ***Show HN*** posts to show the Hacker News community a project, product, or just generally something interesting. Below are a couple of examples:

```
Show HN: Wio Link  ESP8266 Based Web of Things Hardware Development Platform'
Show HN: Something pointless I made
Show HN: Shanhu.io, a programming playground powered by e8vm
```

We'll compare these two types of posts to determine the following:

- Do Ask HN or Show HN receive more comments on average?
- Do posts created at a certain time receive more comments on average?

In [17]:
from csv import reader

opened_file = open('hacker_news.csv')
read_file = reader(opened_file)
full = list(read_file)
headers = full[0]
hn = full[1:]

In [18]:
ask_posts = []
show_posts = []
other_posts = []

In [19]:
for row in hn:
    title = row[1].lower()
    if title.startswith('ask'):
        ask_posts.append(row)
    elif title.startswith('show'):
        show_posts.append(row)
    else:
        other_posts.append(row)
        
        

In [20]:
print(len(ask_posts), len(show_posts), len(other_posts))

1756 1164 17180


In [24]:
total_ask_comments = 0

for row in ask_posts:
    num_comments = int(row[4])
    total_ask_comments += num_comments
    
avg_ask_comments = total_ask_comments / len(ask_posts)

In [26]:
avg_ask_comments

14.203302961275627

In [27]:
total_show_comments = 0

for row in show_posts:
    num_comments = int(row[4])
    total_show_comments += num_comments
    
avg_show_comments = total_show_comments / len(show_posts)

In [28]:
avg_show_comments

10.323024054982818

From the result above, we can conclude that on average, ***ask HN*** receives more comments than ***show HN***.

Now that we know ***ask HN*** receives more comments, we will focus on ask posts than show posts on what time it receives the most comments.

In [29]:
import datetime as dt

In [33]:
result_list = []

for row in ask_posts:
    result_list.append([row[6],
                    int(row[4])])

In [39]:
counts_by_hour = {}
comments_by_hour = {}

for row in result_list:
    date = dt.datetime.strptime(row[0],"%m/%d/%Y %H:%M")
    if date.hour not in counts_by_hour:
        counts_by_hour[date.hour] = 1
        comments_by_hour[date.hour] = row[1]
    else:
        counts_by_hour[date.hour] += 1
        comments_by_hour[date.hour] += row[1]

In [40]:
counts_by_hour

{0: 55,
 1: 60,
 2: 58,
 3: 55,
 4: 48,
 5: 46,
 6: 44,
 7: 34,
 8: 48,
 9: 45,
 10: 60,
 11: 58,
 12: 73,
 13: 86,
 14: 110,
 15: 116,
 16: 109,
 17: 100,
 18: 109,
 19: 112,
 20: 80,
 21: 109,
 22: 71,
 23: 70}

In [41]:
comments_by_hour

{0: 447,
 1: 683,
 2: 1381,
 3: 459,
 4: 339,
 5: 464,
 6: 397,
 7: 267,
 8: 492,
 9: 251,
 10: 794,
 11: 641,
 12: 687,
 13: 1254,
 14: 1420,
 15: 4477,
 16: 1949,
 17: 1146,
 18: 1439,
 19: 1295,
 20: 1722,
 21: 1745,
 22: 479,
 23: 713}

In [42]:
avg_by_hour = []

for hour in comments_by_hour:
    avg_by_hour.append([hour, 
                        comments_by_hour[hour]/ counts_by_hour[hour]])
    
avg_by_hour

[[0, 8.127272727272727],
 [1, 11.383333333333333],
 [2, 23.810344827586206],
 [3, 8.345454545454546],
 [4, 7.0625],
 [5, 10.08695652173913],
 [6, 9.022727272727273],
 [7, 7.852941176470588],
 [8, 10.25],
 [9, 5.5777777777777775],
 [10, 13.233333333333333],
 [11, 11.051724137931034],
 [12, 9.41095890410959],
 [13, 14.581395348837209],
 [14, 12.909090909090908],
 [15, 38.5948275862069],
 [16, 17.880733944954127],
 [17, 11.46],
 [18, 13.20183486238532],
 [19, 11.5625],
 [20, 21.525],
 [21, 16.009174311926607],
 [22, 6.746478873239437],
 [23, 10.185714285714285]]

It's little hard to see which posts within certain hour gets the most comments. Let's sort it. 

In [44]:
swap_avg_by_hour = []

for row in avg_by_hour:
    swap_avg_by_hour.append([row[1], row[0]])

In [47]:
sorted_swap = sorted(swap_avg_by_hour, reverse = True)

Top 5 Hours for Ask Posts Comments

In [49]:
sorted_swap[0:5]

[[38.5948275862069, 15],
 [23.810344827586206, 2],
 [21.525, 20],
 [17.880733944954127, 16],
 [16.009174311926607, 21]]

In [51]:
template = "{hour}:00: {avg_com:.2f} average comments per post"

In [55]:
for row in sorted_swap[0:5]:
    print(template.format(hour = row[1], avg_com = row[0]))

15:00: 38.59 average comments per post
2:00: 23.81 average comments per post
20:00: 21.52 average comments per post
16:00: 17.88 average comments per post
21:00: 16.01 average comments per post


From the result above we can conclude that if you want to get the most response, it's probably 3pm would be the best time to post in Hacker News. 