# App Data Analysis
Goal: Using Google Play Store and Apple App Store data, determine which types of apps attract the most users.


Data sets I'll be using are as follows:

   1. https://www.kaggle.com/lava18/google-play-store-apps
   
   2. https://www.kaggle.com/ramamet4/app-store-apple-data-set-10k-apps

## First, let's open our datasets and define a function we'll use to easily explore the data.

In [2]:
from csv import reader

#Google
opened_file = open('googleplaystore.csv')
read_file = reader(opened_file)
android = list(read_file)
android_header = android[0]
android = android[1:]

#Apple
opened_file = open('applestore.csv')
read_file = reader(opened_file)
ios = list(read_file)
ios_header = ios[0]
ios = ios[1:]

In [3]:
# define function which returns a range of rows and also the number of rows and columns in the data.
def explore_data(dataset, start, end, rows_and_columns=False):
    dataset_slice = dataset[start:end]    
    for row in dataset_slice:
        print(row)
        print('\n') # adds a new (empty) line after each row

    if rows_and_columns:
        print('Number of rows:', len(dataset))
        print('Number of columns:', len(dataset[0]))

Next, let's explore our android data.

In [4]:
print(android_header)
print('\n')
explore_data(android, 0, 3, True)

['App', 'Category', 'Rating', 'Reviews', 'Size', 'Installs', 'Type', 'Price', 'Content Rating', 'Genres', 'Last Updated', 'Current Ver', 'Android Ver']


['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['Coloring book moana', 'ART_AND_DESIGN', '3.9', '967', '14M', '500,000+', 'Free', '0', 'Everyone', 'Art & Design;Pretend Play', 'January 15, 2018', '2.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


Number of rows: 10841
Number of columns: 13


We can see the field headers of our data in the first row, then the first three apps listed in the data.

Looks like we've got data on 10,841 apps across 13 fields. Metrics related to popularity are likely 'Rating', 'Reviews', and 'Installs'.

Next, let's explore our iOS data.

In [5]:
print(ios_header)
print('\n')
explore_data(ios, 0, 3, True)

['id', 'track_name', 'size_bytes', 'currency', 'price', 'rating_count_tot', 'rating_count_ver', 'user_rating', 'user_rating_ver', 'ver', 'cont_rating', 'prime_genre', 'sup_devices.num', 'ipadSc_urls.num', 'lang.num', 'vpp_lic']


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'Instagram', '113954816', 'USD', '0.0', '2161558', '1289', '4.5', '4.0', '10.23', '12+', 'Photo & Video', '37', '0', '29', '1']


['529479190', 'Clash of Clans', '116476928', 'USD', '0.0', '2130805', '579', '4.5', '4.5', '9.24.12', '9+', 'Games', '38', '5', '18', '1']


Number of rows: 7197
Number of columns: 16


Looks like we've got 7,197 apps across 16 fields. The field names aren't the most self explanatory, so I looked here: https://www.kaggle.com/ramamet4/app-store-apple-data-set-10k-apps to determine what each field holds. Metrics related to popularity are likely 'rating_count_tot' and 'user_rating'. Notably, there's no data included on number of installs.

## Now, let's get started cleaning the data. 

## First, let's remove inaccurate data.

In the discussion section of the google play data is mention of a wrong rating for entry 10472: https://www.kaggle.com/lava18/google-play-store-apps/discussion/66015.

In [6]:
print(android_header)
print(android[10472])
print(android[0])

['App', 'Category', 'Rating', 'Reviews', 'Size', 'Installs', 'Type', 'Price', 'Content Rating', 'Genres', 'Last Updated', 'Current Ver', 'Android Ver']
['Life Made WI-Fi Touchscreen Photo Frame', '1.9', '19', '3.0M', '1,000+', 'Free', '0', 'Everyone', '', 'February 11, 2018', '1.0.19', '4.0 and up']
['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


Yes, looks like row 10472 is missing the category field. The 'Life Made WI-Fi Touchscreen Photo Frame' app is likely not in the '1.9' category. **Let's delete this row.**

In [7]:
del android[10472]

In [8]:
print(len(android))

10840


Previously was 10,841 rows, so looks like our delete was completed successfully. 

## Next, let's look for any possible duplicate entries. 

Looking at the discussion page (https://www.kaggle.com/lava18/google-play-store-apps/discussion), seems like Instagram might be one app with duplicate entries.

In [9]:
for i in android:
    name = i[0]
    if name == 'Instagram':
        print(i)

['Instagram', 'SOCIAL', '4.5', '66577313', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']
['Instagram', 'SOCIAL', '4.5', '66577446', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']
['Instagram', 'SOCIAL', '4.5', '66577313', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']
['Instagram', 'SOCIAL', '4.5', '66509917', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']


Indeed, Instagram has duplicate entries. It looks like the 4th column (number of reviews) is slightly different for each row. Let's look for all cases of duplicate entries.

In [10]:
duplicate_apps = []
unique_apps = []

for i in android:
    name = i[0]
    if name in unique_apps:
        duplicate_apps.append(name)
    else:
        unique_apps.append(name)

print('Number of duplicate apps:', len(duplicate_apps))
print('\n')
print('Examples of duplicate apps:', duplicate_apps[:15])

Number of duplicate apps: 1181


Examples of duplicate apps: ['Quick PDF Scanner + OCR FREE', 'Box', 'Google My Business', 'ZOOM Cloud Meetings', 'join.me - Simple Meetings', 'Box', 'Zenefits', 'Google Ads', 'Google My Business', 'Slack', 'FreshBooks Classic', 'Insightly CRM', 'QuickBooks Accounting: Invoicing & Expenses', 'HipChat - Chat Built for Teams', 'Xero Accounting Software']


Now, let's delete duplicate entries. We'll only keep the records with the largest amounts of ratings for each app.

In [11]:
reviews_max = {} #this will hold the the name of the app and most reviews associated with that app for all entries of the app

for i in android:
    name = i[0]
    n_reviews = float(i[3])
    if name in reviews_max and reviews_max[name] < n_reviews:
        reviews_max[name] = n_reviews
    elif name not in reviews_max:
        reviews_max[name] = n_reviews

print('Expected length: ', len(android) - 1181)
print('Actual length:', len(reviews_max))

Expected length:  9659
Actual length: 9659


In [12]:
#Now, we'll loop through the data to return only the records referenced in our dictionary

android_clean = []
already_added = [] # necessary for if multiple duplicate entries have the same number of reviews. we'll just keep the first if the app has already been added.

for i in android:
    name = i[0]
    n_reviews = float(i[3])
    if (n_reviews == reviews_max[name]) and (name not in already_added):
        android_clean.append(i)
        already_added.append(name)

explore_data(android_clean, 0,  3, True)

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 9659
Number of columns: 13


Looks good! We have 9659 rows as expected. Let's check for duplicate entries in the iOS data.

In [13]:
duplicate_apps = []
unique_apps = []

for i in ios:
    name = i[0]
    if name in unique_apps:
        duplicate_apps.append(name)
    else:
        unique_apps.append(name)

print('Number of duplicate apps:', len(duplicate_apps))

Number of duplicate apps: 0


## Now, let's remove non-English apps.

See below for some examples:

In [14]:
print(ios[813][1])
print(ios[6731][1])

print(android_clean[4412][0])
print(android_clean[7940][0])

爱奇艺PPS -《欢乐颂2》电视剧热播
【脱出ゲーム】絶対に最後までプレイしないで 〜謎解き＆ブロックパズル〜
中国語 AQリスニング
لعبة تقدر تربح DZ


We'll be using the ord function to find non-English characters. If a string contains a character with an ord value greater than 127, the string is likely not in English (because it's outside ASCII range).

In [15]:
def is_english(string):
    for character in string:
        if ord(character) > 127:
            return False

    return True

print(is_english('Instagram'))
print(is_english('爱奇艺PPS -《欢乐颂2》电视剧热播'))
print(is_english('Docs To Go™ Free Office Suite'))
print(is_english('Instachat 😜'))

True
False
False
False


Hmm. Well,  let's slightly adjust our criteria so we can identify 'TM' and 😜 as English. We'll say that the app name must have more than 3 characters with an ord value greater than 127 to be considered not English.

In [16]:
def is_english(string):
    non_ASCII = 0
    for character in string:
        if ord(character) > 127:
            non_ASCII += 1
        
    if non_ASCII > 3:
        return False
    else:
        return True

print(is_english('Instagram'))
print(is_english('爱奇艺PPS -《欢乐颂2》电视剧热播'))
print(is_english('Docs To Go™ Free Office Suite'))
print(is_english('Instachat 😜'))

True
False
True
True


Now let's use this function to remove non-English apps.

In [17]:
android_clean_english = []
ios_clean_english = []

for row in android_clean:
    name = row[0]
    if is_english(name):
        android_clean_english.append(row)

for row in ios:
    name = row[1]
    if is_english(name):
        ios_clean_english.append(row)

explore_data(android_clean_english, 0, 3, True)
print('\n')
explore_data(ios_clean_english, 0, 3, True)

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 9614
Number of columns: 13


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'Instagram', '113954816', 'USD', '0.0', '2161558', '1289', '4.5', '4.0', '10.23', '12+', 'Photo & Video', '37', '0', '29', '1']


['529479190', 'Clash of Clans', '116476928', 'USD', '0.0', '2130805', '579', '4.5', '4.5', '9.24.12', '9+', 

We have 9,614 andorid apps remaining and 6,183 iOS apps remaining. 

## Our last data cleaning step will be isolating the free apps.

In [18]:
android_free = []
ios_free = []

for app in android_clean_english:
    price = app[7]
    if price == '0':
        android_free.append(app)

for app in ios_clean_english:
    price = app[4]
    if price == '0.0':
        ios_free.append(app)

explore_data(android_free, 0, 3, True)
print('\n')
explore_data(ios_free, 0, 3, True)

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 8864
Number of columns: 13


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'Instagram', '113954816', 'USD', '0.0', '2161558', '1289', '4.5', '4.0', '10.23', '12+', 'Photo & Video', '37', '0', '29', '1']


['529479190', 'Clash of Clans', '116476928', 'USD', '0.0', '2130805', '579', '4.5', '4.5', '9.24.12', '9+', 

We're left with 8,864 android apps and 3,222 iOS apps.

## Now, let's analyze our data.

As stated previously, our goal is to determine the kinds of apps that are likely to attract the most users in the android and iOS markets. Let's begin by getting a sense of what the genres of app are most developed in each market. The relevant genre columns are 'Category' and 'Genres' in andorid and 'prime_genre' in iOS. We'll build a frequency table to summarize number of apps in the app store by the unqiue values in each of these fields.

In [19]:
def freq_table(dataset, index):
    table = {}
    total = 0
    for row in dataset:
        value = row[index]
        total +=1
        if value in table:
            table[value] += 1
        else:
            table[value] = 1
    
    table_percentages = {}
    
    for i in table:
        percentage = (table[i] / total) * 100
        table_percentages[i] = percentage
    return table_percentages

In [20]:
def display_table(dataset, index):
    table = freq_table(dataset, index)
    table_display = []
    for key in table:
        key_val_as_tuple = (table[key], key)
        table_display.append(key_val_as_tuple)

    table_sorted = sorted(table_display, reverse = True)
    for entry in table_sorted:
        print(entry[1], ':', entry[0])

In [21]:
display_table(android_free, 1)

FAMILY : 18.907942238267147
GAME : 9.724729241877256
TOOLS : 8.461191335740072
BUSINESS : 4.591606498194946
LIFESTYLE : 3.9034296028880866
PRODUCTIVITY : 3.892148014440433
FINANCE : 3.7003610108303246
MEDICAL : 3.531137184115524
SPORTS : 3.395758122743682
PERSONALIZATION : 3.3167870036101084
COMMUNICATION : 3.2378158844765346
HEALTH_AND_FITNESS : 3.0798736462093865
PHOTOGRAPHY : 2.944494584837545
NEWS_AND_MAGAZINES : 2.7978339350180503
SOCIAL : 2.6624548736462095
TRAVEL_AND_LOCAL : 2.33528880866426
SHOPPING : 2.2450361010830324
BOOKS_AND_REFERENCE : 2.1435018050541514
DATING : 1.861462093862816
VIDEO_PLAYERS : 1.7937725631768955
MAPS_AND_NAVIGATION : 1.3989169675090252
FOOD_AND_DRINK : 1.2409747292418771
EDUCATION : 1.1620036101083033
ENTERTAINMENT : 0.9589350180505415
LIBRARIES_AND_DEMO : 0.9363718411552346
AUTO_AND_VEHICLES : 0.9250902527075812
HOUSE_AND_HOME : 0.8235559566787004
WEATHER : 0.8009927797833934
EVENTS : 0.7107400722021661
PARENTING : 0.6543321299638989
ART_AND_DESIGN : 

Within android free, English app, on the 'Category' column, the most popular app category is 'Family' at close to 19%, then 'Game' at close to 10%, and 'Tools' in third at around 8%. It looks like a pretty even spread from there on out, with 20 categories ranging from 1-5% that include 'Business', 'Finance', and 'Medical'. Although the 'Family' category is mostly games for children according to internet research, these apps seem evenly spread between practical, productivity categories.

In [22]:
display_table(android_free, 9)

Tools : 8.449909747292418
Entertainment : 6.069494584837545
Education : 5.347472924187725
Business : 4.591606498194946
Productivity : 3.892148014440433
Lifestyle : 3.892148014440433
Finance : 3.7003610108303246
Medical : 3.531137184115524
Sports : 3.463447653429603
Personalization : 3.3167870036101084
Communication : 3.2378158844765346
Action : 3.1024368231046933
Health & Fitness : 3.0798736462093865
Photography : 2.944494584837545
News & Magazines : 2.7978339350180503
Social : 2.6624548736462095
Travel & Local : 2.3240072202166067
Shopping : 2.2450361010830324
Books & Reference : 2.1435018050541514
Simulation : 2.0419675090252705
Dating : 1.861462093862816
Arcade : 1.8501805054151623
Video Players & Editors : 1.7712093862815883
Casual : 1.7599277978339352
Maps & Navigation : 1.3989169675090252
Food & Drink : 1.2409747292418771
Puzzle : 1.128158844765343
Racing : 0.9927797833935018
Role Playing : 0.9363718411552346
Libraries & Demo : 0.9363718411552346
Auto & Vehicles : 0.9250902527075

Within android free, English app, on the 'Genres' column, the story looks similar. The most popular app category is 'Tools' at 8%, then 'Entertainment' at 6%, and 'Education' in third at around 5%. It looks much more granular - there are many more categories than previously. It looks like a pretty even spread though, with over 20 genres ranging from 1-5%. Weirdly, a category related to games doesn't appear until 20 genres down with 'Arcade' at 1.9%.

In [23]:
display_table(ios_free, 11)

Games : 58.16263190564867
Entertainment : 7.883302296710118
Photo & Video : 4.9658597144630665
Education : 3.662321539416512
Social Networking : 3.2898820608317814
Shopping : 2.60707635009311
Utilities : 2.5139664804469275
Sports : 2.1415270018621975
Music : 2.0484171322160147
Health & Fitness : 2.0173805090006205
Productivity : 1.7380509000620732
Lifestyle : 1.5828677839851024
News : 1.3345747982619491
Travel : 1.2414649286157666
Finance : 1.1173184357541899
Weather : 0.8690254500310366
Food & Drink : 0.8069522036002483
Reference : 0.5586592178770949
Business : 0.5276225946617008
Book : 0.4345127250155183
Navigation : 0.186219739292365
Medical : 0.186219739292365
Catalogs : 0.12414649286157665


Within free, English apps, the most popular category with over 58% of apps is the 'Games' category. This is unlike andorid, where we saw a more even distribution of apps across . categories. The second most popular app category is 'Entertainment' at close to 8%. My general impression is that apps designed for fun, i.e. games and intertainment, are more numerous on the app store. However, and this is a reminder for both android and iOS - this isn't data on the number of downloads per category. So, although these categories of apps are most popularly developed, they're not necessarily the most popular categories downloaded. There are also generally much less categories of apps on iOS compared to android.

Seems to me that we've learned the google play store apps are generally evenly spread between games / entertainment and productivity apps. The iOS store is different - games dominate the app store, with productivity apps trailing far behind.

Now. let's average the number of installs (android) and ratings (iOS - doesn't have an install-like field) for apps per category.

In [38]:
ios_categories = freq_table(ios_free, 11)

table = {}
for genre in ios_categories:
    count_of_genre_user_ratings = 0
    number_of_apps = 0
    for row in ios_free:
        genre_app = row[11]
        if genre_app == genre:
            count_of_user_ratings = float(row[5])
            count_of_genre_user_ratings += count_of_user_ratings
            number_of_apps += 1
    average_count_of_user_ratings = count_of_genre_user_ratings / number_of_apps
    table[genre] = average_count_of_user_ratings

In [39]:
def sort_table(dataset):
    table = dataset
    table_display = []
    for key in table:
        key_val_as_tuple = (table[key], key)
        table_display.append(key_val_as_tuple)

    table_sorted = sorted(table_display, reverse = True)
    for entry in table_sorted:
        print(entry[1], ':', entry[0])

sort_table(table)

Navigation : 86090.33333333333
Reference : 74942.11111111111
Social Networking : 71548.34905660378
Music : 57326.530303030304
Weather : 52279.892857142855
Book : 39758.5
Food & Drink : 33333.92307692308
Finance : 31467.944444444445
Photo & Video : 28441.54375
Travel : 28243.8
Shopping : 26919.690476190477
Health & Fitness : 23298.015384615384
Sports : 23008.898550724636
Games : 22788.6696905016
News : 21248.023255813954
Productivity : 21028.410714285714
Utilities : 18684.456790123455
Lifestyle : 16485.764705882353
Entertainment : 14029.830708661417
Business : 7491.117647058823
Education : 7003.983050847458
Catalogs : 4004.0
Medical : 612.0


So this tells a different story from our previous frequency tables. For one, 'Games' aren't until the 14th category down. The category with the largest average number or ratings is 'Navigation', with ~86,000, followed by 'Reference' at ~75,000 and then 'social networking' with ~72,000. 