# Profitable App Profiles for the App Store and Google Play Markets

The aim of this project is to find mobile app profiles that are profitable for the App store and Google Play markets. Our role in this project is to work as data analyst for a company who build Android and ios mobile apps. We help the team of app developers to make data driven decisions with respect to the apps they build.

Our company only builds apps that are free to download and install and our main source of revenue is the in-app ads. Therefore, Our revenue depends on the number of users who use our app. The goal of this project is to analyze data to figure out what kind of apps attract most users in both markets.


## Opening and Exploring the data

As of September 2018, there were approximately 2 million ios apps available on the App store and 2.1 million Android apps available on Google Play. Collecting data for over 4 million apps requires lot of time and money,so we will try to analyze a sample of data instead. To avoid spending resources with collecting new data ourselves , we should always try to see whether we can find any relevant existing data at no cost. There are 2 datasets which is already available in kaggle for this purpose. 

1. [Googleplay dataset](https://www.kaggle.com/lava18/google-play-store-apps) contains data about approximately ten thousand android apps from Google play.
2. [Appstore dataset](https://www.kaggle.com/ramamet4/app-store-apple-data-set-10k-apps) contains data about approximately seven thousand iOS apps from the App Store.

We will begin by opening the two datasets and then start to explore data

In [1]:
from csv import reader
def open_file(filename):
    opened_file=open(filename, errors='ignore', encoding='utf8')
    read_file=reader(opened_file)
    apps_data=list(read_file)
    apps_header=apps_data[0]
    apps_data=apps_data[1:]
    
    return apps_header,apps_data





In [2]:

apps_ios_header,apps_ios_data=open_file('AppleStore.csv')
apps_android_header,apps_android_data=open_file('googleplaystore.csv')

In [3]:
def explore_data(dataset,start,end, rows_and_columns=False):
    dataset_slice=dataset[start:end]
    for row in dataset_slice:
        print(row)
        print('\n')
    if rows_and_columns :
        print('Number of rows:',len(dataset))
        print('number of columns:',len(dataset[0]))
print(apps_android_header)
print('\n')
explore_data(apps_android_data,0, 3, True)

        
    

['App', 'Category', 'Rating', 'Reviews', 'Size', 'Installs', 'Type', 'Price', 'Content Rating', 'Genres', 'Last Updated', 'Current Ver', 'Android Ver']


['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['Coloring book moana', 'ART_AND_DESIGN', '3.9', '967', '14M', '500,000+', 'Free', '0', 'Everyone', 'Art & Design;Pretend Play', 'January 15, 2018', '2.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


Number of rows: 10841
number of columns: 13


From the results we can see that Google play data set has 10841 rows and 13 columns. The columns that are helpful for our analysis would be 
1. 'App': 'App' contains app name
2. 'Category': 'Category' contains information about what kind of apps. since our aim is to find out what kind of apps attracts the most users, this is an important column
3. Reviews: 'Reviews' shows how how many users have used the app and their reviews on it
4. Installs: 'Installs' shows how many users have installed the app. This is very significant information for our reserach because the revenue of the app depends on number of users who installed it
5. 'Type': 'Type' indicates whether app is free or not 
6. 'Price': 'Price' tells whether the app is free or not. Here we are looking mostly into free apps
7. 'Genres':'Genres' indicate what category the apps belongs to

From analysis we can say that 'Price' and 'Type' almost contains same information and you can choose any one according to your convenience. I would choose 'Type' because it is easy to filter. Similarily 'Category' and 'Genres' relate to same information. You can choose either 'Category' or 'Genres'. I would choose 'Category' because the number of categories is less compared to Genres. Categories give the relevant information for our analysis.


In [4]:
print(apps_ios_header)
print('\n')
explore_data(apps_ios_data,0, 3, True)

['id', 'track_name', 'size_bytes', 'currency', 'price', 'rating_count_tot', 'rating_count_ver', 'user_rating', 'user_rating_ver', 'ver', 'cont_rating', 'prime_genre', 'sup_devices.num', 'ipadSc_urls.num', 'lang.num', 'vpp_lic']


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'Instagram', '113954816', 'USD', '0.0', '2161558', '1289', '4.5', '4.0', '10.23', '12+', 'Photo & Video', '37', '0', '29', '1']


['529479190', 'Clash of Clans', '116476928', 'USD', '0.0', '2130805', '579', '4.5', '4.5', '9.24.12', '9+', 'Games', '38', '5', '18', '1']


Number of rows: 7197
number of columns: 16


From the results we can see that App Store dataset contains 7197 rows and 16 columns. The columns that are helpful for our analysis would be :
1. 'track_name': 'track_name' contains the name of app
2. 'price': 'price' indicates whether app is free or not
3. rating_count_tot: 'rating_count_tot' can be considered as the number of installs per app
4. prime_genre: 'prime_genre' contains the genre in which that app belongs to


## Deleting Wrong Data

The Google Play dtaaset has a [discussion section](https://www.kaggle.com/lava18/google-play-store-apps/discussion) and as per one of the discussions there is an error for row 10472. Let us print this row and compare it against the header and another row that is correct.

In [5]:
print(apps_android_data[10472])   # incorrect row
print('\n')
print(apps_android_header)        # header
print('\n')
print(apps_android_data[0])

['Life Made WI-Fi Touchscreen Photo Frame', '1.9', '19', '3.0M', '1,000+', 'Free', '0', 'Everyone', '', 'February 11, 2018', '1.0.19', '4.0 and up']


['App', 'Category', 'Rating', 'Reviews', 'Size', 'Installs', 'Type', 'Price', 'Content Rating', 'Genres', 'Last Updated', 'Current Ver', 'Android Ver']


['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


The row 10472 corresponds to the app 'Life Made Wi-Fi Touchscreen Photo Frame' has recorded the wrong rating. The row shows  a rating of 19 which is wrong because maximum rating of google play apps is 5(as mentioned in the [discussions section](https://www.kaggle.com/lava18/google-play-store-apps/discussion/66015)). This problem is caused because there is a missing value. Therefore we will delete this row. 

In [6]:
print(len(apps_android_data))
del apps_android_data[10472]   # run only once otherwise it would delete more rows
print(len(apps_android_data))

10841
10840


The above results show that we have deleted the wrong data entry and now the total number of rows is 10480 i.e we have data of 10840 samples (before it was 10841 rows)

While reading the [discussions of App store dataset](https://www.kaggle.com/ramamet4/app-store-apple-data-set-10k-apps/discussion) we understand there is no wrong data for App Store dataset. However people mention about duplicates but it seems that the apps are different.

## Removing Duplicate Entries

### Section 1: Finding duplicate rows

When we explore google play dataset we can find there are some apps that have multiple entries. This means that there are duplicate entries for  particular app and it is not good to have duplicate entries in a dataset.For example, if we look at the app Instagram, we can check how many entries of instagram is available. From discussions we understand that there are duplicates for Instagram. Let us check and see whether data has duplicates for instagram. In code, we check whether the app name is instagram. If the app name is instagram print the entire row.


In [7]:
# checking the number of duplicates for Instagram
for apps in apps_android_data:
    app_name=apps[0]
    if app_name =='Instagram':
        print(apps)
    

['Instagram', 'SOCIAL', '4.5', '66577313', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']
['Instagram', 'SOCIAL', '4.5', '66577446', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']
['Instagram', 'SOCIAL', '4.5', '66577313', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']
['Instagram', 'SOCIAL', '4.5', '66509917', 'Varies with device', '1,000,000,000+', 'Free', '0', 'Teen', 'Social', 'July 31, 2018', 'Varies with device', 'Varies with device']


From the above code we found that there are 4 duplicates for 'Instagram'. Now lets find how many duplicate rows are available in the data. Generally a data analyst would not be aware of which duplicates are available. In the previous case we came to know 'Instagram' has duplicate rows through  [discussions](https://www.kaggle.com/lava18/google-play-store-apps). One cannot go through thousands of rows and identify which data is duplicate. The process is tedious and we might miiss some duplicates too. But there is an easy way to count the number of duplicates and identify which data has duplicates through code. The code for counting the number of duplicate rows is given below.



In [8]:
unique_apps=[]
duplicate_apps=[]
for apps in apps_android_data:
    app_name=apps[0]
    if app_name in unique_apps:
        duplicate_apps.append(app_name)
    else:
        unique_apps.append(app_name)
print("The number of duplicates available is :", len(duplicate_apps))
print('\n')
print("Examples of duplicate apps:", duplicate_apps[:15])       

The number of duplicates available is : 1181


Examples of duplicate apps: ['Quick PDF Scanner + OCR FREE', 'Box', 'Google My Business', 'ZOOM Cloud Meetings', 'join.me - Simple Meetings', 'Box', 'Zenefits', 'Google Ads', 'Google My Business', 'Slack', 'FreshBooks Classic', 'Insightly CRM', 'QuickBooks Accounting: Invoicing & Expenses', 'HipChat - Chat Built for Teams', 'Xero Accounting Software']


There are a total of 1181 duplicates in the dataset. We dont want to count certain apps more than once when we analyze data and therefore we have to remove the duplicate entries and keep o nly one entry per app. We can either remove the duplicates randomly or by applying any condition relevant to data.

For the instagram data we found that there are 4 duplicates. This means that we can only keep one which is relevant and most significant and delete other rows. For instagram data we can see that the duplicates are only different by its number of reviews. The different numbers shows that the data might be collected at different times. Lets consider the row with most reviews because this might be the latest data. i.e we will only keep Instagram data with 66577446 reviews and delete other rows.

Similarily lets check another app 'Slack' for duplicate rows. From the above code output we found that Slack has duplicate rows. Lets identify that and see what crieteria can we use for removing the duplicate rows.

In [9]:
# checking duplicates
for apps in apps_android_data:
    app_name=apps[0]
    if app_name =='Slack':
        print(apps)

['Slack', 'BUSINESS', '4.4', '51507', 'Varies with device', '5,000,000+', 'Free', '0', 'Everyone', 'Business', 'August 2, 2018', 'Varies with device', 'Varies with device']
['Slack', 'BUSINESS', '4.4', '51507', 'Varies with device', '5,000,000+', 'Free', '0', 'Everyone', 'Business', 'August 2, 2018', 'Varies with device', 'Varies with device']
['Slack', 'BUSINESS', '4.4', '51510', 'Varies with device', '5,000,000+', 'Free', '0', 'Everyone', 'Business', 'August 2, 2018', 'Varies with device', 'Varies with device']


From checking the duplicates for 'slack' we found that 'Slack' also has the same issue like 'instagram'. So we can use the same strategy of considering the row with the highest reviews and deleting the other rows.

After finding the number of duplicates we need to eliminate the duplicate entries. As discussed above this can be achieved by keeping only the entries with the maximum number of reviews and discarding other rows. For removing the duplicates we have to create an empty dictionary where each dictionary key is a unique app name and the corresponding dictionary value is the highest number of reviews of that app. With the information stored in the dictionary we create a new dataset which have only one entry per app i.e the entry with maximum reviews if there  are duplicates for that app.

### Section 2: Removing duplicate rows

In [10]:
#create an empty dictionary that maps the app name to the maximum review value of that app
reviews_max={}
# we are discarding the header row since data is from second row
for apps in apps_android_data:
    name=apps[0]
    n_reviews=float(apps[3])
    # To check whether the app name already exists in reviews_max
    # and check whether the value stored in the dictionary is the max number of reviews
    # if not update the number of reviews for that entry in the dictionary
    if name in reviews_max and reviews_max[name]< n_reviews:
        reviews_max[name]=n_reviews
    # if name not in reviews_max dict create an entry for that app name
    elif name not in reviews_max:
        reviews_max[name]=n_reviews

        
        

The code above create an empty dictionary named reviews_max and checks whether the app name already exists in reviews_max and check whether the value stored in the dictionary is the max number of reviews. If the condition is wrong we update the number of reviews for that entry in the dictionary. The above code snippet just identifies the app names and what are the maximum reviews for that app.

In [11]:
print("Expected length:", len(apps_android_data)-1181)
print("Actual length:", len(reviews_max))

Expected length: 9659
Actual length: 9659


Once we figured out the app names and its maximum reviews its time to eliminate the rows with reviews not equal to maximum reviews in the reviews_max dictionary. For that we create an empty list named 'android_clean' to create a list of rows whose number of reviews for an app name equals n_reviews of 'reviews_max' dictionary. We also create an empty list 'already_added' that contains names of apps that we already added.

In [12]:
android_clean=[]
already_added=[]
for apps in apps_android_data:
    name=apps[0]
    n_reviews=float(apps[3])
    if n_reviews==reviews_max[name] and name not in already_added:
        android_clean.append(apps)
        already_added.append(name)
print(len(android_clean))
print(len(already_added))

9659
9659


The above results show that in the new list we only have data with maximum reviews for apps. This means that the android_clean list contains only rows that does not have duplicates. We have found that the actual length of dataset after removing duplicates is 9659. By verifying with the result from the code above, we found that length of android_clean list is 9659. Therefore our analysis is correct. To explore the new dataset android_clean we can use explore_data function.

In [13]:
explore_data(android_clean, 0,3, True)

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 9659
number of columns: 13


While reading the [discussions of App store dataset](https://www.kaggle.com/ramamet4/app-store-apple-data-set-10k-apps/discussion) there are some discussion on duplicates and then it was clarified that the apps are different and hence no duplicates. But for our confirmation let's check the number of duplicates in App store dataset. From the results below we can confirm that there are no duplicate entries for app store dataset. 

In [14]:
#checking for duplicates in apps_data_apple
unique_apps=[]
duplicate_apps=[]
for apps in apps_ios_data:
    id=apps[0]
    if id in unique_apps:
        duplicate_apps.append(id)
    else:
        unique_apps.append(id)
print(" The number of unique apps:" , len(unique_apps))
print(" The number of duplicate entries :", len(duplicate_apps))

 The number of unique apps: 7197
 The number of duplicate entries : 0


## Removing Non-english Apps

### Section 1:

If we explore the datasets we can see some of the Apps are non-English apps. Our aim is to build English apps. So we can rwemove these non-English apps. 

One way to remove the non English apps is to remove each apps whose name contains  a symbol that is not commonly used in english text. English text usually includes letters from English alphabet, numbers composed of digits from 0 to 9, punctuation marks (., !, ?, ;, etc.), and other symbols (+, /, etc.). All these characters that are specific to English text are encoded using the ASCII standard. Each ASCII character has a corresponding number between 0 and 127 associated with it, and we can take advantage of that to build a function that checks an app name and tells us whether it contains non-ASCII characters.

The function **is_english** checks whether the string is english or not. The built in function ord() returns the ASCII encoding of that character. The English characters has an ASCII encoding between 0 and 127.

In [15]:
def is_english(name_string):
   
    for character in name_string:
        if ord(character) >127: 
            return False
     
    return True
print(is_english('Instagram'))
print(is_english('爱奇艺PPS -《欢乐颂2》电视剧热播'))    

True
False


In [16]:
def is_english(string):
    non_ascii = 0
    
    for character in string:
        if ord(character) > 127:
            non_ascii += 1
    
    if non_ascii > 3:
        return False
    else:
        return True
print(is_english('Instagram'))
print(is_english('爱奇艺PPS -《欢乐颂2》电视剧热播'))
print(is_english('Docs To Go™ Free Office Suite'))
print(is_english('Instachat 😜'))

True
False
True
True


In [17]:
android_english = []
ios_english = []

for app in android_clean:
    name = app[0]
    if is_english(name):
        android_english.append(app)
        
for app in apps_ios_data:
    name = app[1]
    if is_english(name):
        ios_english.append(app)
        
explore_data(android_english, 0, 3, True)
print('\n')
explore_data(ios_english, 0, 3, True)

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 9614
number of columns: 13


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'Instagram', '113954816', 'USD', '0.0', '2161558', '1289', '4.5', '4.0', '10.23', '12+', 'Photo & Video', '37', '0', '29', '1']


['529479190', 'Clash of Clans', '116476928', 'USD', '0.0', '2130805', '579', '4.5', '4.5', '9.24.12', '9+', 

The above function **is_english** seems to work fine, but some English app names use emojis or other symbols (™, — (em dash), – (en dash), etc.) that fall outside of the ASCII range. Because of this, we'll remove useful apps if we use the function in its current form. Therefore we are going to use another filter condition to remove non-English apps. 

In the below function, we check the count of english characters in an app name and if the app name contains more than three non-engish characters we would conclude that the app is a non-English app. This function is not perfect because sometimes we might get some non-English apps. But we are ignoring this currently.

## Isolating the free apps

Since our aim is to build only apps that are free to download and install, we are only considering the free apps and the non-free apps will be removed. Below, we seperate the free apps for both datasets.

In [18]:
android_final = []
ios_final = []

for app in android_english:
    price = app[7]
    if price == '0':
        android_final.append(app)
        
for app in ios_english:
    price = app[4]
    if price == '0.0':
        ios_final.append(app)
        
print(len(android_final))
print(len(ios_final))

8864
3222


Our aim is to build apps that are free to download and install and we are only aiming at building english apps. We have performed data cleaning 

1. For removing inaccurate date from the googleplaystore.csv dataset. There was no inaccurate data in AppleStore.csv
2. For removing duplicate app entries in both datasets: we found there were **1181** duplicate entries in googleplaystore.csv and **0** duplicates for AppleStore.csv
3. Removing non-English apps in both datasets: we found that there were **9614 english apps in android_clean and 6183 english apps in apps_ios_data**

After isolating free apps in the above code we found that there are **8864 free english apps in googleplaystore.csv and 3222 free english apps in AppleStore.csv**

Our aim is to determine the kinds of apps that attract more number of users because our revenue depends on number of users who use the app.The validation strategy for an app idea is to
1. Build a minimum version of android app and add it to googleplay
2. If the app develops good response the app will be developed further
3. If the app is profitable after 6 months, we build an ios version of the app and add it to App Store

Since our end goal is to add the app on both google play and App store we need to identify which genre of apps attracts the most users in both googleplay and appstore.

## Most common apps by genre

We'll build two functions we can use to analyze the frequency tables:

1. One function to generate frequency tables that show percentages
2. Another function that we can use to display the percentages in a descending order



In [19]:
def freq_table(dataset,index):
    frequency_column={}
    frequency_percentage={}
    total_apps=len(dataset)
    for apps in dataset:
        column=apps[index]
        if column in frequency_column:
            frequency_column[column]+=1
        
        else:
            frequency_column[column]=1
    for column in frequency_column:
        frequency_percentage[column]=(frequency_column[column]/total_apps)*100
    return frequency_percentage

        
    
    

In [20]:
# display the frequency table 
def display_table(dataset, index):
    table = freq_table(dataset, index)
    table_display = []
    for key in table:
        key_val_as_tuple = (table[key], key)
        table_display.append(key_val_as_tuple)

    table_sorted = sorted(table_display, reverse = True)
    for entry in table_sorted:
        print(entry[1], ':', entry[0])

In [21]:
display_table(ios_final,-5)

Games : 58.16263190564867
Entertainment : 7.883302296710118
Photo & Video : 4.9658597144630665
Education : 3.662321539416512
Social Networking : 3.2898820608317814
Shopping : 2.60707635009311
Utilities : 2.5139664804469275
Sports : 2.1415270018621975
Music : 2.0484171322160147
Health & Fitness : 2.0173805090006205
Productivity : 1.7380509000620732
Lifestyle : 1.5828677839851024
News : 1.3345747982619491
Travel : 1.2414649286157666
Finance : 1.1173184357541899
Weather : 0.8690254500310366
Food & Drink : 0.8069522036002483
Reference : 0.5586592178770949
Business : 0.5276225946617008
Book : 0.4345127250155183
Navigation : 0.186219739292365
Medical : 0.186219739292365
Catalogs : 0.12414649286157665


We can see that among the free English apps, more than a half (58.16%) are games. Entertainment apps are close to 8%, followed by photo and video apps, which are close to 5%. Only 3.66% of the apps are designed for education, followed by social networking apps which amount for 3.29% of the apps in our data set.

The general impression is that App Store (at least the part containing free English apps) is dominated by apps that are designed for fun (games, entertainment, photo and video, social networking, sports, music, etc.), while apps with practical purposes (education, shopping, utilities, productivity, lifestyle, etc.) are more rare. However, the fact that fun apps are the most numerous doesn't also imply that they also have the greatest number of users — the demand might not be the same as the offer.

Let's continue by examining the Genres and Category columns of the Google Play data set (two columns which seem to be related).

In [22]:
display_table(android_final,-4)

Tools : 8.449909747292418
Entertainment : 6.069494584837545
Education : 5.347472924187725
Business : 4.591606498194946
Productivity : 3.892148014440433
Lifestyle : 3.892148014440433
Finance : 3.7003610108303246
Medical : 3.531137184115524
Sports : 3.463447653429603
Personalization : 3.3167870036101084
Communication : 3.2378158844765346
Action : 3.1024368231046933
Health & Fitness : 3.0798736462093865
Photography : 2.944494584837545
News & Magazines : 2.7978339350180503
Social : 2.6624548736462095
Travel & Local : 2.3240072202166067
Shopping : 2.2450361010830324
Books & Reference : 2.1435018050541514
Simulation : 2.0419675090252705
Dating : 1.861462093862816
Arcade : 1.8501805054151623
Video Players & Editors : 1.7712093862815883
Casual : 1.7599277978339352
Maps & Navigation : 1.3989169675090252
Food & Drink : 1.2409747292418771
Puzzle : 1.128158844765343
Racing : 0.9927797833935018
Role Playing : 0.9363718411552346
Libraries & Demo : 0.9363718411552346
Auto & Vehicles : 0.9250902527075

When looking at genres in android_final, we see that genres column is more granular. But we can see that distribution of common apps is much different from ios_final. Lets look at category column to gain more information.

In [23]:
display_table(android_final,1)

FAMILY : 18.907942238267147
GAME : 9.724729241877256
TOOLS : 8.461191335740072
BUSINESS : 4.591606498194946
LIFESTYLE : 3.9034296028880866
PRODUCTIVITY : 3.892148014440433
FINANCE : 3.7003610108303246
MEDICAL : 3.531137184115524
SPORTS : 3.395758122743682
PERSONALIZATION : 3.3167870036101084
COMMUNICATION : 3.2378158844765346
HEALTH_AND_FITNESS : 3.0798736462093865
PHOTOGRAPHY : 2.944494584837545
NEWS_AND_MAGAZINES : 2.7978339350180503
SOCIAL : 2.6624548736462095
TRAVEL_AND_LOCAL : 2.33528880866426
SHOPPING : 2.2450361010830324
BOOKS_AND_REFERENCE : 2.1435018050541514
DATING : 1.861462093862816
VIDEO_PLAYERS : 1.7937725631768955
MAPS_AND_NAVIGATION : 1.3989169675090252
FOOD_AND_DRINK : 1.2409747292418771
EDUCATION : 1.1620036101083033
ENTERTAINMENT : 0.9589350180505415
LIBRARIES_AND_DEMO : 0.9363718411552346
AUTO_AND_VEHICLES : 0.9250902527075812
HOUSE_AND_HOME : 0.8235559566787004
WEATHER : 0.8009927797833934
EVENTS : 0.7107400722021661
PARENTING : 0.6543321299638989
ART_AND_DESIGN : 

From analysing the frequency tables of both applestore and googleplaystore, we can understand that in AppStore , English free apps designed for fun was dominant and in googleplaystore both practical and fun apps are almost in equal percent. Since the company can get revenue only from maximum number of users, frequency table alone will not give us much idea about which apps can contribute to the revenue of the company. Therefore we should consider the number of installs for each app or determine the most popular app to get a better idea of what genre of app would be the best for market.

## Most Popular Apps by Genre

### Section 1: App store

To determine the apps with most downloads we have to look at the number of installs information in the App store. But the App store data does not have a column related to number of installs. As a workaround we will take the total number of user ratings as a proxy, which we can find in the rating_count_tot app. We find the most popular app by calculating the average number of userratings per app genre. The app genre with the highest average user rating would be the most popular genre.


The code below finds the average user rating per app genre  and gives result in the descending order.

In [24]:
genres_ios = freq_table(ios_final, -5)
genre_list=[]
# iterating through each genre in genre_freq dictionary
for genre in genres_ios:
    # total to store the sum of user ratings for a particular genre
    total=0
    # len_genre stores number of apps in that genre
    len_genre=0
    # loop over the appstore dataset
    for apps in ios_final:
        # retrieving the genre of app
        genre_app=apps[-5]
        # checking whether genre matches genre in genre_freq
        if genre_app==genre:
            n_user_ratings=float(apps[5])
            # take sum of all userratings of that app
            total+=n_user_ratings
            # increment number of apps specific to that genre
            len_genre+=1
    # once you calculated total userratings and number of apps in that genre 
    # calculate avg_user_ratings
    avg_user_ratings=total/len_genre
    # to receive the output as a sorted list we made a tuple 
    # of avg-user_ratings and genre
    # This will create  a list of tuples
   
    genre_tuple=(avg_user_ratings,genre)
    genre_list.append(genre_tuple)
# to sort the genre_list   
genre_list_sorted=sorted(genre_list, reverse=True) 
for entry in genre_list_sorted:
    print(entry[1], ':', entry[0])
    

Navigation : 86090.33333333333
Reference : 74942.11111111111
Social Networking : 71548.34905660378
Music : 57326.530303030304
Weather : 52279.892857142855
Book : 39758.5
Food & Drink : 33333.92307692308
Finance : 31467.944444444445
Photo & Video : 28441.54375
Travel : 28243.8
Shopping : 26919.690476190477
Health & Fitness : 23298.015384615384
Sports : 23008.898550724636
Games : 22788.6696905016
News : 21248.023255813954
Productivity : 21028.410714285714
Utilities : 18684.456790123455
Lifestyle : 16485.764705882353
Entertainment : 14029.830708661417
Business : 7491.117647058823
Education : 7003.983050847458
Catalogs : 4004.0
Medical : 612.0


From the results we can recommend that Navigation apps has the highest number of user reviews. To have  a better understanding lets look at rows with app_genre 'Navigation'.

In [26]:
for apps in ios_final:
    if apps[-5]=='Navigation':
        print(apps[1], ':', apps[5]) # print name and number of ratings

Waze - GPS Navigation, Maps & Real-time Traffic : 345046
Google Maps - Navigation & Transit : 154911
Geocaching® : 12811
CoPilot GPS – Car Navigation & Offline Maps : 3582
ImmobilienScout24: Real Estate Search in Germany : 187
Railway Route Search : 5


On average, navigation apps have the highest number of user reviews, but this figure is heavily influenced by Waze and Google Maps, which have close to half a million user reviews together.The same pattern applies to social networking apps, where the average number is heavily influenced by a few giants like Facebook, Pinterest, Skype, etc. Same applies to music apps, where a few big players like Pandora, Spotify, and Shazam heavily influence the average number.

Our aim is to find popular genres, but navigation, social networking or music apps might seem more popular than they really are. The average number of ratings seem to be skewed by very few apps which have hundreds of thousands of user ratings, while the other apps may struggle to get past the 10,000 threshold. We could get a better picture by removing these extremely popular apps for each genre and then rework the averages, but we'll leave this level of detail for later.

In [27]:
for apps in ios_final:
    if apps[-5]=='Reference':
        print(apps[1], ':', apps[5])

Bible : 985920
Dictionary.com Dictionary & Thesaurus : 200047
Dictionary.com Dictionary & Thesaurus for iPad : 54175
Google Translate : 26786
Muslim Pro: Ramadan 2017 Prayer Times, Azan, Quran : 18418
New Furniture Mods - Pocket Wiki & Game Tools for Minecraft PC Edition : 17588
Merriam-Webster Dictionary : 16849
Night Sky : 12122
City Maps for Minecraft PE - The Best Maps for Minecraft Pocket Edition (MCPE) : 8535
LUCKY BLOCK MOD ™ for Minecraft PC Edition - The Best Pocket Wiki & Mods Installer Tools : 4693
GUNS MODS for Minecraft PC Edition - Mods Tools : 1497
Guides for Pokémon GO - Pokemon GO News and Cheats : 826
WWDC : 762
Horror Maps for Minecraft PE - Download The Scariest Maps for Minecraft Pocket Edition (MCPE) Free : 718
VPN Express : 14
Real Bike Traffic Rider Virtual Reality Glasses : 8
教えて!goo : 0
Jishokun-Japanese English Dictionary & Translator : 0


Reference apps have 74,942 user ratings on average, but it's actually the Bible and Dictionary.com which skew up the average rating.However, this niche seems to show some potential. One thing we could do is take another popular book and turn it into an app where we could add different features besides the raw version of the book. This might include daily quotes from the book, an audio version of the book, quizzes about the book, etc. On top of that, we could also embed a dictionary within the app, so users don't need to exit our app to look up words in an external app.

This idea seems to fit well with the fact that the App Store is dominated by for-fun apps. This suggests the market might be a bit saturated with for-fun apps, which means a practical app might have more of a chance to stand out among the huge number of apps on the App Store.

Other genres that seem popular include weather, book, food and drink, or finance. The book genre seem to overlap a bit with the app idea we described above, but the other genres don't seem too interesting to us:

 - Weather apps — people generally don't spend too much time in-app, and the chances of making profit from in-app adds are low. Also, getting reliable live weather data may require us to connect our apps to non-free APIs.

 - Food and drink — examples here include Starbucks, Dunkin' Donuts, McDonald's, etc. So making a popular food and drink app requires actual cooking and a delivery service, which is outside the scope of our company.

 - Finance apps — these apps involve banking, paying bills, money transfer, etc. Building a finance app requires domain knowledge, and we don't want to hire a finance expert just to build an app [1] 
 
The above analysis was taken from [dataquest solution notebook](https://github.com/dataquestio/solutions/blob/master/Mission350Solutions.ipynb)

Now let's analyze the Google Play market a bit.

### Section 2: Google play

For google play dataset we have information about number of installs. So we are using number of installs column in google play dataset to gain insights on most popular app genre. While we look at the 'Installs' column of google play dataset we can see that the number of installs are expressed as 1,000,000+, 100,000 + , etc. So we need to remove that '+' from the string. Here instead of expressing the number of installs as float the data was collected as string. We need to remove '+' from number of installs and convert the number of installs into float data type. 

The code below computes average number of installs per app genre:

In [29]:
# To calculate average number of installs per app genre
categories_android = freq_table(android_final, 1)
category_android_list=[]
for category in categories_android:
    total=0
    len_category=0
    
    for apps in android_final:
        category_app=apps[1]
        if category_app==category:
            n_installs=apps[5]
            n_installs=n_installs.replace('+','')
            n_installs=n_installs.replace(',', '')
            
            n_installs=float(n_installs)
            total+=n_installs
            len_category+=1
         
    avg_n_installs=total/len_category
    category_tuple=(avg_n_installs,category)
    category_android_list.append(category_tuple)
category_android_sorted=sorted(category_android_list, reverse=True)
for entry in category_android_sorted:
    print(entry[1], ':', entry[0])
    

COMMUNICATION : 38456119.167247385
VIDEO_PLAYERS : 24727872.452830188
SOCIAL : 23253652.127118643
PHOTOGRAPHY : 17840110.40229885
PRODUCTIVITY : 16787331.344927534
GAME : 15588015.603248259
TRAVEL_AND_LOCAL : 13984077.710144928
ENTERTAINMENT : 11640705.88235294
TOOLS : 10801391.298666667
NEWS_AND_MAGAZINES : 9549178.467741935
BOOKS_AND_REFERENCE : 8767811.894736841
SHOPPING : 7036877.311557789
PERSONALIZATION : 5201482.6122448975
WEATHER : 5074486.197183099
HEALTH_AND_FITNESS : 4188821.9853479853
MAPS_AND_NAVIGATION : 4056941.7741935486
FAMILY : 3695641.8198090694
SPORTS : 3638640.1428571427
ART_AND_DESIGN : 1986335.0877192982
FOOD_AND_DRINK : 1924897.7363636363
EDUCATION : 1833495.145631068
BUSINESS : 1712290.1474201474
LIFESTYLE : 1437816.2687861272
FINANCE : 1387692.475609756
HOUSE_AND_HOME : 1331540.5616438356
DATING : 854028.8303030303
COMICS : 817657.2727272727
AUTO_AND_VEHICLES : 647317.8170731707
LIBRARIES_AND_DEMO : 638503.734939759
PARENTING : 542603.6206896552
BEAUTY : 51315

From analyzing the results of the category with the highest number of installs, we can find that category 'Communication' has the highest number of installs in googleplay market. This number is heavily skewed up by a few apps that have over one billion installs (WhatsApp, Facebook Messenger, Skype, Google Chrome, Gmail, and Hangouts), and a few others with over 100 and 500 million installs:



In [30]:
for app in android_final:
    if app[1] == 'COMMUNICATION' and (app[5] == '1,000,000,000+'
                                      or app[5] == '500,000,000+'
                                      or app[5] == '100,000,000+'):
        print(app[0], ':', app[5])

WhatsApp Messenger : 1,000,000,000+
imo beta free calls and text : 100,000,000+
Android Messages : 100,000,000+
Google Duo - High Quality Video Calls : 500,000,000+
Messenger – Text and Video Chat for Free : 1,000,000,000+
imo free video calls and chat : 500,000,000+
Skype - free IM & video calls : 1,000,000,000+
Who : 100,000,000+
GO SMS Pro - Messenger, Free Themes, Emoji : 100,000,000+
LINE: Free Calls & Messages : 500,000,000+
Google Chrome: Fast & Secure : 1,000,000,000+
Firefox Browser fast & private : 100,000,000+
UC Browser - Fast Download Private & Secure : 500,000,000+
Gmail : 1,000,000,000+
Hangouts : 1,000,000,000+
Messenger Lite: Free Calls & Messages : 100,000,000+
Kik : 100,000,000+
KakaoTalk: Free Calls & Text : 100,000,000+
Opera Mini - fast web browser : 100,000,000+
Opera Browser: Fast and Secure : 100,000,000+
Telegram : 100,000,000+
Truecaller: Caller ID, SMS spam blocking & Dialer : 100,000,000+
UC Browser Mini -Tiny Fast Private & Secure : 100,000,000+
Viber Mess

If we removed all the communication apps that have over 100 million installs, the average would be reduced roughly ten times:

In [31]:
under_100_m = []

for app in android_final:
    n_installs = app[5]
    n_installs = n_installs.replace(',', '')
    n_installs = n_installs.replace('+', '')
    if (app[1] == 'COMMUNICATION') and (float(n_installs) < 100000000):
        under_100_m.append(float(n_installs))
        
sum(under_100_m) / len(under_100_m)

3603485.3884615386

We see the same pattern for the video players category, which is the runner-up with 24,727,872 installs. The market is dominated by apps like Youtube, Google Play Movies & TV, or MX Player. The pattern is repeated for social apps (where we have giants like Facebook, Instagram, Google+, etc.), photography apps (Google Photos and other popular photo editors), or productivity apps (Microsoft Word, Dropbox, Google Calendar, Evernote, etc.).

Again, the main concern is that these app genres might seem more popular than they really are. Moreover, these niches seem to be dominated by a few giants who are hard to compete against.

The game genre seems pretty popular, but previously we found out this part of the market seems a bit saturated, so we'd like to come up with a different app recommendation if possible.

The books and reference genre looks fairly popular as well, with an average number of installs of 8,767,811. It's interesting to explore this in more depth, since we found this genre has some potential to work well on the App Store, and our aim is to recommend an app genre that shows potential for being profitable on both the App Store and Google Play.

Let's take a look at some of the apps from this genre and their number of installs:

In [32]:
for app in android_final:
    if app[1] == 'BOOKS_AND_REFERENCE':
        print(app[0], ':', app[5])

E-Book Read - Read Book for free : 50,000+
Download free book with green book : 100,000+
Wikipedia : 10,000,000+
Cool Reader : 10,000,000+
Free Panda Radio Music : 100,000+
Book store : 1,000,000+
FBReader: Favorite Book Reader : 10,000,000+
English Grammar Complete Handbook : 500,000+
Free Books - Spirit Fanfiction and Stories : 1,000,000+
Google Play Books : 1,000,000,000+
AlReader -any text book reader : 5,000,000+
Offline English Dictionary : 100,000+
Offline: English to Tagalog Dictionary : 500,000+
FamilySearch Tree : 1,000,000+
Cloud of Books : 1,000,000+
Recipes of Prophetic Medicine for free : 500,000+
ReadEra – free ebook reader : 1,000,000+
Anonymous caller detection : 10,000+
Ebook Reader : 5,000,000+
Litnet - E-books : 100,000+
Read books online : 5,000,000+
English to Urdu Dictionary : 500,000+
eBoox: book reader fb2 epub zip : 1,000,000+
English Persian Dictionary : 500,000+
Flybook : 500,000+
All Maths Formulas : 1,000,000+
Ancestry : 5,000,000+
HTC Help : 10,000,000+
E

From the results it seems that there is still a smalll number of extremely popular apps that skew the average. However, it looks like there are only a few very popular apps, so this market still shows potential. Let's try to get some app ideas based on the kind of apps that are somewhere in the middle in terms of popularity (between 1,000,000 and 100,000,000 downloads):

In [34]:
for app in android_final:
    if app[1] == 'BOOKS_AND_REFERENCE' and (app[5] == '1,000,000,000+'
                                            or app[5] == '500,000,000+'
                                            or app[5] == '100,000,000+'):
        print(app[0], ':', app[5])

Google Play Books : 1,000,000,000+
Bible : 100,000,000+
Amazon Kindle : 100,000,000+
Wattpad 📖 Free Books : 100,000,000+
Audiobooks from Audible : 100,000,000+


In [35]:
for app in android_final:
    if app[1] == 'BOOKS_AND_REFERENCE' and (app[5] == '1,000,000+'
                                            or app[5]=='5,000,000+'
                                            or app[5] == '10,000,000+'
                                            or app[5] == '50,000,000+'):
        print(app[0], ':', app[5])

Wikipedia : 10,000,000+
Cool Reader : 10,000,000+
Book store : 1,000,000+
FBReader: Favorite Book Reader : 10,000,000+
Free Books - Spirit Fanfiction and Stories : 1,000,000+
AlReader -any text book reader : 5,000,000+
FamilySearch Tree : 1,000,000+
Cloud of Books : 1,000,000+
ReadEra – free ebook reader : 1,000,000+
Ebook Reader : 5,000,000+
Read books online : 5,000,000+
eBoox: book reader fb2 epub zip : 1,000,000+
All Maths Formulas : 1,000,000+
Ancestry : 5,000,000+
HTC Help : 10,000,000+
Moon+ Reader : 10,000,000+
English-Myanmar Dictionary : 1,000,000+
Golden Dictionary (EN-AR) : 1,000,000+
All Language Translator Free : 1,000,000+
Aldiko Book Reader : 10,000,000+
Dictionary - WordWeb : 5,000,000+
50000 Free eBooks & Free AudioBooks : 5,000,000+
Al-Quran (Free) : 10,000,000+
Al Quran Indonesia : 10,000,000+
Al'Quran Bahasa Indonesia : 10,000,000+
Al Quran Al karim : 1,000,000+
Al Quran : EAlim - Translations & MP3 Offline : 5,000,000+
Koran Read &MP3 30 Juz Offline : 1,000,000+
H

This niche seems to be dominated by software for processing and reading ebooks, as well as various collections of libraries and dictionaries, so it's probably not a good idea to build similar apps since there'll be some significant competition.

We also notice there are quite a few apps built around the book Quran, which suggests that building an app around a popular book can be profitable. It seems that taking a popular book (perhaps a more recent book) and turning it into an app could be profitable for both the Google Play and the App Store markets.

However, it looks like the market is already full of libraries, so we need to add some special features besides the raw version of the book. This might include daily quotes from the book, an audio version of the book, quizzes on the book, a forum where people can discuss the book, etc.

The above analysis was taken from [dataquest solution notebook](https://github.com/dataquestio/solutions/blob/master/Mission350Solutions.ipynb)

## Conclusions
In this project, we analyzed data about the App Store and Google Play mobile apps with the goal of recommending an app profile that can be profitable for both markets.

We concluded that taking a popular book (perhaps a more recent book) and turning it into an app could be profitable for both the Google Play and the App Store markets. The markets are already full of libraries, so we need to add some special features besides the raw version of the book. This might include daily quotes from the book, an audio version of the book, quizzes on the book, a forum where people can discuss the book, etc.