## Data analysis to understand which kind of mobile apps are more attractive to users (US)
---

## Análise de dados para entender quais tipos de  aplicativos móveis são mais atraentes para os usuários (PT-BR)

(US)
* What the project is about

The project is about to help developers to understand the user's behavior and the app's preference on the Google Play Store and App Store.

* What your goal is in this project

The goal is to help developers understand what kinds of apps are likely to attract more users on Google Play and the App Store.



(PT-BR)
* Sobre o que é o projeto:

O projeto é sobre auxiliar os desenvolvedores a compreenderem o comportamento dos usuários e a prefência de aplicativos na Google Play Store e App Store.

* Objetivo do projeto:

O objetivo é demonstrar através da análise exploratória constatar quais são os aplicativos mais utilizados e com isso focar em possíveis planos de rentabilidade.


In [None]:
def open_dataset(file_name):
    from csv import reader
    opened_file = open(file_name, encoding='utf8')
    read_file = reader(opened_file)
    data = list(read_file)
    return data

---
(US)
We'll start by opening and exploring these two data sets. To make it easier for you to explore them, we created a function named explore_data() you can use repeatedly to print rows in a more readable way.

In [None]:
def explore_data(dataset, start, end, rows_and_columns=False):
    dataset_slice = dataset[start:end]    
    for row in dataset_slice:
        print(row)
        print('\n') # adds a new (empty) line after each row
    if rows_and_columns:
        print('Number of rows:', len(dataset))
        print('Number of columns:', len(dataset[0]))

In [None]:
google_data = open_dataset('googleplaystore.csv') 
apple_data = open_dataset('AppleStore.csv') 


(US) Below, exploring both data sets using the "explore_data()" function and analising the headers, the first four rows and numbers of
rows and collumns of data sets.

(BR) Abaixo, explorando ambos data sets usando a função "explore_data()" e analisando os cabeçalhos, as primeiras quatro linhas e o número de linhas e colunas dos data sets.

In [32]:
print('***********************************************************************************************************************')

print('Google Play Store')
print('Documentation: ''https://www.kaggle.com/lava18/google-play-store-apps/home')
print('-----------------------------------------------------------------------------------------------------------------------')
print(explore_data(google_data,0,4,True))

print('***********************************************************************************************************************')

print('App Store')
print('Documentation: ''https://www.kaggle.com/ramamet4/app-store-apple-data-set-10k-apps/home')
print('-----------------------------------------------------------------------------------------------------------------------')
print(explore_data(apple_data,0,4,True))

print('***********************************************************************************************************************')

***********************************************************************************************************************
Google Play Store
Documentation: https://www.kaggle.com/lava18/google-play-store-apps/home
-----------------------------------------------------------------------------------------------------------------------
['App', 'Category', 'Rating', 'Reviews', 'Size', 'Installs', 'Type', 'Price', 'Content Rating', 'Genres', 'Last Updated', 'Current Ver', 'Android Ver']


['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['Coloring book moana', 'ART_AND_DESIGN', '3.9', '967', '14M', '500,000+', 'Free', '0', 'Everyone', 'Art & Design;Pretend Play', 'January 15, 2018', '2.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design',

(BR) Abaixo foi verificado que a linha "10473" do data set contém dado missing, no qual comparamos a linha errada com uma linha correta conforme o cabeçalho onde é notado a falta de dado na coluna "Category".

In [None]:
print(google_data[0]) # header
print('\n')
print(google_data[10473]) # wrong data
print('\n')
print(google_data[1]) # correct data

In [None]:
print(len(google_data))
del google_data[10473] #deleting wrong row 
print(len(google_data))

(BR) O Data set Google Play Store possui alguns dados duplicados devido somente a atualização do número de "Reviews" dos apps como "Instagram" por exemplo, conforme "print" abaixo:

In [None]:
print(google_data[0]) # header google_data
print('\n')
for app in google_data:
    name = app[0]
    if name == 'Instagram':
        print(app)

In [None]:
duplicate_apps = []
unique_apps = []

for app in google_data:
    name = app[0]
    if name in unique_apps:
        duplicate_apps.append(name)
    else:
        unique_apps.append(name)
        
print('Number of duplicate apps: ', len(duplicate_apps))
print('\n')
print('Examples of duplicate apps: ', duplicate_apps[:15])
print('\n')
print('Number of unique apps: ', len(unique_apps[1:]))

(US) I could use this information to build a criterion for removing the duplicates. The higher the number of reviews, the more recent the data should be. Rather than removing duplicates randomly, I'll only keep the row with the highest number of reviews and remove the other entries for any given app.

(PT-BR) Eu poderia usar essas informações para criar um critério para remover as duplicatas. Quanto maior o número de comentários, mais recentes devem ser os dados. Em vez de remover duplicatas aleatoriamente, só manterei a linha com o maior número de reviews e removerei as outras entradas de qualquer aplicativo.

In [9]:
print('Expected length: ', len(google_data[1:]) - 1181)

Expected length:  9659


(US) Below I created a dictionary where each key is a unique app name and the corresponding dictionary value is the highest number of reviews of that app.

Start by creating an empty dictionary named reviews_max and Loop through the Google Play data set.

In [10]:
reviews_max = {}

for app in google_data[1:]:
    name = app[0]
    n_reviews = float(app[3])
    if name in reviews_max and reviews_max[name] < n_reviews:
        reviews_max[name] = n_reviews
    elif name not in reviews_max:
        reviews_max[name] = n_reviews
        
print(len(reviews_max))
    
    

9659


(US) Below I used the dictionary that I have created above to remove the duplicate rows.

Start by creating two empty lists: android_clean (which will store the new cleaned data set) and already_added (which will just store app names).

In [11]:
android_clean = []
already_added = []


for app in google_data[1:]:
    name = app[0]
    n_reviews = float(app[3])
    if n_reviews == reviews_max[name] and name not in already_added:
        android_clean.append(app)
        already_added.append(name)

print(explore_data(android_clean,1,5,True))


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


['Pixel Draw - Number Art Coloring Book', 'ART_AND_DESIGN', '4.3', '967', '2.8M', '100,000+', 'Free', '0', 'Everyone', 'Art & Design;Creativity', 'June 20, 2018', '1.1', '4.4 and up']


['Paper flowers instructions', 'ART_AND_DESIGN', '4.4', '167', '5.6M', '50,000+', 'Free', '0', 'Everyone', 'Art & Design', 'March 26, 2017', '1.0', '2.3 and up']


Number of rows: 9659
Number of columns: 13
None


In the previous step, I managed to remove the duplicate app entries in the Google Play data set. Remember that the language I use for the apps is English, and I'd like to analyze only the apps that are directed toward an English-speaking audience. However, if I explore the data long enough, I'll find that both data sets have apps whose name suggests that they are not directed toward an English-speaking audience.

The numbers corresponding to the characters we commonly use in an English text are all in the range 0 to 127, according to the ASCII (American Standard Code for Information Interchange) system. Based on this number range, I can build a function that detects whether a character belongs to the set of common English characters or not. If the number is equal to or less than 127, then the character belongs to the set of common English characters, otherwise it doesn't.

So if an app name contains a character that is greater than 127, then it probably means that the app has a non-English name.

In [12]:
def is_english(string):
    non_ascii = 0
    for character in string:
        if ord(character) > 127:
            non_ascii += 1
        
    if non_ascii > 3:
        return False
    else:
        return True

print(is_english('Docs To Go™ Free Office Suite'))
print(is_english('Instachat 😜'))
print(is_english('爱奇艺PPS -《欢乐颂2》电视剧热播'))

True
True
False


Above, I wrote a function that detects non-English app names, but we saw that the function couldn't identify correctly certain English app names like 'Docs To Go™ Free Office Suite' and 'Instachat 😜'. This is because emojis and some characters like ™ fall outside the ASCII range and have corresponding numbers that are over 127.
If we're going to use the function I've created, I'll lose useful data since many English apps will be incorrectly labeled as non-English. To minimize the impact of data loss, I'll only remove an app if its name has more than three characters with corresponding numbers falling outside the ASCII range. This means all English apps with up to three emoji or other special characters will still be labeled as English. The filter function is still not perfect, but it should be fairly effective.

In [13]:
android_english = []
ios_english = []

for app in android_clean:
    name = app[0]
    if is_english(name):
        android_english.append(app)

for app in apple_data:
    name = app[1]
    if is_english(name):
        ios_english.append(app)

print(explore_data(android_english,0,3,True))
print('\n')
print(explore_data(ios_english,0,3,True))         

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 9614
Number of columns: 13
None


['id', 'track_name', 'size_bytes', 'currency', 'price', 'rating_count_tot', 'rating_count_ver', 'user_rating', 'user_rating_ver', 'ver', 'cont_rating', 'prime_genre', 'sup_devices.num', 'ipadSc_urls.num', 'lang.num', 'vpp_lic']


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'In

So far in the data cleaning process:

    Removed inaccurate data
    Removed duplicate app entries
    Removed non-English apps

As I mentioned in the introduction, I only analyze apps that are free to download and install, and the main source of revenue consists of in-app ads. Our data sets contain both free and non-free apps, and I'll need to isolate only the free apps for analysis.

Isolating the free apps will be the last step in the data cleaning process. In the next cells, I'll going to start analyzing the data.

In [14]:
android_final = []
ios_final = []

for app in android_english:
    price = app[7]
    if price == '0':
        android_final.append(app)

for app in ios_english:
    price = app[4]
    if price == '0.0':
        ios_final.append(app)

print(explore_data(android_final,0,3,True))
print('\n')
print(explore_data(ios_final,0,3,True))     

['Photo Editor & Candy Camera & Grid & ScrapBook', 'ART_AND_DESIGN', '4.1', '159', '19M', '10,000+', 'Free', '0', 'Everyone', 'Art & Design', 'January 7, 2018', '1.0.0', '4.0.3 and up']


['U Launcher Lite – FREE Live Cool Themes, Hide Apps', 'ART_AND_DESIGN', '4.7', '87510', '8.7M', '5,000,000+', 'Free', '0', 'Everyone', 'Art & Design', 'August 1, 2018', '1.2.4', '4.0.3 and up']


['Sketch - Draw & Paint', 'ART_AND_DESIGN', '4.5', '215644', '25M', '50,000,000+', 'Free', '0', 'Teen', 'Art & Design', 'June 8, 2018', 'Varies with device', '4.2 and up']


Number of rows: 8864
Number of columns: 13
None


['284882215', 'Facebook', '389879808', 'USD', '0.0', '2974676', '212', '3.5', '3.5', '95.0', '4+', 'Social Networking', '37', '1', '29', '1']


['389801252', 'Instagram', '113954816', 'USD', '0.0', '2161558', '1289', '4.5', '4.0', '10.23', '12+', 'Photo & Video', '37', '0', '29', '1']


['529479190', 'Clash of Clans', '116476928', 'USD', '0.0', '2130805', '579', '4.5', '4.5', '9.24.12', '

So far, I spent a good amount of time on cleaning data, and:

    Removed inaccurate data
    Removed duplicate app entries
    Removed non-English apps
    Isolated the free apps

As I mentioned in the introduction,the aim is to determine the kinds of apps that are likely to attract more users because the revenue is highly influenced by the number of people using their apps.

Let's begin the analysis by getting a sense of what are the most common genres for each market. For this, I'll need to build frequency tables for a few columns in our data sets.

In [53]:
def freq_table(dataset, index):
    table = {}
    total = 0
    
    for row in dataset:
        total += 1
        value = row[index]
        if value in table:
            table[value] += 1
        else:
            table[value] = 1
    
    table_percentages = {}
    for key in table:
        percentage = (table[key] / total) * 100
        table_percentages[key] = percentage 
    
    return table_percentages


def display_table(dataset, index):
    table = freq_table(dataset, index)
    table_display = []
    for key in table:
        key_val_as_tuple = (table[key], key)
        table_display.append(key_val_as_tuple)
        
    table_sorted = sorted(table_display, reverse = True)
    for entry in table_sorted:
        print(entry[1], ':', entry[0])


In [54]:
print(display_table(android_final,1))

FAMILY : 18.907942238267147
GAME : 9.724729241877256
TOOLS : 8.461191335740072
BUSINESS : 4.591606498194946
LIFESTYLE : 3.9034296028880866
PRODUCTIVITY : 3.892148014440433
FINANCE : 3.7003610108303246
MEDICAL : 3.531137184115524
SPORTS : 3.395758122743682
PERSONALIZATION : 3.3167870036101084
COMMUNICATION : 3.2378158844765346
HEALTH_AND_FITNESS : 3.0798736462093865
PHOTOGRAPHY : 2.944494584837545
NEWS_AND_MAGAZINES : 2.7978339350180503
SOCIAL : 2.6624548736462095
TRAVEL_AND_LOCAL : 2.33528880866426
SHOPPING : 2.2450361010830324
BOOKS_AND_REFERENCE : 2.1435018050541514
DATING : 1.861462093862816
VIDEO_PLAYERS : 1.7937725631768955
MAPS_AND_NAVIGATION : 1.3989169675090252
FOOD_AND_DRINK : 1.2409747292418771
EDUCATION : 1.1620036101083033
ENTERTAINMENT : 0.9589350180505415
LIBRARIES_AND_DEMO : 0.9363718411552346
AUTO_AND_VEHICLES : 0.9250902527075812
HOUSE_AND_HOME : 0.8235559566787004
WEATHER : 0.8009927797833934
EVENTS : 0.7107400722021661
PARENTING : 0.6543321299638989
ART_AND_DESIGN : 

In [58]:
print(display_table(android_final,-4))

Tools : 8.449909747292418
Entertainment : 6.069494584837545
Education : 5.347472924187725
Business : 4.591606498194946
Productivity : 3.892148014440433
Lifestyle : 3.892148014440433
Finance : 3.7003610108303246
Medical : 3.531137184115524
Sports : 3.463447653429603
Personalization : 3.3167870036101084
Communication : 3.2378158844765346
Action : 3.1024368231046933
Health & Fitness : 3.0798736462093865
Photography : 2.944494584837545
News & Magazines : 2.7978339350180503
Social : 2.6624548736462095
Travel & Local : 2.3240072202166067
Shopping : 2.2450361010830324
Books & Reference : 2.1435018050541514
Simulation : 2.0419675090252705
Dating : 1.861462093862816
Arcade : 1.8501805054151623
Video Players & Editors : 1.7712093862815883
Casual : 1.7599277978339352
Maps & Navigation : 1.3989169675090252
Food & Drink : 1.2409747292418771
Puzzle : 1.128158844765343
Racing : 0.9927797833935018
Role Playing : 0.9363718411552346
Libraries & Demo : 0.9363718411552346
Auto & Vehicles : 0.9250902527075

In [59]:
print(display_table(ios_final,-5))

Games : 58.16263190564867
Entertainment : 7.883302296710118
Photo & Video : 4.9658597144630665
Education : 3.662321539416512
Social Networking : 3.2898820608317814
Shopping : 2.60707635009311
Utilities : 2.5139664804469275
Sports : 2.1415270018621975
Music : 2.0484171322160147
Health & Fitness : 2.0173805090006205
Productivity : 1.7380509000620732
Lifestyle : 1.5828677839851024
News : 1.3345747982619491
Travel : 1.2414649286157666
Finance : 1.1173184357541899
Weather : 0.8690254500310366
Food & Drink : 0.8069522036002483
Reference : 0.5586592178770949
Business : 0.5276225946617008
Book : 0.4345127250155183
Navigation : 0.186219739292365
Medical : 0.186219739292365
Catalogs : 0.12414649286157665
None
