- Đặng Châu Anh - 22127008
- Nguyễn Kim Anh - 22127014
- Đỗ Minh Huy - 22127147
- Trần Dịu Huyền - 22127170
- Analyze the top trending spotify music data.
- Presenting the findings in a way that provides insights into local music preferences.
- The Vietnamese music industry has significant growth in 2024, with a surge in both quantity and quality of music products. This project uses Spotify data to better understand the changing tastes of Vietnamese listeners. By identifying popular songs, albums, and new trends, the project will provide useful insights to help industry professionals, artists, and fans stay updated on current music preferences.
We planned and discussed the problems together, then record it at Trello.
Both raw and cleaned data can be found in the data and cleaned_data folders, respectively in github repository.
https://github.com/pililover/IntroDS-Group-HAHA.git
During the implementation of this project, we encountered many obstacles:
- 22127008: Providing insights for questions is also a difficulty. I had to spend a lot of time thinking about how to ask meaningful questions and how to answer them. Furthermore, choosing correct chart types to visualize the data is also a difficulty.
- 22127014: Considering how to ask meaningful questions also took a lot of time to implement. Request for more features is also a difficulty cause one website can only provide a limited number of features and we need to scrape data from multiple sources to get more features. So how to combine them is also a difficulty.
- 22127147: I had difficulty in finding the data source and the data collection process. I had to learn how to use selenium to scrape data from the website cause some websites are dynamic.
- 22127170: I had some trouble with preprocessing the data and visualizing the data. I had to consider carefully which method to clean the data in particular. Providing insights for questions to make it meaningful is also a challenge.
- Through this project, we learned how to organize and divide work appropriately. Through the process of completing the project, we effectively analyzed and extracted insights from raw datasets. Additionally, we gained valuable experience in managing our time efficiently for each phase of the project.
- 22127008: After completing this project, I learned how to choose a model and use metrics to predict the required results. In addition, understanding and visualizing data into understandable insights is also an interesting thing that I learned during the process of doing this project.
- 22127014: After this project, parsing data from the website is no longer a difficulty for me. I also learned how to frame the problems directly and provide questions that are meaningful and insightful which is very important in data analysis.
- 22127147: After completing this project, I have learned a lot about data analysis, but the most impressive thing is that I have learned how to evaluate the regression model and how to choose the right model for the data.
- 22127170: Thanks to this project, I learned how to preprocess and clean data effectively. I also gained experience in choosing the right visualization techniques to present data insights meaningfully. This project helped me understand the importance of data preprocessing and visualization in the data analysis process.
- If our group had more time, we would have liked to delve deeper into the dataset to uncover more insights and scrape more data from different sources to enrich the dataset cause we think it would be more interesting and informative. Furthermore, we would like to use various machine learning algorithms to predict the number of cases in the future.