-
Notifications
You must be signed in to change notification settings - Fork 0
06. Evaluation
Develop customer segmentation models based on purchasing behavior.
The human knowledge-based model aligns well with the business objective of creating interpretable and actionable customer segments. By grouping customers according to clear behaviors and spending habits, the model provides a straightforward way for the business to identify and cater to different customer needs. Each segment, such as Discount Seekers or Loyal High-Spenders, reflects a specific behavioral pattern that allows for tailored marketing strategies, targeted promotions, and customized customer retention programs. This segmentation framework enables the business to efficiently allocate resources and design initiatives that address key customer motivations, contributing to overall customer satisfaction and loyalty.
While this human knowledge-based model offers interpretability and actionable insights, it has some limitations:
-
Coarse Segmentation: This approach relies heavily on pre-defined rules based on current understanding of customer behavior. The rule-based segmentation might oversimplify complex customer behavior patterns. For example, high-spending customers may have distinct underlying motivations that are not captured by predefined segments.
-
Lack of Scability: As customer data grows, maintaining and updating rule-based segments could become time-consuming. Additionally, human knowledge-based approach may miss emerging customer trends that could require new segments or updated definitions.
-
Potential Overlap Between Segments: Some customers may exhibit behaviors that align with multiple segments (e.g., a tech-savvy user who is also a discount seeker), which may lead to overlaps and reduce the precision of targeted strategies.
To address these limitations, future improvements could include:
-
Hybrid Approach with Machine Learning: Introduce a hybrid model that combines human knowledge with machine learning techniques, such as clustering with PCA or HDBSCAN, to validate and enhance existing segments. This could allow the model to discover emerging patterns or segments based on data-driven insights while retaining business interpretability.
-
Periodic Re-evaluation of Segments: As customer behavior and market conditions change, regularly re-evaluating and updating the segments will ensure that the model remains relevant and responsive to current trends.
-
Hierarchical Segmentation: Develop a hierarchical model with broad segments (e.g., high-level behavioral categories) that branch into sub-segments for finer distinctions. This would capture both broad patterns and specific behavioral nuances within each customer group.
- The model (ROI) effectively identify the most profitable marketing channels (e.g., email and mobile push) and provide actionable insights into customer behavior and engagement patterns. This supports the objective of maximizing revenue and improving customer satisfaction by tailoring marketing strategies to align with customer preferences.
- The calculation of ROI for different channels helps in understanding cost-effectiveness, aligning with the goal of reducing operational costs and enhancing profitability.
- Insights into peak engagement times and effective features (e.g., "discount," "saleout") directly inform marketing strategies, supporting customer retention and lifetime value objectives.
- For e-commerce business owners and executives, this approach provides understanding of revenue drivers and sales trends, aiding strategic decisions.
- Marketing teams benefit from insights into effective channels and campaign features, enhancing the ROI of marketing efforts and improving customer acquisition and retention strategies.
- The approach effectively uses data integration and feature engineering to extract insights on customer behavior and campaign performance, directly aligning with the business objective of understanding customer purchasing behavior.
- By analyzing conversion rates and purchase volumes across campaigns, the analysis provides actionable insights into the effectiveness of various marketing strategies
- Data Limitations: Due to the large size of the original dataset, a sampling method was employed, resulting in an incomplete representation. Consequently, only 3 out of 4 channels were analyzed, as data from the SMS channel was not included in the sample, limiting the comprehensiveness of the analysis. Additionally, the multichannel strategy analysis is based on limited data points, restricting the ability to draw robust conclusions about its performance and potential synergies with other channels.
- Feature Interaction: Our current approach may not fully capture complex interactions between features and external market factors, which could limit the depth of insights regarding customer behavior and engagement.
- Temporal and Market Dynamics: Our current approach might not account for seasonal variations or broader market trends that could affect consumer behavior and pricing strategies.
- Cost Estimation: The reliance on standard pricing models may not accurately reflect the unique cost structures and operational expenses of the business, potentially affecting the precision of ROI calculations.
- Potential Class Imbalance: The high accuracy of 99.97% in the Random Forest model suggests potential class imbalance, meaning it may predominantly predict the majority class rather than providing nuanced insights into different purchasing behaviors.
- Temporal Scope: The analysis focuses on specific time features (hour, month, day of the week) which might not account for longer-term trends or variations across different periods, such as holidays or economic shifts.
- Lack of Customer Segmentation: While targeting strategies were noted, our current approach does not explicitly incorporate detailed customer segmentation, which could refine marketing strategies further and enhance personalized customer experiences.
- Lack of Comprehensive Metric Evaluation: The evaluation mainly focuses on accuracy; however, additional metrics like precision, recall, and F1-score would provide a more comprehensive understanding of model performance, especially in situations with class imbalance.
- Expand Data Collection: Increase the volume and diversity of data collected, particularly for multichannel and sms channels, for random sample, to validate current findings and explore new insights into channel performance and interactions.
- Adopt Advanced Analytical Techniques: Implement more sophisticated models that can account for interaction effects between features and adapt to dynamic market conditions. Techniques such as machine learning could be employed to better capture these complexities.
- Incorporate Temporal and Market Factors: Enhance models by integrating seasonal trends and market dynamics to provide more accurate and context-sensitive insights into customer behavior and pricing strategies.
- Refine Cost Models: Develop customized cost models that reflect the specific operational expenses of the business, leading to more accurate ROI evaluations and better resource allocation decisions.
- Address Class Imbalance: Introduce techniques such as resampling (oversampling the minority class, undersampling the majority class), or using algorithms like SMOTE (Synthetic Minority Over-sampling Technique) to balance the data classes for more accurate model predictions.
- Expand Feature Set: Incorporate additional features that might affect purchase behavior, such as customer demographics, past purchase history, or engagement with previous campaigns, to refine models and provide richer insights.
- Enhanced Model Assessment: Use additional evaluation metrics such as precision, recall, and the F1-score to better assess the model's ability to balance sensitivity and specificity.
- Integrate Customer Segmentation: Develop detailed customer segmentation analyses to personalize marketing strategies further and improve the targeting accuracy of campaigns.
- Longer-Term Analysis: Expand the temporal scope of the analysis to include trends over multiple years or seasons, providing context for how broader market dynamics might affect engagement and sales patterns.
The models, Seasonal ARIMA, linear programming and stock level tracking, support cost-effective inventory management, allowing for proactive ordering and inventory adjustments to maintain service levels. These insights help managers make informed decisions on order timing and quantities, minimizing unnecessary holding costs while meeting demand.
-
SARIMA’s reliance on historical data may limit its adaptability to sudden demand changes, such as unexpected surges or declines in demand. For instance, it is unable to predict economic recessions that may result in lower consumption.
-
The linear programming model assumes fixed costs for holding and restocking, which oversimplifies real-world scenarios where costs fluctuate based on supply chain dynamics and supplier pricing changes. It also fails to model that products' value decrease as it is kept in inventory.
-
The current models also have limited flexibility for product variety as it currently only evaluates 3 categories. It may struggle to handle a broad range of products with different demand patterns and cost structures, requiring customization for each product type.
-
We can establish a monitoring and retraining schedule to adapt the model to changing demand patterns, seasonality, supplier performance and global market. Regular updates ensure the model remains aligned with business objectives and adapts to environmental shifts.
-
We could introduce a dynamic cost function that accounts for fluctuations in holding and restocking costs based on supplier contracts, seasonal pricing, or discount opportunities. This would more accurately reflect and enhance inventory cost considerations.
-
We can develop a segmented approach for different product categories, such as high-demand vs. low-demand items or fast-moving vs. slow-moving products. Applying separate demand forecasting and optimization techniques for each segment would improve overall accuracy and cost-effectiveness across product types and reduce the issue of complexity.
With revenue maximising as our business objective, the dynamic pricing model can aid our ecommerce business in deciding the best price for a product. While our price predictions are not 100% accurate, it provides a reasonable range of prices which are close to optimal for a particular product category.
-
Model is trained on all product categories: This reduces accuracy for individual product categories, as different categories is likely to have different prices even if they have the same demand and supply characteristics. For example, clothing will very likely cost less than electronics even if they have similar demand and supply characteristics.
-
Lack of Forecast for Real-Time Competitor Data: If we want to forecast and decide for our prices in the future, it is highly likely that we will not have access to our competitor's prices. Given the importance of competitor's prices when deciding on our price, our model will not be able to reach an optimal pricing prediction.
-
Separate Models for Each Product Category: Separate models allow for more tailored pricing strategies for each category, making it easier to optmise revenue. This is because different product categories may have different price distributions despite having similar demand and supply characteristics. Products in the same category are more likely to have similar price distributions if they follow similar demand and supply characteristics.
-
Additional Model for Prediction of Competitor's Prices: In the event where competitor's prices are not known, we can develop a separate model solely to predict our competition's prices. This can be done by looking at the historical price trends of our competition. This prediction can then be fed into our pricing model in order for our business to forecast future prices.
Our logistic regression model demonstrated strong performance, achieving an accuracy of 97.51% and a high F1-score for delayed orders, aligning well with our primary business objectives of predicting delays accurately and identifying supply chain bottlenecks. By achieving nearly perfect recall for delayed orders, the model effectively flags potential delays, enabling preemptive actions to mitigate these issues.
From a business perspective, this high level of predictive accuracy supports timely interventions, such as prioritizing high-risk orders and reallocating resources to streamline fulfillment. With these insights, stakeholders can make informed decisions that help reduce delivery times and enhance the overall efficiency of the supply chain.
While the logistic regression model performed well, several limitations are noted:
-
Feature Engineering Constraints: The model relies heavily on the quality of manually engineered features, such as scheduled days for shipping and late delivery risk. These features may not capture all complex interactions that impact order delays, potentially limiting model performance.
-
Assumption of Linearity: Logistic regression assumes a linear relationship between features and the log-odds of delay, which may oversimplify real-world dependencies within the supply chain. Non-linear models, such as decision trees or ensemble methods, might capture these relationships more effectively.
-
Imbalanced Data Handling: Although the model handles class imbalance reasonably well, further techniques such as Synthetic Minority Over-sampling Technique (SMOTE) or additional sampling methods might enhance its performance on underrepresented classes, especially if the imbalance were to increase over time.
-
Limited Scalability: Logistic regression’s interpretability comes with limited scalability for complex, high-dimensional data. As the dataset grows in complexity, this model may struggle to capture nuanced patterns compared to more sophisticated methods.
To address the limitations and further improve model performance, we suggest the following enhancements:
-
Advanced Feature Engineering: Introduce features that capture non-linear relationships or complex interactions within the supply chain data, such as interaction terms or time-based trends, to improve predictive power.
-
Alternative Models: Experiment with more advanced models, such as Random Forest or XGBoost, which can handle non-linear relationships and might capture subtle dependencies within the data. These models could potentially yield higher performance, albeit with reduced interpretability.
-
Enhanced Imbalance Techniques: Implement sampling techniques, such as SMOTE or under-sampling of the majority class, to ensure the model remains effective as the dataset grows or the class imbalance shifts. This would also allow a broader exploration of data points in underrepresented classes, improving the model's recall for delays.
-
Continuous Monitoring and Model Updating: Regularly monitor the model’s performance and retrain it as needed to adapt to changes in the supply chain environment, such as seasonal trends or supplier performance fluctuations, ensuring the model remains aligned with evolving business objectives.
-
Incorporation of External Data: To improve accuracy, consider incorporating external data sources, such as weather conditions or real-time traffic data, which may affect delivery times and influence the likelihood of delays.
1. What is the potential of using natural language processing to analyze customer reviews and feedback?
-
Extraction Of Common Feedback: Common feedback is not presented as a column in our Customer Reviews dataset. We extracted common feedback from the reviews themselves, hoping to find words that revolve around feedback, and the common issues gathered. The output may not be accurate, but we can still roll with it.
-
Limited scalability: We did not run NER with SpaCy to identify common issues across all reviews as this will result in a long list of words. For simplicity, we ran NER only on negative and neutral reviews. We may have missed out on issues found in positive reviews, and the common issues we extracted are only from a subset of customer reviews.
Alternative Model Consideration: Usage of a more accurate model compared to NER using SpaCy that seeks to extract common issues from all customer reviews without getting a long list of words. The same may be said to find a more accurate model than VADER for customer review sentiment analysis, if there is.
2. What are the key insights about product features, quality, and customer satisfaction from customer review analysis? (Use an LLM)
As we have written in our report, our system is not totally foolproof. We explained that our implementation is as such because of the nature of the reviews. There are definitely better systems that can be implemented for aspect-based sentiment analysis (ABSA). However, our system is still usable to provide insights to help us improve our relationship with customers by identifying the biggest issues and working on them.
-
Regex Searching Incapabilities: Our LLM system actively searches for entities within the text to check if the review has mentioned that aspect, using regex searching based off a list of keywords. This is not highly accurate, especially if the reviews are better structured and more comprehendible.
-
Customer Satisfaction As An Overall Sentiment: For our implementation, we assume customer satisfaction to be captured by the overall sentiment of the review. This is a bold assumption. In our case, due to the nature of the reviews, we have adjusted our implementation as such to increase accuracy. This is not a viable solution when we analyze reviews that can be separated clearly into portions of text talking about each of our business question’s key areas.
Future improvements include:
-
Building Our Own Model: Developing our own LLM-based system to implement ASBA on customer reviews. These pre-trained models from the Internet may not be accurate when they are run on a data set with poorly structured customer reviews, which was what we had. Developing an LLM and training it on these reviews is an improvement from these transformer models that were pre-trained with better quality reviews. Additionally, we need to have a test data set that classifies each review properly for training.
-
Customer Review Improvement: Employing methods to collect reviews that are short and simple, addressing all aspects. The nature of reviews is very important, as good quality reviews are easier to comprehend, and the output is more accurate.
-
Alternative Ways To Tidy Reviews: Explore more complex ways of tidying reviews. Given the
review_contentcolumn, we can chop these full reviews into different parts; removing irrelevant parts and retaining portions of the text vital for each aspect. After retaining these portions, create separate columns “Product Features”, “Quality” and “Customer Satisfaction”, and iteratively add the portions of text under their respective columns for each review. -
Alternative Models Consideration: Changing the models and techniques in place. For instance, the model and regex search can make way for a more complex ASBA model that performs all these tasks concurrently.