I am interested in quantitative research at the intersection of mathematical modeling, statistical learning, optimization, and reliable computation.
My projects range from financial econometrics and probabilistic race modeling to counterfactual retrieval, credit-risk modeling, topological analysis of high-dimensional data, and applied matching systems. I enjoy turning quantitative questions into reproducible experiments and using computational evidence to test where an idea works, fails, or needs stronger assumptions.
| Project | Focus | Methods |
|---|---|---|
| Home Credit Default Risk | Ranking repayment difficulty from application and historical credit data | LightGBM, relational feature engineering, 5-fold OOF validation, Optuna, SHAP |
| Jockey Predictor | Estimating Hong Kong race outcomes and comparing model probabilities with market odds | NGBoost, probabilistic regression, Monte Carlo simulation, Selenium, GitHub Actions |
| S&P 500 Pairs Trading Research | Testing whether cointegrated equity pairs support mean-reversion after costs | Engle-Granger, ADF, FDR control, rolling OLS, chronological backtesting |
| Approximate Counterfactual Preservation | Reducing named features while preserving retrieval-based counterfactuals | Exact search, beam search, ablations, non-monotone optimization |
| Topology of the Million Song Dataset | Comparing acoustic and behavioral structure in music data | Persistent homology, UMAP, HDBSCAN, graph analysis |
| Imperial Marriage Pact | Preference-based matching from survey data | Vector similarity, constraints, ranking, batch data pipelines |
- Statistical learning and high-dimensional data
- Probabilistic modeling and Monte Carlo methods
- Time-series analysis and financial econometrics
- Counterfactual reasoning and model interpretability
- Probability, numerical methods, and algorithmic reasoning
Python · NumPy · pandas · SciPy · statsmodels · scikit-learn · LightGBM · NGBoost · Optuna · SHAP · scientific computing · experiment design · data visualization