π Master of Information & Data Science from UC Berkeley School of Information (MIDS '26)
π Dual B.S. in Psychology & Criminal Justice from University of Central Florida
πΌ Data Scientist at KPMG US β bridging risk analytics, compliance data, and machine learning
π 5+ years turning complex data into actionable insights across audit, compliance, fintech, and public health
π± Building at the intersection of AI/ML, causal inference, and human-centered design.
Languages: Python, SQL, R
ML & AI: Scikit-learn, TensorFlow, PyTorch, XGBoost, Random Forest, NLP (spaCy, LangChain), Deep Learning, Computer Vision, Causal Inference
Data & Viz: Pandas, NumPy, Altair, Matplotlib, Seaborn, Plotly, Tableau, Power BI
Backend & Databases: Flask, PostgreSQL, DuckDB, MongoDB, BigQuery, REST APIs, Chrome Extension APIs
Cloud & Tools: AWS, Google Cloud, Git, Docker, Jupyter, Neo4j
Statistics: Hypothesis Testing, A/B Testing, Propensity Score Matching, Synthetic Control, Econometric Modeling, Regression Analysis
- π¬ Build end-to-end ML pipelines β from EDA and feature engineering to model deployment and monitoring
- π Design interactive dashboards (Flask, Altair, Tableau) that turn complex data into narratives
- π§ Apply causal inference, NLP, and deep learning to real-world problems in fintech, public health, and policy
- π‘οΈ Bring compliance rigor to AI β SOC 1/2, SOX, risk assessment, and controls testing inform how I think about bias, fairness, and robustness in ML systems
I speak English, Haitian Creole, French, and Spanish β because great insights should travel across languages, not just dashboards.
UC Berkeley MIDS Capstone | Jan β Apr 2026
A browser extension that detects deceptive UI patterns (dark patterns) on websites in real time using a two-tier ML detection system. A lightweight client-side gatekeeper (regex rules + ONNX/WASM inference) filters candidate elements before sending them to cloud-hosted XGBoost models trained per pattern type. Includes a nightly online retraining pipeline with Golden Set evaluation, versioning, and rollback.
Tech: Python, Flask, NLP, Chrome Extension APIs, ONNX, XGBoost, DOM Parsing
Highlights: Real-time detection β’ Online retraining β’ Model versioning & rollback β’ Performance dashboard
π GitHub
UC Berkeley MIDS Capstone | 2025
An ML pipeline to predict California wildfire risk using satellite imagery, weather data, and historical fire records. Engineered geospatial features and trained XGBoost, Random Forest, and Neural Network models to forecast fire probability at the county level. Evaluated model performance with precision-recall analysis for imbalanced disaster prediction.
Tech: Python, XGBoost, Random Forest, Neural Networks, Geopandas, Scikit-learn, Geospatial Analysis
Highlights: Geospatial ML β’ Imbalanced classification β’ Emergency response planning
π GitHub | Medium Article
UC Berkeley MIDS | 2025
An interactive Flask website visualizing mental health and substance use trends across the United States using NSDUH public health data. Features Altair-powered interactive charts, geospatial analysis, and narrative storytelling that transforms complex health data into accessible insights.
Tech: Python, Flask, Altair, DuckDB, Pandas, Docker
Highlights: Live deployed site β’ Public health focus β’ Interactive visualizations
π Live Site | Medium
UC Berkeley DATASCI 281 | 2024
A progressive feature engineering pipeline for art genre classification, evolving from handcrafted visual descriptors (HOG, LAB color, LBP, edge detection) to deep learning feature extractors (ResNet50). Achieved 73% accuracy across 7 art movements with minimal performance loss after PCA-based dimensionality reduction.
Tech: Python, PyTorch, ResNet50, PCA, CNN, Transfer Learning, OpenCV
Highlights: Computer vision β’ Feature engineering β’ Dimensionality reduction
π GitHub
Published on Medium
Analyzed key factors influencing peer-to-peer credit pricing using Lending Club data. Built an OLS regression model to analyze borrower characteristics and loan attributes. Discovered that Lending Club's internal credit grade alone explained over 90% of interest rate variance, revealing grade as a near-black-box pricing mechanism.
Tech: Python, Regression, Risk Analysis, Statistical Modeling
π Read on Medium
UC Berkeley | 2024
Analyzed global refugee movement patterns using Neo4j graph analytics to uncover migration routes, host country networks, and temporal trends. Visualized complex relational data to surface insights about displacement flows and resettlement patterns across regions and time periods.
Tech: Python, Neo4j, Graph Analytics, Geospatial Visualization, Pandas
Highlights: Graph databases β’ Network analysis β’ Geospatial data
I'm currently exploring opportunities in Data Science, AI Engineering, and Analytics where I can build impactful, human-centered AI.
π§ paulbianka@gmail.com
π linkedin.com/in/bianka-paul-004b03194
π medium.com/@paulbianka
Made with π», β, and a lot of curiosity.