I design and support secure, scalable cloud data platforms for analytics, reporting, machine learning, and generative AI workloads.
I have 6+ years of experience developing production data pipelines, building governed data layers, improving data quality, automating infrastructure and deployments, optimizing cloud costs, and supporting business-critical workloads in financial services and healthcare environments.
Python SQL PySpark Apache Spark Apache Airflow AWS Glue Azure Data Factory ETL/ELT
AWS Azure Databricks Snowflake Amazon Redshift PostgreSQL MySQL
Terraform Docker Kubernetes Helm GitHub Actions Jenkins ArgoCD
Prometheus Grafana CloudWatch Azure Monitor ELK
Data Quality Data Validation Reconciliation IAM/RBAC Encryption Audit Logging
A production-style financial transaction pipeline implementing bronze, silver, quarantine, and gold data layers.
Technologies: Python, PySpark, Apache Airflow, SQL, PostgreSQL, Docker and GitHub Actions.
Project highlights:
- Bronze, silver, quarantine, and gold data layers
- Automated data-quality validation
- Invalid-record quarantine with error reasons
- Source-to-target reconciliation
- Airflow-based orchestration
- SQL reporting models
- Automated testing and GitHub Actions CI/CD
- 250 synthetic source records processed
- 238 valid records and 12 quarantined records
- 95.2% data-quality pass rate
View the project repository β
This project uses synthetic data and contains no employer code, client data, credentials, or proprietary architecture.
A production-style synthetic healthcare claims pipeline implementing bronze, silver, quarantine, and gold lakehouse layers.
Technologies: Python, PySpark, Apache Airflow, SQL, Docker and GitHub Actions.
Project highlights:
- Synthetic healthcare claims ingestion
- Duplicate and invalid-claim detection
- Patient-ID pseudonymization
- Raw patient IDs and names removed from the curated silver layer
- Invalid-record quarantine with error reasons
- Provider, diagnosis, and monthly-utilization analytics
- Source-to-target reconciliation
- Automated testing and GitHub Actions CI/CD
- 300 synthetic claims processed
- 284 valid masked claims and 16 quarantined claims
- 94.67% data-quality pass rate
View the project repository β
All patient, provider, diagnosis, and claim information is synthetic. The project contains no employer code, client data, credentials, or proprietary architecture.
- AWS Certified Solutions Architect β Associate
- Microsoft Certified: Azure Administrator Associate
- HashiCorp Certified: Terraform Associate
- Databricks Academy Accreditation β Generative AI Fundamentals
My current areas of technical development include:
- Cloud data-platform observability
- Infrastructure as Code for data workloads
- Pipeline reliability and operational monitoring
- Governed datasets for analytics and AI/ML workloads
- Secure deployment automation with Kubernetes and CI/CD
- Portfolio: sandya-beemaneni.vercel.app
- LinkedIn: Sandya Beemaneni
- Email: bsandya1926@gmail.com
All public projects use synthetic or publicly available data and are independent reference implementations.
