Skip to content
View SandyaBeemaneni's full-sized avatar

Block or report SandyaBeemaneni

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SandyaBeemaneni/README.md

Hi, I'm Sandya Beemaneni πŸ‘‹

Senior Cloud Data Platform Engineer | Senior Data Engineer

I design and support secure, scalable cloud data platforms for analytics, reporting, machine learning, and generative AI workloads.

I have 6+ years of experience developing production data pipelines, building governed data layers, improving data quality, automating infrastructure and deployments, optimizing cloud costs, and supporting business-critical workloads in financial services and healthcare environments.


πŸ› οΈ Technical Skills

Data Engineering

Python SQL PySpark Apache Spark Apache Airflow AWS Glue Azure Data Factory ETL/ELT

Cloud and Data Platforms

AWS Azure Databricks Snowflake Amazon Redshift PostgreSQL MySQL

Platform Engineering

Terraform Docker Kubernetes Helm GitHub Actions Jenkins ArgoCD

Monitoring and Reliability

Prometheus Grafana CloudWatch Azure Monitor ELK

Data Governance and Security

Data Quality Data Validation Reconciliation IAM/RBAC Encryption Audit Logging


πŸš€ Featured Projects

A production-style financial transaction pipeline implementing bronze, silver, quarantine, and gold data layers.

Technologies: Python, PySpark, Apache Airflow, SQL, PostgreSQL, Docker and GitHub Actions.

Project highlights:

  • Bronze, silver, quarantine, and gold data layers
  • Automated data-quality validation
  • Invalid-record quarantine with error reasons
  • Source-to-target reconciliation
  • Airflow-based orchestration
  • SQL reporting models
  • Automated testing and GitHub Actions CI/CD
  • 250 synthetic source records processed
  • 238 valid records and 12 quarantined records
  • 95.2% data-quality pass rate

View the project repository β†’

This project uses synthetic data and contains no employer code, client data, credentials, or proprietary architecture.

A production-style synthetic healthcare claims pipeline implementing bronze, silver, quarantine, and gold lakehouse layers.

Technologies: Python, PySpark, Apache Airflow, SQL, Docker and GitHub Actions.

Project highlights:

  • Synthetic healthcare claims ingestion
  • Duplicate and invalid-claim detection
  • Patient-ID pseudonymization
  • Raw patient IDs and names removed from the curated silver layer
  • Invalid-record quarantine with error reasons
  • Provider, diagnosis, and monthly-utilization analytics
  • Source-to-target reconciliation
  • Automated testing and GitHub Actions CI/CD
  • 300 synthetic claims processed
  • 284 valid masked claims and 16 quarantined claims
  • 94.67% data-quality pass rate

View the project repository β†’

All patient, provider, diagnosis, and claim information is synthetic. The project contains no employer code, client data, credentials, or proprietary architecture.


πŸ“œ Certifications

  • AWS Certified Solutions Architect – Associate
  • Microsoft Certified: Azure Administrator Associate
  • HashiCorp Certified: Terraform Associate
  • Databricks Academy Accreditation – Generative AI Fundamentals

πŸ“Œ Current Engineering Focus

My current areas of technical development include:

  • Cloud data-platform observability
  • Infrastructure as Code for data workloads
  • Pipeline reliability and operational monitoring
  • Governed datasets for analytics and AI/ML workloads
  • Secure deployment automation with Kubernetes and CI/CD

🀝 Connect With Me


All public projects use synthetic or publicly available data and are independent reference implementations.

Pinned Loading

  1. financial-data-platform-demo financial-data-platform-demo Public

    Production-style financial transaction pipeline with Python, PySpark, SQL, Airflow, data quality, Docker and CI/CD.

    Python

  2. healthcare-claims-lakehouse-demo healthcare-claims-lakehouse-demo Public

    Synthetic healthcare claims lakehouse pipeline with PySpark, data quality, PHI masking, bronze-silver-gold layers and analytics.

    Python