Skip to content
View shbhamdbey's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report shbhamdbey

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
shbhamdbey/README.md

Hi πŸ‘‹, I'm Shubham Dubey

AI Data Engineer by day. Distributed Systems tinkerer by night. Pizza volume expert 24/7.

shbhamdbey


🧠 What I Do

I build AI-powered data pipelines that make LLMs actually useful in production. Over the past 4+ years, I've shipped systems for Pfizer, Covetrus, and NJIT β€” from event-driven streaming architectures to GPU-clustered LLM inference engines.

  • πŸŽ“ M.S. Computer Science β€” New Jersey Institute of Technology
  • πŸ€– Current role: AI Data Engineer β€” I don't just move data, I make it think
  • ⚑ Superpower: Turning research papers into running code (see my "From Scratch" projects below)

πŸš€ Recent Work: LLM-Powered Healthcare Intelligence

5 LLM inference pipelines. 2 fine-tuning loops. 2 eval gates. One very busy GPU cluster.

I architected and shipped a production LLM stack for a healthcare data platform that uses Gemma-4 models (4B β†’ 31B parameters) to classify, normalize, and QA-score patient records at scale:

  • 🏷️ Taxonomy Classification β€” FAISS-GPU + EmbeddingGemma vector search β†’ Gemma-4-E4B-it breed/species classification with 3-stage QA scoring
  • ⚧ Gender Normalization β€” LLM-driven standardization of raw gender descriptions into 7 clinical categories with plausibility scoring
  • πŸ“ Transcription QA β€” Gemma-4-31B-IT-QAT evaluating raw clinical transcripts for quality, plausibility, and completeness
  • πŸ“‹ SOAP Note QA β€” Faithfulness scoring of AI-generated summaries against source transcripts (hallucination detection, omission tracking, medication error flags)
  • πŸ”„ Continuous Fine-Tuning β€” Closed-loop eval gates: train β†’ holdout eval β†’ alias promotion only if accuracy beats production by margin

Stack: vLLM Β· Ray on Spark Β· FAISS-GPU Β· dbt Β· MLflow Β· L40S GPUs Β· Unity Catalog


πŸ› οΈ My "From Scratch" Projects

"Don't just use the tool. Understand the tool. Then build the tool."

I got tired of just using Kafka and Spark, so I built them from the ground up in Python β€” straight from the original research papers. No libraries. No shortcuts. Just pure computer science.

πŸ”₯ Kafka from Scratch

Built from: "Kafka: a Distributed Messaging System for Log Processing" (Kreps et al., LinkedIn, 2011)

A working TCP broker with:

  • βœ… Append-only log storage with .log + sparse .index files
  • βœ… Fixed-header binary format ([4B Length | 8B Offset | 8B Timestamp | Payload])
  • βœ… Consumer-driven offsets β€” the broker is completely stateless
  • βœ… O(log n) binary search on the sparse index for fast seeks
  • βœ… PRODUCE / FETCH wire protocol over TCP
  • βœ… ~500 lines of Python, zero dependencies

Because reading the paper is cool. Making the paper actually run is cooler.

⚑ Spark from Scratch

Built from: "Resilient Distributed Datasets" (Zaharia et al., UC Berkeley, 2012)

A mini Spark engine featuring:

  • βœ… RDD lineage graph β€” fault tolerance without replication
  • βœ… Lazy evaluation β€” transformations build a DAG, actions trigger execution
  • βœ… Narrow dependencies β€” pipelined with Python generators (no intermediate storage!)
  • βœ… Wide dependencies β€” shuffle stages with HashPartitioner
  • βœ… DAG Scheduler β€” automatically splits jobs at shuffle boundaries
  • βœ… Thread-pool executor with retry logic (lineage-based recovery)
  • βœ… ~600 lines of Python, zero dependencies

I now understand why groupByKey is a shuffle boundary on a spiritual level.


πŸ’¬ Ask Me About

LLMs     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Production inference, fine-tuning, eval gates
Python   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Data pipelines & "from scratch" engines
Kafka    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Event streaming & log-centric architectures
Spark    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Distributed computing & DAG optimization
SQL      β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  The OG data language
NoSQL    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      When relationships get complicated
AWS      β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      Cloud infrastructure
Azure    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      More cloud infrastructure
Tableau  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      Making data pretty
CI/CD    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      Shipping things that don't break
Java     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ      Building robust backends
ML       β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ          Teaching machines to be slightly less dumb

🌱 Currently Exploring

  • AI/LLM Engineering β€” inference optimization, structured output, hallucination detection
  • Data Engineering β€” streaming, ELT, dbt, modern data stacks
  • Distributed Systems β€” one paper implementation at a time

πŸ“« Reach Me


⚑ Fun Fact

A pizza that has radius "z" and height "a" has volume Ο€ Γ— z Γ— z Γ— a.

Yes, I will bring this up in every technical interview. No, I will not apologize. πŸ•


python java aws docker kafka spark mongodb mysql linux

Built with caffeine, curiosity, and a stubborn refusal to accept "it just works" as an answer.

Pinned Loading

  1. kafka-Implementation kafka-Implementation Public

    My Own Kafka

    Python

  2. Pfizer_Capstone Pfizer_Capstone Public

    Jupyter Notebook

  3. Sales-Predictions-Using-Marketing-Performance- Sales-Predictions-Using-Marketing-Performance- Public

    Jupyter Notebook

  4. spark-Implementation spark-Implementation Public

    My Own Spark

    Python

  5. Spotify-Kafka-Flink-Streaming-Data-Pipline Spotify-Kafka-Flink-Streaming-Data-Pipline Public

    Dockerfile