AI engineer and IEEE-published researcher with 2+ years of experience in LLM evaluation, enterprise AI deployment, and multi-agent systems. Currently pursuing my MS in Artificial Intelligence at Northeastern University (GPA: 3.91).
Writing a preprint introducing a long-context, multi-turn benchmark for evaluating LLMs in adversarial multi-agent settings β measuring cooperation, deception, persuasion, and social reasoning under uncertainty. Preliminary results across 600+ games and 40+ models show that long-context interaction histories improve collaborative reasoning over compressed baselines.
Comparative study of face recognition models across live-action and animated video content. Evaluated multiple architectures trained from scratch on both domains, benchmarking recognition accuracy, detection rate, and domain gap across frame-level and video-level metrics. Findings highlighted how stylized animated features cause systematic failures in models designed for photorealistic faces.
Extended BRT-Align to show that latent-space geometry governs the effectiveness of RL-based steering for correcting misaligned LLM outputs under adversarial attacks. Benchmarked transformer, CNN, and classical ML classifiers across diverse threat scenarios, demonstrating that high detection accuracy doesn't reliably translate to robust alignment.
Benchmarked single- and multi-node distributed training using PyTorch DDP with NCCL across CNNs, ResNet, and Vision Transformers on CIFAR-10. Quantified tradeoffs between batch size, gradient sync cost, and scaling efficiency across 1-GPU, 2-GPU, and multi-node configurations.
Built a hybrid CNN model combining DenseNet121, ResNet50, and InceptionV3 to detect and classify multiple lung diseases from CT images β achieving 98.61% accuracy. Published at IEEE ICNC 2024.
Python PyTorch TensorFlow LLMs RAG Multi-agent Systems Computer Vision NLP Distributed Training React Java MySQL
