Engineering low-latency agentic workflows, real-time voice streaming systems, and production-grade RAG architectures.
I am an Applied AI Engineer focused on bridging generative models with reliable, production-ready backend systems. My recent work centers on sub-second voice-to-voice agents, event-driven code governance pipelines, and automated application security testing, emphasizing deterministic evaluation, low latency, and clean systems architecture.
- 🌐 Available For: Full-Time / Contract Applied AI Engineer & Founding AI Engineer roles (Worldwide Remote — US & EU timezone overlap).
- 🎯 Core Focus: Full-Stack AI Engineering, Real-Time Bidirectional WebSockets, Multi-Agent Orchestration, and RAG Platforms.
An autonomous, voice-driven AI agent for real-time infrastructure triage and remediation.
- Tech Stack: Python, FastAPI, WebSockets, Next.js, Groq, Pytest (Eval Harness).
- Key Highlight: Implemented strict Eval-Driven Development (EDD) to guarantee zero-hallucination tool calling, allowing the agent to safely execute
fetch_logsandexecute_rollbackcommands during simulated outages.
A high-throughput WebSocket API for deploying low-latency conversational voice agents.
- Tech Stack: Node.js, Express, WebSockets, FAISS Vector Search, Streaming STT/TTS Pipelines.
- Key Highlight: Sub-800ms full-duplex voice response latency with integrated vector knowledge retrieval and connection-level WebSocket authentication.
An automated dynamic application security testing (DAST) suite for auditing web endpoints and APIs against OWASP Top 10 vulnerabilities.
- Tech Stack: Python, Flask, Next.js, Tailwind CSS, Cryptographic Domain Verification.
- Key Highlight: Executes 18 modular attack analyzers (SQLi, SSRF, BAC, CSRF, CSP) with automated DNS/meta-tag ownership validation and structured risk reporting.
An automated, serverless code analysis pipeline that reviews GitHub Pull Requests via live webhooks.
- Tech Stack: FastAPI (Python), HMAC-SHA256 Auth, Unified Diff Token-Chunking Engine, GitHub Apps API.
- Key Highlight: Evaluates multi-file unified diffs (up to 32k tokens) and dispatches structured, line-level security and performance feedback in under 15 seconds.
An autonomous application pipeline engine featuring multi-agent orchestration and a terminal UI.
- Tech Stack: Go (Terminal UI / Catppuccin Theme), Node.js, Bash Automation, Multi-Agent Workflows.
- Key Highlight: Full zero-data-loss pipeline with deterministic deduplication, automated CI/CD with CodeQL/SBOM, and real-time operational observability.
A test-driven Retrieval-Augmented Generation document analysis engine.
- Tech Stack: Python, Pytest, Custom Semantic Extractors, Vector Similarity Retrieval.
- Key Highlight: Complete test coverage across chunking, normalization, and semantic query parsing pipelines.
| Domain | Technologies & Frameworks |
|---|---|
| Applied AI & LLMs | OpenAI API, Anthropic Claude, LangChain, LangGraph, RAG, Prompt Engineering, FAISS, Embeddings |
| Real-Time Voice & Audio | WebSockets, Twilio Media Streams, G.711 / PCM Audio Streaming, Speech-to-Text (STT), Text-to-Speech (TTS) |
| Backend & APIs | Python (FastAPI, Flask), Node.js (Express), Go, RESTful APIs, OpenAPI/Swagger, HMAC Auth |
| Databases & State | PostgreSQL, Supabase, Redis, Local Vector Stores, In-Memory Caching |
| Frontend & UI | Next.js, React, TypeScript, Tailwind CSS, Terminal UIs (TUI) |
| DevOps & Security | Docker, GitHub Actions CI/CD, CodeQL, DAST Scanning, Linux / Bash Scripting, Git |
- Latency Over Bloat: Real-time AI must feel immediate. Optimizing payload sizes, streaming audio chunks, and trimming middleware latency always beats adding heavier models.
- Deterministic Guardrails: LLM outputs must be parsed into strict schemas (Pydantic / Zod) before interacting with databases, APIs, or user interfaces.
- Observability & Reliability: Production AI requires continuous testing, robust error fallback states, and structured JSON logs to diagnose hallucinations and rate-limit bottlenecks.
- LinkedIn: linkedin.com/in/rehantariqbhatti
- GitHub: github.com/reh1t


