I am a Test Architect and AI Quality Engineering professional with experience in enterprise test strategy, intelligent test automation, AI-powered quality engineering, and modern software-testing practices.
My current focus is on designing and building quality-engineering solutions using Agentic AI, Retrieval-Augmented Generation (RAG), Large Language Models, Playwright, multi-agent workflows, and automated AI evaluation.
- AI and LLM application testing
- Agentic AI quality-engineering workflows
- RAG application testing and evaluation
- LLM-as-a-Judge validation
- Hallucination, faithfulness, groundedness, and relevance testing
- Playwright UI and API automation
- BDD and Cucumber frameworks
- Test architecture and automation strategy
- API, performance, accessibility, and security testing
- CI/CD quality gates
- AI observability, latency, token, and cost monitoring
- Responsible AI testing and governance
- Building AI-powered test-design solutions
- Generating test scenarios from requirements and user stories
- Evaluating RAG retrieval precision, recall, and relevance
- Measuring faithfulness, groundedness, hallucination, and answer quality
- Creating multi-agent quality-engineering workflows
- Integrating Playwright automation with AI agents
- Implementing AI quality gates in CI/CD pipelines
- Designing human-in-the-loop validation and governance controls
- Tracking model performance, latency, token usage, and cost
Agentic AI Generative AI RAG LLM Evaluation Prompt Engineering LLM-as-a-Judge Responsible AI
Azure OpenAI Amazon Bedrock LangChain LangGraph RAGAS Arize Phoenix ChromaDB Vector Databases
Playwright Selenium Appium Cucumber Pytest Postman Newman REST Assured Katalon
GitHub Actions Jenkins Azure DevOps Docker AWS Git Allure Reports
Python TypeScript JavaScript Java SQL Gherkin
An AI-powered RAG platform that transforms requirement documents and user stories into structured, traceable, and review-ready test scenarios and test cases.
Key capabilities:
- Multi-format requirement-document ingestion
- ChromaDB-based vector storage
- Hybrid semantic and BM25 retrieval
- Project-isolated knowledge management
- Positive, negative, boundary, and alternate-flow test generation
- Requirement-to-test traceability
- BDD and Gherkin scenario generation
- LLM-as-a-Judge quality evaluation
- Jira requirement ingestion
- Token usage and AI observability
- Structured test-case export
- Role-based access and audit logging
Technologies:
Python Flask ChromaDB RAG LLM Evaluation RAGAS Arize Phoenix Azure OpenAI Amazon Bedrock
A multi-agent quality-engineering framework designed to support test analysis, automation implementation, execution, review, security validation, and reporting.
Key capabilities:
- Specialized quality-engineering agents
- Requirement and test-analysis workflows
- AI-assisted test design
- Automation implementation support
- Playwright-based testing capabilities
- Test execution and failure analysis
- Security and quality-review workflows
- Reusable agent skills, prompts, and templates
- Human-in-the-loop approval controls
- Documentation and governance guidance
- Extensible multi-agent architecture
Technologies:
Agentic AI JavaScript Node.js Playwright Multi-Agent Systems Quality Engineering Security Testing
A reusable evaluation framework for validating the quality, reliability, performance, and cost of RAG and LLM applications.
Planned evaluation capabilities:
- Context precision and recall
- Retrieval relevance
- Faithfulness and groundedness
- Answer relevance
- Hallucination detection
- Citation correctness
- LLM-as-a-Judge evaluation
- Ground-truth comparison
- Prompt and model comparison
- Configurable quality thresholds
- Latency measurement
- Token-consumption tracking
- Cost estimation
- JSON, CSV, and HTML reports
- CI/CD quality gates
- Phoenix and OpenTelemetry tracing
Requirement Documents
β
RAG-Based Knowledge Retrieval
β
AI-Powered Test Design
β
Multi-Agent Automation Workflow
β
Playwright UI and API Execution
β
LLM and RAG Quality Evaluation
β
Observability, Governance, and Reporting