Dhiraj Gupta — Senior QA Automation Engineer 17+ years QA experience | AT&T Labs | AI Testing Specialist
A QA testing framework for Generative AI applications. Tests LLM evaluation, RAG validation, prompt injection security, semantic similarity, and API endpoints.
- Python 3.11
- Playwright
- PyTest
- Google Gemini API
- pytest-html
| File | Tests | Description |
|---|---|---|
| test_llm_evaluation_mock.py | 5 | LLM evaluation — no API needed |
| test_rag_validation_mock.py | 7 | RAG retrieval validation |
| test_prompt_injection_mock.py | 7 | Security injection testing |
| test_semantic_similarity_mock.py | 7 | Meaning consistency testing |
| test_api_endpoints_mock.py | 10 | REST API endpoint testing |
| test_llm_evaluation.py | 9 | Live Gemini API — batched |
Total: 36 mock tests + 9 live API tests
# Install dependencies
pip install -r requirements.txt
# Run all mock tests
pytest tests/ -v --html=reports/test_report.html --self-contained-html
# Run specific suite
pytest tests/test_rag_validation_mock.py -v -s# Add GOOGLE_API_KEY to .env file first
# Run one batch at a time to save free quota
pytest tests/test_llm_evaluation.py -m batch1 -v -s
pytest tests/test_llm_evaluation.py -m batch2 -v -s
pytest tests/test_llm_evaluation.py -m batch3 -v -s
pytest tests/test_llm_evaluation.py -m batch4 -v -s
pytest tests/test_llm_evaluation.py -m batch5 -v -sAll 36 mock tests passing. 0 failures.
- LLM Evaluation — Response accuracy, hallucination prevention
- RAG Testing — Document retrieval, context accuracy
- Prompt Injection — Security attack prevention
- Semantic Similarity — Meaning consistency validation
- API Testing — Status codes, schema, performance SLA
See all automation projects: https://github.com/rinkugupta3?tab=repositories