Open to QA leadership & AI-quality roles

Quality engineering,
with AI built in.

I'm Rahul Prajapati — QA & Automation Analyst (Acting QA Lead) at BDO Canada. I make enterprise platforms and AI agents trustworthy, and I build the agents that make testing faster.

Rahul Prajapati, QA & Automation Analyst in Toronto
Toronto, ON5+ yrs enterprise SaaS QA
40 → 0%automation coverage in 7 months
0%faster regression cycles
0%faster test design with AI agents
0QA engineers led & mentored

01 — Strengths

Four ways I create impact

The at-a-glance version. Every number below is one I stand behind in interviews.

My core strength in AI: I bring ideas to life — closing the gap between people who have a problem or an idea and a working solution built with AI tools, then automating it to increase efficiency and scale.

QA Engineering

Coverage 40 → 70% · 30% less defect leakage

  • Selenium/TestNG framework architecture
  • CI/CD quality gates: 50% faster regression
  • Release governance & go-live sign-off

AI & LLM Testing

25% higher agent accuracy · 30% fewer bad outputs

  • 100+ prompt variations, hallucination detection
  • Copilot agents for test generation (87% faster)
  • Blind pairwise LLM evals (Mercor LLMArena)

Leadership

Team of 5 · 20% productivity gain

  • Onboarding time halved via SOPs & mentoring
  • UAT coordination across business stakeholders
  • 2× Gold Award + Standing Ovation

Product Thinking

373% ROI option chosen over 17%

  • Build-vs-Buy CBA: 6-sheet, 3-year model
  • PRDs & product concepts, validation scoring
  • Shipping AI apps in public (RunningOak Labs)

02 — Case studies

Real problems, measured outcomes

Details anonymised where client confidentiality requires. Every number is one I can defend in an interview.

01
AI agent testing

Making an enterprise AI agent tell the truth

Problem. An ITSM virtual agent (Microsoft Teams + ServiceNow) answered employee requests using live enterprise data — and sometimes answered wrong. No playbook existed for testing it.

Approach. Built the AI/LLM testing strategy from scratch: 100+ structured prompt variations, hallucination detection, response-accuracy scoring, data-retrieval validation against ServiceNow records, and regression-safe prompt refinement.

Result. A measurably more accurate and safer agent in production, plus a reusable evaluation playbook the team still runs.

25%higher response accuracy
30%fewer erroneous outputs
100+prompt variations tested
02
AI-assisted testing

Copilot agents that write test cases

Problem. Test design was the bottleneck: manual case creation ate hours every sprint, and ambiguous requirements kept surfacing late as rework.

Approach. Engineered internal Copilot-based agents for automated test-case generation and early requirement validation — flagging ambiguity before development, not after.

Result. Test design stopped being the bottleneck; requirement gaps now surface in refinement instead of QA.

87%faster test design
15 hrsreclaimed each week
35%reduction in dev rework
03
Quality strategy · business case

Build vs buy: a QA automation platform decision

Problem. Should we build an automation framework (Playwright MCP + GitHub Copilot) or buy an enterprise platform (Katalon / AccelQ)? Opinions were cheap; the decision wasn't.

Approach. Built a six-sheet cost-benefit workbook: 3-year horizon, vendor evaluation (security certs, Entra auth, CRM integration, AI capabilities), sensitivity analysis, and a risk register — executive-ready.

Result. Clear, defensible numbers: Buy returned ~373% ROI with year-one payback against ~17% for Build. The decision made itself.

373%ROI on the chosen path
3-yearfinancial model
6-sheetexecutive workbook
04
LLM evaluation

Blind LLM judging under sprint conditions — Mercor LLMArena

Problem. Which of two anonymous model responses is actually better? Rank them — accurately, consistently, at speed, in a 24-hour evaluation sprint.

Approach. RLHF-style pairwise preference ranking with rubric-driven scoring: accuracy, completeness, structure, instruction-following, hallucination risk — every judgment backed by a written rationale.

Result. Hands-on experience with the exact evaluation methodology behind modern model training, from the human side of the loop.

Pairwisepreference ranking
Rubricdriven scoring
24 hrsevaluation sprint
05
Design study · RAG

Designing a personal RAG system

Problem. A year of knowledge scattered across nine AI accounts. How do you build one queryable brain out of it — without leaking secrets into embeddings?

Approach. Full architecture: Supabase schema (messages + chunks), 1536-dim embeddings with an HNSW cosine index, a redaction module scrubbing keys/tokens/PII before persistence, and CI/CD config.

Result. An honest lesson in product thinking: a structured markdown system solved 80% of the problem with none of the code — so that's what shipped. Full design available on request.

80%solved without code
Vectorsearch architecture
Redactionpipeline design

03 — Toolkit

The tools I actually reach for

Honest levels — I'd rather show you where I'm strong than pad a logo wall.

QA & Test Engineering

  • Manual · functional · regression · UAT
  • Test strategy & defect management
  • Release management & go-live sign-off
  • Accessibility · NVDA · WCAG
  • Performance & load · NeoLoad · JMeter
  • Security · OWASP ZAP · Veracode

Automation & CI/CD

  • Selenium · TestNG · Cucumber
  • Playwright + Playwright MCP
  • Postman · API testing
  • Azure DevOps · JIRA · Zephyr
  • GitHub Actions · Copilot
  • SQL · MS SQL · MySQL
  • TypeScript

AI × Quality

  • AI agent testing
  • LLM evaluation · pairwise · rubrics
  • AI-assisted test generation
  • Prompt engineering
  • RAG · vector search · embeddings
  • MCP · Claude Agent SDK

AI Build & Product

  • Antigravity · Copilot Studio
  • Custom GPTs & Gems · n8n
  • PRD / TRD · product concepts
  • Business case · ROI modeling
  • Vendor evaluation
PlatformsServiceNow · Workday · Azure AD · Power Automate · Supabase · Vercel · Testim · Reflect.run · TestLodge · Zephyr Scale
CertificationsAzure AI-900 · AZ-900 · DP-900 · AWS Security Fundamentals

04 — Projects

What I build and ship

Studied in public at github.com/Rahulprajapati99.

Finance Dashboard

↗

JavaScript · Antigravity · Playwright E2E

A personal finance dashboard built end-to-end with AI-assisted development, then covered by a 23-test Playwright suite running green in CI.

View repo →

RunningOak Labs

Building

App studio · Play Store approved

My studio for simple, useful AI-powered apps. Google Play developer account approved; first Android app in build — shipped in public.

Follow on LinkedIn →

Life Centralization System

Design study · RAG · PWA

A personal "central brain" — finance and career dashboard with RAG retrieval over years of notes. Fully architected, then deliberately not over-built.

Design study

Automated opportunity hunt

↗

Open source · Gemini · Telegram

Rebuilt an open-source AI job-search agent around my own workflow: paste a job link, get a fit score and a tailored CV with anti-fabrication guardrails, delivered over Telegram.

View repo →

In the lab

PRDs

Product concepts

Concepts in the pipeline: Color Vibes, TailBlazer, a notary app, and a Bloomberg-style terminal — candidates for RunningOak's first ship.

In discovery

The workshop

↗

GitHub · AI-agent ecosystem

Hands-on across agent skills, LLM-evaluation harnesses and AI tooling — staying current with how modern AI systems are actually built and tested.

View GitHub →

05 — About

How I got here

Five years ago I started in QA testing municipal web apps. Today I sign off on enterprise releases and design testing strategies for AI agents that answer with live business data. The common thread: I like finding out where systems break before users do.

At BDO Canada I own quality strategy for enterprise billing and AI-enabled platforms — coordinating UAT with business stakeholders, architecting automation frameworks, implementing CI/CD quality gates, and leading a team of five QA analysts. Three internal awards along the way (two Gold Awards and a Standing Ovation).

What I'm building toward: quality leadership for AI products — the person who decides not just whether it works, but whether it should ship.

Off hours I'm usually planning the next road trip (Newfoundland icebergs, most recently), building apps, or working toward a stubborn life goal of visiting every country on Earth.

06 — Contact

Let's talk

Open to conversations about senior AI quality engineering roles, AI agent testing, and QA strategy.

rahul.connectx@gmail.com