Attorney Share

QA at team scale, run with Claude

Legal marketplace

QA & Test AutomationPlaywright E2ECI/CDObservabilityClaude Code

~500

E2E tests

across ~25 suites

94.3%

pass rate

42,789 runs

83.1s

avg suite

per execution

80

flaky tests

surfaced & tracked

Brief format

One engineer, team-scale output

The only QA engineer on this engagement. Claude Code is the co-engineer that lets a one-person QA function operate at team scale: I plan and review, Claude builds, and every change is mine to own. All Playwright E2E — ~500 tests, ~25 product areas, ~90 spec files, 27 page models, ~47 utilities — none of which existed before this engagement.

How I work with Claude

A co-engineer, not autocomplete. Big efforts start as a phased written plan Claude holds and re-plans the moment reality contradicts an assumption. Work lands as many small, reviewable pull requests instead of one risky batch. Verdicts come from end-to-end tests against deployed environments — the way the user actually hits them — not isolated unit checks. And a persistent context file carries decisions and conventions across repos and sessions, so Claude keeps the project instead of relearning it cold.

The system, raised in four moves

Built from scratch — the Playwright framework, page models, fixtures, auth and CI against live preview environments. Accelerated — two disciplined refactors stripped redundancy and roughly halved runtime (~24 → ~12 min), shipped as small reviewable changes, not one rewrite. Made intelligent — Smart Test Selection reads each PR's diff and recommends the tests it needs, live in shadow mode (skips nothing yet), guarded by drift detection, confidence thresholds and an auto-revert circuit breaker. Made measurable — Exolar turns failures, flaky signals and every selection decision into data.

Exolar — observability I built with Claude

Every test reports to Exolar: failures, stack traces and flaky-test detection, plus every decision the smart selector logs. Exolar also exposes an MCP, so Claude can read the run logs itself and tell a real bug from a flaky test without me pulling logs by hand. Reliability came with it — 80 flaky tests surfaced and fixed at the root cause, not masked with retries. Quality and CI cost are proven with data, not guessed.

Said plainly

The cheap, per-PR inference inside Smart Test Selection runs on a small, low-cost non-Claude model — by design, because it runs on every pull request. Claude (Claude Code) is the engineer that designed and built the system: the framework, the analytics and the safeguards. I review and own every change. The value is sustained engineering judgment, not a single API call.

Have a similar challenge?

Get a senior engineering lead on a call and a concrete plan in days, not months.

Talk to Distillery