QA / E2E agent guide
This is the operating manual for tests/e2e/, qa/audit/, and the Playwright infrastructure
around them. It's written for a coding agent picking up this repo cold, not just a human
contributor — read this before exploring the app manually.
Why this exists
Before this infrastructure, checking BoardReadyOps' UI required manually clicking through every
route. That's how the 2026-09-01 UI/UX audit findings (policy bypass, fake persistence, broken
focus management, 404s, duplicated titles...) got found by hand, one route at a time. This
system turns that into pnpm qa:audit — one command that checks every known route, every
critical viewport, for the failure classes that audit found, and reports actionable findings
instead of requiring someone to notice them.
tests/unit/web/*.test.ts (axe-core against renderToStaticMarkup/happy-dom) already covers
component-level accessibility and behavior; this layer exists for what only a real browser can
catch: actual hydration timing, real navigation and URL state, real computed CSS (contrast,
overflow, touch target size), and cross-page link integrity.
Quick reference
| Command | What it does | When to run it |
|---|---|---|
pnpm qa:smoke |
Desktop-only, critical-route subset of the audit | Every PR (also runs in CI) |
pnpm qa:audit |
Full route audit, 3 viewports (375/768/1440), one browser | Before a UI-heavy PR, or pnpm qa |
pnpm qa:e2e |
Review lifecycle, modal contract, tabs contract, regression suite | After touching review/modal/tab UI |
pnpm qa:a11y |
Just the axe pass from the audit, desktop, critical routes | Quick accessibility check |
pnpm qa:visual |
Screenshot regression against tests/e2e/*-snapshots/ |
After a visual change (also ci / visual on every web-UI PR) |
pnpm qa:visual:update |
Regenerates visual baselines | Linux only — use the qa-visual-baselines workflow, and review the render yourself |
pnpm qa:cross-browser |
Full audit + qa:e2e across Chromium/Firefox/WebKit | Nightly (also qa-nightly.yml) |
pnpm qa:route-coverage |
Fails if a page.tsx has no qa/audit/routes.ts entry |
Runs as part of qa:audit's file set |
pnpm qa:production-smoke |
Read-only synthetic checks against a real deployment | Manually, with PLAYWRIGHT_BASE_URL set |
pnpm qa:typecheck |
Type-checks tests/e2e/** and qa/** |
Part of pnpm run typecheck already |
pnpm test:e2e (pre-existing) still runs the whole tests/e2e/ directory with no filtering —
useful for "run literally everything once" locally, not for CI (too slow for a PR gate).
Authenticated tests
tests/e2e/global-setup.ts mints a signed brops_session cookie directly — no live GitHub
OAuth round trip — using encodeUserSession() from apps/web/lib/user-session.ts, the exact
function the real callback route uses. QA_SESSION_SECRET (or SESSION_SECRET) can override the
local QA key, but it is not required: both playwright.config.ts and global-setup.ts fall back
to the same fixed local-only placeholder. The storage state is written directly as JSON, so
Firefox/WebKit matrix jobs do not need Chromium installed just to mint the cookie, and routes
marked auth: "authenticated" are actually exercised signed in by default.
A spec that needs to be signed in does:
import { authenticatedStorageState } from "./fixtures/auth.js";
test.use({ storageState: authenticatedStorageState });
This session has no GitHub App installations (installationIds: []), so it satisfies
routes that only check "is someone signed in" (Policies, My Work, Settings). It does not
unlock /dashboard or /repositories/:id — those also need DATABASE_URL pointed at a
Postgres instance with a real installation/repository seeded, because
apps/web/lib/repository-dashboard.ts's loadRepositoryDetail() has no demo fallback at all
(see qa/audit/routes.ts's requiresDb/skipWithoutDb flags — those routes are skipped, not
faked, when DATABASE_URL isn't set). Standing up that seeded Postgres tenant for full
authenticated DB-backed E2E is intentionally not built by this infrastructure yet — see
"What's not done" below.
Adding a route
- Add the
page.tsx. - Add an entry to
qa/audit/routes.ts(id,pathwith fixture ids already substituted,auth,requiresDb).pnpm qa:route-coveragefails the build otherwise — that's deliberate, matching the task's "route exists but has no coverage" requirement. - If it's a route worth a screenshot baseline, add its
idtovisualRoutesin the same file and runpnpm qa:visual:updateonce, then commit the new.png.
Adding a new component/interaction state
- If it's a modal/dialog, run it through
expectDialogContract()fromqa/audit/dialog-contract.ts(seetests/e2e/modal-contract.spec.tsfor the pattern) — it checks the same WAI-ARIA dialog contract (initial focus, focus trap, Escape, focus restore) every modal in this app is expected to satisfy. - If it's a tab pattern, mirror
tests/e2e/tabs-contract.spec.ts: roving tabindex, arrow keys, URL-backing, back/forward. - Otherwise, a targeted
test()in the relevant spec file is usually enough. Don't reach for a Page Object abstraction for a one-off interaction — see "Test quality" below.
Updating visual baselines
pnpm qa:visual:update overwrites tests/e2e/*-snapshots/*.png. Never run this to make a
failing CI run pass without first opening the diff (playwright show-report after a failed run
shows before/after/diff images) and confirming the change is the intended UI change, not a
regression. Commit baseline updates in the same PR as the UI change that caused them, with a
one-line note on why in the commit message.
Where the baselines are checked
ci / visual runs pnpm qa:visual on every pull request that touches the web UI, gated on the
same needs_accessibility signal as the accessibility and qa-e2e jobs. qa-nightly's visual
job runs the same suite on a schedule.
The pull request check is the one that matters. Before it existed, a change could alter the page and merge with a stale baseline while every check was green, and the failure surfaced the next morning on a run nobody was watching — pointing at a page that was correct. That happened once; hence the job.
Regenerating them
Playwright names snapshots per-OS (<id>-chromium-<platform>.png) and the committed baselines
are Linux-only (-chromium-linux.png), because that is what CI runs on. A Windows or macOS
checkout cannot produce them: a local pnpm qa:visual:update writes -win32 or -darwin
files that CI never reads, and the Linux baseline stays stale.
So regenerate through the qa-visual-baselines workflow instead:
- Push the UI change to a branch.
gh workflow run qa-visual-baselines.yml --ref <branch>— it runsqa:visual:updateonubuntu-24.04and uploads the snapshot directory as an artifact. It hascontents: readonly, so it cannot commit; you decide what lands.gh run download <run-id>and compare each file against what is committed. The workflow regenerates all baselines, so check the hashes and commit only the ones that actually changed — otherwise an unrelated drift rides along unreviewed.- Open the new screenshot and look at it before committing. A baseline is worth exactly as much as the render inside it; committing a broken layout teaches the suite that the breakage is correct. Cropping the changed region and reading it takes a minute.
- Commit with a note on what changed and why, ideally citing the run id.
Inspecting a failure
Every CI run uploads playwright-report/ as an artifact (ci / qa-e2e on PRs,
qa-nightly-report-<browser> nightly). Locally:
pnpm exec playwright show-report
opens the last HTML report — timeline, screenshots on failure, and (on retry) a full trace you
can step through frame by frame with pnpm exec playwright show-trace <trace.zip>.
Exploratory / agent-driven QA (Playwright MCP)
There's no --agent CLI flag in the installed Playwright version (1.55.1) for the
Planner/Generator/Healer workflow the original task envisioned — that's a separate, newer
Playwright feature this repo doesn't currently depend on. The practical equivalent, wired up
here, is Playwright MCP (.mcp.json at the repo root, npx @playwright/mcp@latest): any
MCP-capable coding agent (Claude Code, etc.) opened in this repo can drive a real browser
directly — navigate, read the accessibility tree, click, fill forms, screenshot — without
writing a script first.
The intended agent workflow:
- Explore. Use the Playwright MCP browser tools to navigate to the route/state in question
against
pnpm dev(or let Playwright's ownwebServerstart one). Read the accessibility tree and DOM rather than guessing from the source. - Generate. Once you've confirmed the expected behavior by hand, write it as a real
Playwright Test spec under
tests/e2e/—getByRole/getByLabelselectors, nowaitForTimeoutsleeps where an assertion-based wait will do (see "No flaky tests" below). Add the route toqa/audit/routes.tsif it's new. - Heal. When an existing spec starts failing, don't loosen the assertion. Run it with
--debugor inspect its trace, confirm whether the app changed on purpose or regressed, and either update the spec to match an intended change (with the same review discipline as a visual baseline update) or fix the regression.
Production synthetic monitoring
tests/e2e/production-smoke.spec.ts + playwright.production.config.ts — read-only checks
against a real deployed instance:
PLAYWRIGHT_BASE_URL=https://boardreadyops.com pnpm run qa:production-smoke
playwright.production.config.ts refuses to run at all without PLAYWRIGHT_BASE_URL set (no
default that could accidentally point at production), has no webServer (it's not starting a
local instance), and every test in that file only performs GET/navigation checks —
qa/audit/production-guard.ts's guardProductionSafety() is called at the top of the file as a
standing marker that nothing in it may mutate production state. If you ever add a check that
needs to write anything, it cannot go in this file.
This suite is structured to be portable to Checkly or similar synthetic-monitoring platforms
later — each test() is self-contained with no shared mutable state between them, matching how
a browser-check platform runs them. No Checkly account is configured in this repo; this is the
local/CI equivalent until one is.
Cross-browser
playwright.config.ts only defines the chromium project by default (PRs stay fast). Setting
QA_CROSS_BROWSER=1 (which pnpm qa:cross-browser does via scripts/qa-cross-browser.mjs —
a small Node wrapper rather than cross-env, since this repo avoids adding a dependency for one
script and needs it to work on Windows too) also defines firefox and webkit projects. BrowserStack
isn't configured — use: { ...devices[...] } from @playwright/test is what each project uses, so
pointing a project at BrowserStack's remote Chromium/Safari/Edge later is a connectOptions
change to playwright.config.ts, not a rewrite; see
Playwright's BrowserStack guide if that becomes necessary.
Visual regression: native Playwright only
tests/e2e/visual.spec.ts uses expect(page).toHaveScreenshot() — no Chromatic/Percy
dependency. If Storybook is added later (see "What's not done"), Chromatic pairs naturally with
it; if a different SaaS is chosen, make sure it isn't just duplicating what this file already
covers.
No flaky tests
- Selectors:
getByRole/getByLabelfirst,getByTextwhen there's no better role, adata-testid/class selector only when nothing semantic exists (a few already do, matching what's on the actual DOM — e.g..finding-triage-card,.disposition-select). - No
waitForTimeoutas a substitute for an assertion. The fewwaitForTimeoutcalls that exist in this suite (e.g. after opening a review) wait out a known post-hydration settle window documented inline, not a guess at how long an async operation takes — prefer awaitFor({ state: "visible" })orexpect(...).toBeVisible()instead when you can. - Fixture ids (
qa/audit/routes.ts'sdemoReviewId,demoRunId) are stable, checked-in constants, not generated per run — see "What's not done" for the larger seeded-tenant gap this doesn't yet solve for authenticated/DB-backed routes.
Test quality
- One helper function reused 3+ times earns a shared home in
qa/audit/; a one-off interaction stays inline in its spec. - No Page Object Model layer — this app's DOM is stable enough (and Playwright's locators already lazy-resolve) that POM would be an abstraction with no real payoff yet. Revisit if spec files start duplicating the same 10-line interaction verbatim.
- Assert the actual outcome (a value changed, a request was made, state survived reload), not just "the element became visible" where a stronger assertion is available for free.
Security
tests/e2e/.auth/(storageState with the signed test session) is gitignored — never commit it.- The QA session (
global-setup.ts) carries a syntheticuserId/login, never a real credential, and is signed with a secret you provide, not one baked into the repo. qa/audit/checks.ts's console-error allowlist is a short, exact-substring list, on purpose — broadening it to a prefix or regex defeats the point of catching real console errors.- Production-facing tests are read-only by construction (see above); there is no code path in
this infrastructure that can mutate
https://boardreadyops.com.
What's not done
Built deliberately, not by accident — see the setup task's own instruction not to make large unrelated changes while standing this up:
- Seeded Postgres QA tenant (task section 5): the ~30 repository/review/run/policy fixture
combinations the original task describes need real rows in a disposable Postgres database,
not just the existing
DEMO_REVIEWS/buildDemoRunin-memory fixtures this suite reuses.docker-compose.ymlunderdeploy/and theDATABASE_URL=postgresql://boardreadyops@127.0.0.1:5432/boardreadyops_toolchainconvention already used bytest:intare the natural starting point for that follow-up; it's sized as its own task, not a quick addition here. - Playwright Test Agents (Planner/Generator/Healer as a native feature): not available in the installed Playwright version; Playwright MCP (above) is the practical substitute wired up in this pass.
- Storybook: evaluated, not added.
tests/unit/web/*.test.ts(axe +renderToStaticMarkup) and this E2E layer already give the state-heavy components listed in the original task (ApprovalModal, FindingsTab, etc.) real coverage in the context they actually render in; a component gallery would add value for isolated visual QA and design review, but is a standalone toolchain addition (new build config, new CI job, dozens of story files) better scoped as its own task than folded into this one. - Lighthouse CI: see
.github/workflows/lighthouse.ymland.lighthouserc.json— wired up for a handful of key routes, accessibility regressions fail the run, performance/best-practices/SEO are warn-level baselines rather than a score target, per the original task's explicit "don't chase Lighthouse scores" instruction. - Checkly: documented above as the migration target for
production-smoke.spec.ts, not integrated — no account/credentials available in this environment. - DB-backed authenticated coverage:
qa:auditnow always has a deterministic signed-in QA session, so an unexpected 401 is a real P0 instead of an allowlisted local-dev artifact. Routes that require an installation/repository seeded in Postgres are still skipped or degraded whenDATABASE_URLis absent; a disposable seeded tenant remains a separate infrastructure task. - Visual baselines: Linux Chromium baselines are checked in for every
visualRoutesentry and are checked byci / visualon every web-UI pull request, not only nightly. Regenerate them through theqa-visual-baselinesworkflow — a localqa:visual:updatewrites a-win32/-darwinfile CI never reads — and look at the render before committing it.