AI-BUILT WEBSITE VALIDATION
QA for AI-built websites: validate before you advertise
To QA an AI-built website, write down its business rules, check critical journeys in authorized staging, and verify outcomes beyond the visible interface. Whether you use Lovable, Bolt, Polsia, or another tool, the implementation still requires independent verification against your actual business expectations. This guide explains what to check, which blind spots to consider, and how to produce evidence that supports a confident launch decision. ESTECH is independent and is not affiliated with any builder.
Verify against business expectations, not just the UI
Any implementation—generated or hand-written—can produce a plausible interface that passes visual inspection but fails under real usage. Test against written business rules: who can perform each action, what happens when inputs are invalid or missing, how the system recovers from transient failures, and whether state remains consistent across sessions and devices.
Automated tests derived from the same generation context may mirror the implementation rather than challenge it. This is a risk scenario worth checking for in any project. Write acceptance criteria from the user and business perspective first, then verify that the implementation satisfies them. Review the assertions and execution evidence independently, including when AI helped write the tests.
Permissions, sessions, and server-side authorization
Hidden buttons and client-side route guards cannot enforce server authorization. In your authorized staging environment with synthetic test accounts, verify that each agreed privileged action enforces authorization on the server: attempt the action with an unauthenticated session, with a lower-privilege role, and with an expired token. Confirm that the server rejects unauthorized requests with appropriate status codes and does not leak sensitive data in error responses. Ask the engineering owner to help verify the API behavior; these functional role checks do not constitute a penetration test.
Test basic session behavior in your synthetic staging environment: concurrent logins, token refresh, and logout behavior against the agreed session policy.
Payments, retries, and state consistency
Use sandbox or test-mode payment providers during QA. Production payment testing is excluded from this guide. In sandbox mode, verify the conditions applicable to your product: successful payment, declined payment with retry, and duplicate submission prevention. Configure test inboxes and approved notification destinations as well; staging alone does not prevent real emails or charges.
State consistency is a common risk area in e-commerce and SaaS implementations. After a failed payment retry, confirm that discounts, cart contents, and user selections are preserved. After a form resubmission, confirm that the system does not create duplicate records or send duplicate notifications. A single successful checkout does not exercise these recovery paths; verify them explicitly regardless of how the code was produced.
Secrets and environment hygiene
Assign a qualified developer to review generated client bundles, repository history, and configuration files for unintended keys or credentials. Never include actual secrets in reports, logs, or evidence artifacts. If a key may have been exposed, follow your provider’s incident and revocation process immediately. Ensure staging and production use separate credentials, databases, and third-party accounts.
Restrict QA access to the minimum scope needed. Testers should not have production write access, unrestricted database queries, or admin privileges beyond the agreed test scope. Document who has access to what, and revoke it when the test window closes. A standard pilot engagement does not include a full security audit; treat this review as a targeted hygiene check within the agreed scope.
Mobile, forms, loading states, and error handling
Responsive layouts can break at specific breakpoints or under slow network conditions in any implementation. Test critical journeys at mobile viewports, with throttled connections, and with JavaScript disabled where progressive enhancement is claimed. Verify that form validation messages are visible, associated with the correct fields, and do not disappear on re-render.
User-triggered pending operations need useful feedback. Confirm that actions initiated by the user provide clear status indication, handle timeout gracefully, display a user-meaningful error message on failure, and offer a recovery path (retry, cancel, or navigate elsewhere). Not every asynchronous operation requires a visible loader; focus on operations where the user would otherwise be left uncertain about progress or outcome. Empty states should explain what the user can do next, not just that nothing exists.
Launch triage table
Use this adaptable table to decide what to verify before advertising or accepting users. Applicability depends on your product, audience, and risk profile. Basic mobile usability and error recovery are essential if your audience or use case depends on them; other items may be deferred with documented risk acceptance. There is no universal required vs. recommended classification for all products.
| Area | Check | Applicability | Evidence |
|---|---|---|---|
| Authorization | Server rejects unauthorized and expired-session requests | Products with accounts or restricted actions | API response logs + screenshot |
| Payments | Sandbox payment, decline, retry, and duplicate prevention | If applicable | Sandbox transaction IDs + UI trace |
| Data integrity | No duplicate records after resubmit or retry | Actions that create or update records | Authorized record check + request reference |
| Secrets | No exposed private keys; staging/prod credentials separated | Products using credentials or external services | Config review summary |
| Mobile | Critical journeys usable at target viewports with throttled network | Essential if audience-dependent | Screenshots + performance trace |
| Error handling | User-meaningful errors and recovery paths for user-triggered operations | Essential if audience-dependent | Screenshot + console/network log |
| Accessibility | Keyboard navigable; visible focus; basic contrast and labeling | Public forms and interactive controls | Manual walkthrough notes |
Classifying findings and handing off evidence
A pass means the agreed expected result was observed and evidence retained. Classify other results explicitly: confirmed failure (the system behaves incorrectly under valid conditions), blocked or test-environment problem (staging misconfiguration, stale data, flaky selector), inconclusive (behavior observed but not yet explained or reproduced), or not tested (a journey or variant not exercised this cycle). Never equate absence of findings with a pass; untested scope and inconclusive results represent remaining uncertainty that the product owner must weigh.
- Agreed journeys, expected outcomes, test roles, and synthetic dataset references.
- Environment URL, run time, and verified build identifier; flag unknown build identity.
- Results, reproduction steps, restricted evidence links, and an explicit untested-scope list.
- Known risks, their owners, and the follow-up decision or check needed.
Hand off evidence to authorized engineering and product reviewers. Traces and logs may contain sensitive data; restrict access and share only redacted, approved derivatives with marketing, investors, or external stakeholders. Do not distribute raw traces broadly. Reproduce findings using fresh synthetic fixtures and include request correlation references so reviewers can locate supporting logs without accessing unrelated data.
Further reading
- OWASP Authorization Cheat Sheet — general principles for server-side authorization design.
- Playwright best practices — general principles for reliable browser automation.