One dominant action
Can a first-time user identify the primary action within five seconds?
Evidence: A single visual priority, descriptive CTA copy, and no competing primary buttons.
Free evidence-based tool by HorizonX
Score an AI-assisted interface against 15 observable quality gates. No taste contests or framework bias—just the details that separate a convincing prototype from a product people can trust.
The standard
A vibe-coded interface is production-ready when its important flows remain clear, complete, responsive, accessible, trustworthy, consistent, and stable under realistic conditions. Visual refinement matters, but it cannot compensate for missing states or untested behavior.
Interactive scorecard
Choose the strongest answer supported by evidence. “Looks fine” is not evidence.
Can a first-time user identify the primary action within five seconds?
Evidence: A single visual priority, descriptive CTA copy, and no competing primary buttons.
Do headings, spacing, and grouping explain the page before the body copy is read?
Evidence: A consistent heading scale, short sections, and related controls grouped together.
Do labels describe outcomes instead of relying on vague words such as Submit or Continue?
Evidence: Action-led labels, useful empty states, and errors that explain recovery.
Are hover, focus, active, disabled, loading, empty, error, and success states designed?
Evidence: A state inventory or component examples covering every user-visible transition.
Can users understand, confirm, and recover from destructive actions?
Evidence: Clear consequences, confirmation where risk is high, and undo or recovery when practical.
Does the layout adapt when content needs it rather than only at preset device widths?
Evidence: No clipped controls, unreadable cards, or horizontal scrolling from 320px upward.
Are interactive targets large enough and separated enough for touch input?
Evidence: Targets near 44×44px, visible pressed states, and no hover-only functionality.
Does the interface survive long names, translated copy, empty data, and large values?
Evidence: Tests with 2× text length, realistic data, zero states, and documented overflow behavior.
Can every task be completed with a keyboard in a logical focus order?
Evidence: Visible focus, no keyboard traps, semantic controls, and predictable focus after dialogs.
Are status, errors, and selections understandable without relying on color alone?
Evidence: Text or icon reinforcement plus sufficient text and control contrast.
Do controls and dynamic regions expose useful names, roles, and updates?
Evidence: Native HTML first, associated labels, useful alt text, and announced async feedback.
Does the interface explain what happens to user data before sensitive actions?
Evidence: Plain-language privacy cues, permission context, and no surprise collection.
Are limitations, pricing boundaries, and unavailable features presented honestly?
Evidence: No fake urgency, hidden costs, fabricated activity, or misleading disabled controls.
Do spacing, typography, color, radius, and elevation follow a repeatable system?
Evidence: Named tokens or documented scales rather than one-off values across the interface.
Has the real interface been checked for performance, errors, and layout stability?
Evidence: Real-device checks, error monitoring, stable layout, and measured—not assumed—performance.
Methodology
Each criterion has equal weight. Surface polish cannot hide foundational gaps.
Review the real interface, not only its design file or ideal happy path.
Use screenshots, tests, recordings, component examples, or measured runtime data.
Assign 0, 1, or 2. Reserve 2 for criteria that have been demonstrated.
Fix the weakest group first and repeat the review after the interface changes.
Important: This checklist is a practical design and frontend quality gate. It does not replace product research, security testing, legal review, or performance measurement for your system.
Questions
It means there is observable evidence for clarity, complete states, responsive behavior, accessibility, trust, system consistency, and runtime quality—not merely an attractive first screen.
Use 0 when missing, 1 when partially implemented or unverified, and 2 only when evidence confirms the criterion works in the real interface.
No. User testing, security review, performance measurement, and product-specific validation still apply. The checklist is a quality gate, not a launch guarantee.