Agents build UI now. Nobody signs off on it.
The Interaction Conformance Standard fixes that. A one-page sign-off standard for AI-built interfaces. Version 0.1, July 2026. Adopt it as written.
Version 0.1, published July 2026
Agents now build interfaces, and nothing in most governance stacks checks their output before it ships. This standard defines the check: what conformance means, how it's verified, who signs off, and what happens on failure. It's written to survive a change control board and a quarterly budget review. The evidence behind every mechanism sits in the sources on the readback layer page. Take it, adapt it, argue with it. The flag is planted either way.
-
Purpose
Agents now produce interfaces: flows, states, behaviors, and screens. Existing AI governance covers prompts, data access, and tool boundaries; it does not check whether agent output conforms to the design system, meets accessibility law, or behaves like the product. This standard closes that gap.
-
Scope
Applies to all UI and interaction design produced wholly or partly by AI agents for governed surfaces: web, native apps, kiosks, agent desktops, and embedded screens, in every locale shipped. Human-built work continues through existing design review; where humans and agents co-produce, this standard applies to the output regardless of authorship split.
-
Definition
The readback layer is the set of criteria, machinery, sign-offs, and failure handling that verifies agent-built work against the design system before and after ship. It validates behavior, not just pixels.
-
Conformance criteria
Pass or fail, checked in this order.
F. Foundations. Only approved tokens; no hard-coded values; no colors, type styles, or spacing outside the system. Locale token layers respected (script-family type floors, density interactions, RTL mirroring).
C. Components. Real system components imported from system packages; no re-implementations or lookalikes; composition follows published component APIs; deprecated components blocked.
I. Interaction. Flows and states match sanctioned patterns: navigation, forms and validation, error and empty states, loading and feedback, motion. No invented interaction patterns without an approved proposal. This class is the layer's hardest test; machine checks cover pattern and state structure, human review covers judgment.
A. Accessibility. WCAG 2.2 AA; EN 301 549 mapping for EU surfaces; accessibility acceptance criteria present on the originating ticket.
L. Locale. Strings externalized with context notes; ICU syntax valid; expansion tolerances respected; blocking locales translated before merge; no locale forks of components.
-
Gates
V1, automated, every change: token lint, component provenance check (imports from system packages), accessibility engine, visual regression including tall-script, long-word, and RTL screens, string and ICU validation. Failures bounce without consuming human time.
V2, benchmark cadence: the agent context (MCP, rules files, orchestration files) is benchmarked on a schedule, tracking conformance rate and hallucination rate per model and per context format. Context changes are treated as API changes and require human approval.
V3, human sign-off, non-delegable: every merge to a governed surface; every deprecation and breaking change; every exception grant; locale sign-off in regulated communications by market compliance with an auditable approval chain; every rule, MCP, or context change. Sign-off happens on the readback.
V4, audit sampling: a monthly human audit of a sample of shipped agent-built work against all five criteria classes, including raw output behind clean readbacks, because the published record shows foundations drift even when component-level checks pass.
-
The readback
Every run of the readback layer produces a designer-readable artifact: the readback. (The name is aviation's: an instruction is read back before anyone acts on it. Same control, pointed at software.) A checker agent, separate from the builder agent and running on independent context, reads the built application against the system (token usage, component provenance, pattern and state structure) and renders a conformance redline: a stripped, monochrome scaffold of the application, annotated. Approved components carry their system names; lookalikes are flagged. Token usage is mapped, and rogue values and unapproved spot colors are rendered in color against the monochrome scaffold, so drift is visible at a glance. Flows and states are mapped to the sanctioned interaction patterns they claim to follow; unrecognized patterns are marked for human judgment. Open exceptions carry forward from the register.
Two rules keep the readback honest. First, the facts on it come from deterministic checks (lint output, import provenance, structural analysis); the checker agent assembles and narrates, it does not adjudicate. Second, the readback is evidence, not absolution: V3 sign-off happens on the readback, every red item resolves to block, fix, or exception, and V4 audits sample raw output behind clean readbacks to measure the readback's own false-negative rate.
The control shape is maker-checker: one agent builds, a different agent reads it back, a human signs. This is the four-eyes control regulated industries already run, applied to agent-built interaction design. The readback attaches to the change record, giving formal change control the evidence artifact the public record currently lacks.
-
Failure handling
Three outcomes, no fourth: block (CI or reviewer stops the change), fix (agent-drafted or human correction, re-gated), or exception (entered in the exception register with an owner and an expiry date; at review it is promoted into the system, migrated onto an existing component or pattern, or deprecated with a sunset date). Silent drift is treated as a governance incident, not a design preference.
-
Roles
The platform engineer owns V1 and V2 machinery, the checker agent, and the readback pipeline. The core design system lead owns exception grants and V4 audits. Designers review readbacks; the readback is the designer's V3 interface. The accessibility specialist owns class A. Locale owners own class L sign-offs in their markets. The council owns changes to this standard, pattern approvals in class I, and monthly review of the exception register. Auto-merge is permitted only for token-level fixes with passing tests on non-regulated surfaces, and is a council-level policy switch.
-
Metrics
Reported quarterly to the executive sponsor with dollar attribution: conformance rate at first gate; foundation drift rate in audit samples; readback false-negative rate (drift found in V4 samples behind clean readbacks); hallucination rate per benchmark run; mean time to detect drift; exception count and age; coverage as a governed ratio (target around 70 percent of new-feature design from the system, 50 percent for existing surfaces).
-
Cadence
V1 runs on every change. V2 monthly and on any context change. V3 continuous, on the readback. V4 monthly, including raw output behind clean readbacks. Standard reviewed quarterly; criteria classes re-baselined annually against current accessibility law and the token specification.
-
Maturity placement
Adopting this standard is the move from served context to enforced context; running V2 and V4 on cadence with published metrics is the move to an audited loop. It presumes the governance basics have shipped first: published decision rights, contribution criteria, and a deprecation policy. Enforcement without documented decision rights is theater.
Version 0.1, published July 2026. Download the one-page PDF, or ask the notebook about any clause.
Version 0.1, published July 2026