---
name: "standard-design-system"
type: standard
license: CC BY-NC-SA 4.0
description: The quality bar a design system must clear — token architecture, naming and three-way name alignment, component hierarchy, binding discipline, coverage, documentation, accessibility (WCAG 2.2 AA, stated as run checks), and versioning/change discipline — plus the component config record (component.config/v1), the single per-part source of truth Figma and code are both built from. Read before creating, extending, or assessing any venture's or client's design system (skill-brand-system, skill-design-tokens) so quality is judged the same way every run and an existing system can be scored rather than merely used. Sits beside convention-storybook (the code-side instance shape and the three build modes) and convention-figma (the Figma-side file organisation and structural bar). Surfaced by the 2026-07-10 Pixel capture; naming and accessibility added 2026-08-10 (APP-615); binding discipline and the component config 2026-09-04 (APP-804); grown as it is reused.
---

# Design-system standard

What a good design system must contain, stated as a bar the work can pass or fail. This is a *standard*, not a convention: each requirement below is derived, and a system can be scored against it. Owned by Pixel; read by `skill-brand-system` and `skill-design-tokens` before creating or extending any system, and by the read mode in `skill-brand-system` when scoring an existing one. `convention-storybook` holds the code-side instance shape this standard's outputs land in; Sonar's Design System audit (`convention-audit`) is the strategy-level view — this is the craft-level bar beneath it. `convention-figma` holds the Figma-side counterpart — how a design-system file is organised and the structural bar a Figma build is reviewed against. Which side is the reference at any moment — code, the design, or Figma at a refinement pass — is the scope's declared build mode (`convention-storybook`, *Three modes*); this bar holds in every mode.

## The bar — what a complete system contains

* **Token architecture.** Tokens are layered — primitive (raw values) → semantic (purpose-named: `surface`, `text-muted`) → component (scoped where needed) — with a consistent naming scheme stated once and followed everywhere. Theming (dark mode, per-venture variants) is expressed as alternate values on the same semantic tokens, never as a parallel token set. **The rule binds a single-file artefact as much as a system with a library behind it** — one HTML file that ships a hand-rolled set alongside a component library's own (a bespoke `--ink` / `--line` / `--surface` trio beside shadcn's `--sh-*`) carries two parallel sets by construction, and the failure surfaces as a dark toggle that re-values one and leaves the other: half the surface flips, half does not, and neither set is wrong on its own terms. Where an artefact adopts a library's tokens its semantic layer *is* that library's; where it adopts none it mints one set and themes by re-valuing it. Never both. A flat bag of hex values fails this bar.
* **Naming.** One name identifies a part across every surface it appears on — the Figma component name, the Storybook title path, and the code path resolve to the same string — so the design↔code bridge is a lookup rather than a translation. The name states what the part is and where it sits; the layer is carried by the path, not restated as a prefix. A system where the same row is `WorkListRow` in code, `Work / Row` in Figma and `Sections/Work List` in Storybook fails this bar even though each surface is internally consistent.
* **Component hierarchy.** The set is organised in layers — tokens → assets → components → sections → templates (the same hierarchy `convention-storybook` renders) — with each component built from the layer beneath it, not from raw values. A component carrying hardcoded values that bypass the tokens fails. (An illustrative mark is the exception — *The asset record*.)
* **Binding discipline.** Three checkable rules the hierarchy bar implies, stated because each was failed in one shipped part (a1.meirionpritchard.ds, modifier-cell, 2026-08-11 — APP-804). (a) **Documentation names the file's real tokens** — a docs page or config that cites a token cites one that exists in the system, by its real name, never an invented placeholder. (b) **Height derives inside-out** — a row's or cell's height comes from its content plus its padding tokens, never from a fixed-height override; a fixed height that must be kept in step with the padding is two sources of truth. (c) **A part binds existing semantic tokens** — it reaches for the semantic layer the system already has rather than minting a token per part; a new semantic token is added only where no existing one carries the meaning, and is recorded as a system change (*Versioning*). A part shipped with its values un-bound to any token fails all three. Illustrative marks are the one class exempt from this rule and from the hierarchy bar's token-binding clause (*The asset record*, and `convention-figma`, *The structural bar*).
* **Coverage.** The system covers what the venture actually ships: every recurring UI element has a systemised part, and one-off exceptions are named as exceptions. Coverage is stated (what is in, what is deliberately out) rather than implied. A system that documents ten components while the product uses forty fails.
* **Documentation.** Every part carries its usage: when to use it, its props/variants/states, and its don'ts — readable by a human and consumable by a machine (the dual-readable rule in `convention-storybook`). An undocumented part is inventory, not system.
* **Accessibility.** Every part meets **WCAG 2.2 AA**, checked rather than asserted, and carries its accessibility contract alongside its visual states — keyboard model, focus behaviour, accessible name, semantics. A contract that names states but not focus, contrast, motion or semantics is half a contract.
* **Versioning and change discipline.** Changes to the system are deliberate: proposed, recorded (what changed and why), and propagated — never silently restyled in place. Where the system lives in code, the token source is single and changes flow from it (see `skill-design-tokens`); where it lives in Figma, the library is published and consumers update from it.

## Token reuse — by meaning, never by value

*Binding discipline* (c) says a part binds an existing semantic token rather than minting its own. Three rules extend it to the case it does not cover: the token that already carries the meaning is there, and nobody is using it. Each was failed on one pass (a1.meirionpritchard.ds nav tokens, 2026-09-08 — APP-1181).

* **Never delete an orphan as cleanup.** A token whose components have moved on is re-valued into the current set — re-valued to what the work needs, renamed where the meaning has narrowed or split, and re-pointed at — or left in place with a recorded reason. Delete-then-create looks identical in the file and is not: the variable's identity goes with it, so every existing binding breaks and a component that referenced the old token falls silently back to a raw value instead of following the new one. Re-valuing keeps both the identity and the record that the value was once considered, and reaches the same end state.
* **Match on meaning, never on value.** The test is whether a token *means* the thing being bound to it, not whether its current value happens to match. Binding a nav to `space/hero-y` because both hold 20 couples two unrelated parts, and a later change to the hero bar's padding then moves the nav silently. The right reuse was `space/nav-y` — an orphaned nav token, re-valued for nav work.
* **Never collapse two semantic tokens because they resolve to the same primitive.** Two semantic tokens holding one primitive is the tier working as intended, not duplication: two parts can each carry their own semantic token even where both currently resolve to the same primitive, so one part's token can be re-pointed later without dragging the other with it. An audit that deduplicates by value has broken the system rather than tidied it.

The last rule binds the audit as much as the build. *Find the duplicate values* is the obvious way to write a token sweep and it is precisely wrong here: a sweep reports same-primitive pairs as observations, never as findings, and an orphan as a candidate for re-valuing, never for deletion.

## Typography — the ramp as a checkable bar

`skill-brand-system` step 2 already sends a venture's **type scale** here to be scored — "palette, type scale, logo usage, spacing … — to the bar in `standard-design-system`" — and until now this standard stated no bar for it, so a ramp could be built, overridden and shipped with nothing to check it against. These are the properties a ramp is scored on, each **pass / partial / fail** like every other bar above. (Surfaced by one audit — mbn.type-ramp-review.pass-1 on the Meirion storefront, 2026-10-03 — whose findings are **one** gap and not six: the bar was missing, not six bars. APP-1999, folding APP-2000, APP-2002 and APP-2003.)

* **A stated ratio, and ordinality that is checked.** The steps between roles come from **one stated ratio** the system names once — a modular scale (1.2, 1.25, 1.333) or an explicitly stated bespoke set — never from a value chosen per role. And the numbered heading roles are **mechanically checked to descend: h1 ≥ h2 ≥ … ≥ h6 at every breakpoint**, as *rendered* size, so an override cannot silently invert the hierarchy. This is **not** check 5's "heading levels descend without skipping", which governs the *level sequence in the markup*; this governs the *rendered sizes those levels resolve to*, and the two are independently wrong — a document can have faultless heading levels and an inverted ramp. A ramp with no stated ratio fails; a ramp whose rendered order inverts fails whatever its markup does. (Observed: h1 at 15px rendering smaller than h6 at 16px desktop and the same size as body copy, with no ratio governing 56→50→40 or the cliff to 15→13→11→10→12→15/16 — so nothing could have caught the inversion before an audit caught it by eye.)
* **Tracking is relative, never a fixed px offset.** Letter-spacing is expressed in **em or %**, because tracking reads as a proportion of size: one flat px value is imperceptible at the top of a scale and loose at the bottom — a +0.6px offset is ≈0.8% of a 72px display role and ≈4% of a 15px body role, from the same declaration. The system states **family-level defaults** — zero-to-negative at display sizes, near-zero through body and the numbered headings, positive only on small-caps and label-style roles (badges, captions, buttons) where it is doing real legibility work — and a role departing from its family's default carries a reason. A ramp carrying one fixed-px value across the whole scale fails. (APP-2000.)
* **A derived type value is recomputed when the value it derives from changes.** This is *Binding discipline* (b) — "height derives inside-out … a fixed height that must be kept in step with the padding is two sources of truth" — reaching the type ramp: a **line-height fixed in px that must be kept in step with a font-size is the same two sources of truth**. So whenever a font-size override is applied to an existing role, its line-height is **recomputed as a deliberate ratio of the new size**, never left at the value the old size produced. The tell is a run of oddly precise, non-round unitless ratios across adjacent roles (1.267 / 1.385 / 1.455 / 1.6 on h1–h4), which is what a stale fixed-px line-height divided by a shrunken font-size looks like. Treat that pattern as a prompt to read the declaration, not as a finding on its own: on the one observed instance it was scored **assumed** — inferred from the pattern and the baseline's own provenance note, not confirmed from raw CSS — and the rule stands on (b) rather than on that instance. (APP-2002.)
* **Responsive behaviour is stated per role family, not emergent.** The system states, **for each role family, whether it scales by breakpoint or stays fixed, and why** — expressed through the semantic modes the component config's **responsive** section already requires, never as a per-role device variant. An unstated pattern is not a decision, and the tell that it never was one is an inconsistency *inside* a single family. (Observed: h1–h4, `.button`, `.badge`, `.caption-with-letter-spacing`, `.form__label` and the table-cell role fixed across breakpoints while the display roles, h5, h6, body, price and `.caption` scale — h1–h4 fixed and h5–h6 scaling inside the same heading family. APP-2003.)
* **An interactive or UI text role is measured at its box, not at its font-size.** A button, form control or badge is checked against *Accessibility* check 3 (WCAG 2.2, 2.5.8) with its **padding and box dimensions read alongside its font-size in the same pass**, so target size is answerable from the ramp read rather than deferred to a component pass that may never happen. Where such a role renders **below the system's own body size**, that is recorded as a named exception carrying a reason (*Where a bar cannot be met*), never left as an unremarked value. (Observed: `.button` at 10/10px. **A stated minimum px floor for these roles is deliberately not set here** — a floor is a new bar level that neither WCAG nor anything else in this standard derives, so it is an Owner call, open on APP-2001. Until it lands, this clause is the check and the named exception is the remedy.)

**Scoring.** Score the ramp **per property**, as *Accessibility* is scored per check, so a ramp passing ordinality and failing tracking reads as two findings rather than one blurred verdict. Where the only available surface is a computed-styles baseline, the ratio and the controlling unit may not be readable from it at all — tag those **unknown**, never **fail** (*Scoring an existing system*).

*Provenance.* The storefront figures quoted above are mbn.type-ramp-review.pass-1's own readings and are **unverified** from the refine surface, which can read no repo and made no live-page read; that flag covers those quoted px values only. The five rules are derived from this standard's own *Binding discipline* (b), the component config's **responsive** section, *Accessibility* checks 3 and 5, and WCAG — none of which turns on those readings.

## The component config — one record per part (`component.config/v1`)

Every part carries one config record, and it is the **single source of truth Figma and code are both built from** — not a document written after the fact from either. The schema, `component.config/v1`, has ten sections; a record that omits one names the gap rather than leaving it implicit. Seven describe the part; three — **binding**, **measurement** and **provenance** — exist so that a code build can reach parity from the record alone and a reviewer can check parity by measuring rather than by reading.

| Section | What it holds |
| -- | -- |
| **meta** | The part's one name (the three-way-aligned string, *Naming*), its layer (component, section, template), its status, and the family it belongs to. |
| **tokens** | The primitive → semantic → component chain the part consumes, by real token name; any component-scoped token it introduces, and why. |
| **anatomy** | The layer tree — and the same tree is the Figma layer structure and the code's slot structure. One tree, three surfaces. |
| **api** | Variants, props and states with their allowed values and defaults; what is a variant in Figma is a prop in code, under the same name. |
| **responsive** | How the part behaves across breakpoints, expressed through the system's semantic modes rather than a per-part device variant. |
| **a11y** | The accessibility contract (*Accessibility*, check 6): role, accessible name, keyboard model, focus behaviour, state announcements. |
| **docs** | The dual-readable usage — when to use, when not to, don'ts — that renders as the family's Figma doc frame and the Storybook autodocs alike. |
| **binding** | Per token row and per layer, the **code** side beside the Figma side: the code token path or CSS variable, and a flag wherever code uses a literal where Figma binds a variable. |
| **measurement** | What each number is measured on — Figma frame, CSS border-box, rule-to-rule pitch — the expected value per mode, which viewport width selects which mode, and, for type, the line box and trim mechanism per surface. |
| **provenance** | The Figma node ids, the file version and the code commit the record was last reconciled against. |

**Every section states both surfaces, and a divergence is written down rather than discovered.** The seven sections above were written from the Figma side and read as though code follows by construction. Where it does not, the record is not merely incomplete — it is wrong in the one direction a builder will act on, stating a number that is true in Figma and false if applied to code. So each existing section carries the code side explicitly. **anatomy** names, per layer, the DOM element or class, the slot-visibility rule, and the sizing shape: fill or fixed, minimum width, wrap or truncate. **api** names, per Figma property, the code prop and the value map, with "no code equivalent" stated rather than left blank, and, per state, how code implements it and the story that proves it (`convention-storybook`). **responsive** names, per mode axis, what the mode means in Figma, what it means in code, and what the part does under it. And wherever the two surfaces genuinely differ — a deliberate mechanism change, a value bound on one side and literal on the other, a mode that touches one surface and not the other — the record **states the divergence and the reason for it**, because *Naming*'s rule that a Figma variant is a code prop under the same name is the default, and an exception nobody wrote down reads as parity. The test for the three added sections is the same as for the seven: a record that omits one names the gap. (Worked case: A1-604, the cell, parity pass 2026-10-05 — Figma values read from the file, Storybook rendered and measured at 1280, 1024, 640 and 600 wide. Height, padding, gap, border, selector, handle and label size all agreed; every difference was something the record did not state, so each had to be found by reading CSS and measuring. The record lists label line height 30/22 where code sets it to the selector size, 20/16, as a deliberate mechanism change, so a build from the record alone produces a 60px row rather than 50. Heights are given as 50 and 38 — Figma frame heights with a centre-aligned 2px stroke — against a browser border-box of 52 and 40, and the record does not say which is measured. Token names differ per surface (`space/cell-y` against `--cell-pad-y`, `theme/content` against `--ink`, `type/cell-size` against `--fs-nav`), with two code values literal where Figma binds a variable. Five Figma properties map to differently named code props or to none. And dark mode means two things: a Figma theme mode that flips the cell, against code where `body.dark` never touches it. The trim half of **measurement** has a second, independent instance on the Figma side — a cap-height-trimmed style cropped by a frame's Clip content, APP-2069, folded in `convention-figma` and carried here as evidence. Premises unverified — no repo access from this refine surface: the Storybook measurements at `origin/main` 5b0c2b8, the Figma file and its node values, the token names and every px figure above are the filing leg's report, and whether Figma's text-box trim explains the 18px label box at 25px size was not read directly; that flag covers those figures only, not the rule, which rests on this section's own single-source-of-truth statement and on the Owner's 2026-10-05 ask. APP-2070.)

The record is written when a part is added or restructured, and the Figma build and the code build are each checked against it rather than against each other. Worked examples exist for `alert` and `modifier-cell` (a1.meirionpritchard.ds, 2026-08-11). The working form — YAML, one file per part, kept beside the part in the repo where the system lives in code, and mirrored on the family's `doc/config` frame in Figma (`convention-figma`) — is not yet scored; the presence and completeness of the record is. `skill-design-tokens` maintains the records; `skill-brand-system` read mode scores whether they exist.

## The asset record — one record per mark

An asset (an icon, a mark, a spot) is a part with no behaviour, so most of `component.config/v1` has nothing to state for it (api, responsive, binding), and what matters for a mark is not in that record. An asset carries an **asset record** instead, mirrored on the assets page's `doc/config` frame (`convention-figma`, *File organisation*). It states:

* **source** — the exported SVG as the source of truth, and where it lives; "not yet exported" is written rather than left blank.
* **size and stroke** — recorded as values (the main's size, stroke weight, corner radius), each marked bound to a variable or exempt.
* **exemption** — for an illustrative mark, the token-binding exemption (`convention-figma`, *The structural bar*), written on the record. An unbound value with no exemption stated is a defect; one with it is a finding of none.
* **motion** — the animation, where the mark has one, with its reduced-motion behaviour (*Accessibility*, check 4).
* **accessibility** — decorative or labelled, and the accessible name where labelled.
* **usage** — where each mark is used.

A record that omits a field names the gap, as the component record does, and presence and completeness are what is scored. **Illustrative marks are exempt from token binding** (Owner ruling, Aled, 2026-10-09 — APP-2204): bespoke artwork, its values are optical tuning, and the check is that it renders identically to its exported SVG. UI-structural icons (`assets/icon/*`: close, drag-handle, selector) are not marks in this sense. They carry an asset record too, and bind tokens. The value-level question — a register of deliberate literals and derived values across the system — is not settled here and is carried by the open token-section and token-convention proposals (APP-2087, APP-2088). (Owner, Aled, 2026-10-09: the documentation for the assets pages "might need to be different from the component pages as they have different requirements" — APP-2205. The six fields are the filing leg's session observation, adopted here as the drain's wording, and they are not a named schema version. Premise unverified: that the nav marks are consumed in code as exported SVGs and are an animated set rests on the titles of A1-124 and A1-204; no repo or Figma file was read from the refine surface. That flag covers those two points only.)

## Naming — the bar, and the part that is only a convention

Two different things travel under "naming", and only one of them is a standard. Keeping them apart is what stops an arbitrary lexical choice being scored as a quality failure.

**The bar (a standard — one choice is better).** Three-way resolvability, as stated above: one string, three surfaces, derivable without a mapping table. It is a standard because misalignment has a cost that is not a matter of preference — every bind between Figma, Storybook and code becomes a translation, and translation is where the bridge silently rots. It is testable: take ten parts at random and check that the Figma name, the Storybook title path and the code path are the same string. Any part that needs a lookup table fails.

Two further bar-level properties follow from it:

* **The scheme is stated once and covers tokens and components alike** — one scheme per system, written down in the system's own documentation, not inferred from the existing entries.
* **The system's name wins over a build's name.** Where a build ships a part under a different name, the system's name is authoritative and the rename is scoped into the build's work — never absorbed as a permanent alias, which reintroduces the mapping table by another route.

**Only a convention (arbitrary-but-shared).** The lexical form itself — slash- or dot-delimited, lowercase or PascalCase, singular or plural nouns — is a choice where two sensible people could each decide differently and only agreement matters. By the test in `core-operating-model` that makes it a convention, not a standard, so it is **not scored** here. It is written here because the estate reads design-system naming from this artefact and from nowhere else; when the form hardens it belongs in a convention, and moving it there is a later, deliberate step rather than a reason to leave it homeless now.

**The current scheme, held as a trial instance.** The Meirion Work-section blueprint fixed: slash paths, lowercase, singular nouns, the layer carried by the path (`components/project-row`, `sections/work-list`), no layer prefixes. That is the working scheme for that trial and not yet general canon — the blueprint-first model it belongs to is itself under trial, with the verdict due on the A1-231 epic (APP-613). Score a system against the alignment bar; do not score it against these lexical choices until the trial returns a verdict.

## Accessibility — the bar, stated as checks

The named level is **WCAG 2.2 AA**, applied to every part of the system and to every theme it ships: light, dark and per-venture variants are each a separate pass, because a pair that passes on one ground can fail on another. AA is the bar, not the aspiration — a part below it is not finished. `standard-prioritisation` already treats an accessibility failure as **Critical**; this section is the definition that severity has been presupposing.

Seven checks. Each is run and recorded, never asserted.

1. **Contrast, checked at the token pair.** Body text and text-like elements at least 4.5:1 against their ground; large text (24px and above, or 18.66px and above when bold) at least 3:1; non-text essentials — component boundaries, form-field borders, focus indicators, and icons or chart marks that carry meaning — at least 3:1 against what sits next to them. Check the **semantic token pairs** (`text-muted` on `surface-raised`, and so on) rather than screenshots: the pair is the unit that recurs, so checking it once covers every component that uses it, and a new component inherits the result.
2. **Focus visibility and order.** Every interactive element is reachable by keyboard, in an order that matches the visual reading order. The focus indicator is visible, holds at least 3:1 against what surrounds it, and is never removed without a replacement. Focus is not trapped except deliberately inside a modal, which returns focus to its trigger on close.
3. **Target size.** Pointer targets are at least 24 × 24 CSS px with enough spacing that neighbours do not overlap (WCAG 2.2, 2.5.8), or an equivalent-size alternative exists on the same surface. 44 × 44 is the preference for a primary touch target, not the bar. **For a text-only target the font-size is not the measurement** — the target is the box, so read padding and box dimensions **alongside** font-size in the same pass rather than inferring one from the other; a type-ramp read that reports sizes without boxes has not checked this criterion and says so (*Typography — the ramp as a checkable bar*). (APP-2004.)
4. **Motion preference.** Every animation, transition, parallax and autoplay honours `prefers-reduced-motion: reduce`, and the reduced branch is a real branch — the motion is removed or replaced with a plain fade — not the same motion with a shorter duration. The branch lives in the motion/token layer so it applies system-wide rather than being re-implemented per component. Name the mechanism and check it actually runs: a reduced-motion claim offered as a rationale with no mechanism behind it is how this check is failed while reading as passed (worked case: APP-726).
5. **Semantic structure.** Native semantics first — a button is a `button`, heading levels descend without skipping, every control has a programmatically associated label, landmarks appear once each, and images carry alt text or are explicitly marked decorative. ARIA is used only where no native element carries the meaning, and never to relabel a native element that is already correct.
6. **The documented contract.** Each part's documentation states its keyboard model (which keys do what), its focus behaviour, its accessible name, and its role and state contract — in the same dual-readable place as its props (`convention-storybook`), so both a human and the machine layer read it.
7. **Text that survives being resized and re-spaced.** Two AA criteria bind the **type ramp** specifically rather than any one component, and neither is answerable from a table of computed values — both are **live tests**, which is why a ramp read reporting only computed px skips them silently and reads as passing. **1.4.4 Resize Text:** text reaches **200%** without loss of content or function, which turns on the **controlling unit in the source** — a computed-px baseline reports `15px` whether the declaration was `15px`, `0.9375rem` or a clamp, and only some of those resize. So read the **declared** unit, not the computed value, and confirm the enlargement in a browser. **1.4.12 Text Spacing:** with line-height forced to **1.5× font-size**, paragraph spacing to **2× font-size**, letter-spacing to **0.12em** and word-spacing to **0.16em**, no content or function is lost — applied as a **DOM override on the running page**, never reasoned from the static values, because what fails here is overflow and clipping and the values do not carry it. Record both as run, naming the unit read and the override applied; a ramp whose criteria were inferred from a computed-styles baseline alone is scored **unknown** on these two, never **pass** (*Scoring an existing system*). (Surfaced by mbn.type-ramp-review.pass-1, 2026-10-03, where both came back unresolvable from the baseline that pass had — APP-2004. That provenance claim is the audit's own and is **unverified** here; the criteria themselves are published spec and the omission was verified by reading all six checks above.)

**How the checks are run.** Automated first: an axe pass over the story set, which Chromatic already runs beside its visual snapshots on a wired repo (`convention-github`). Then a manual keyboard pass per part — tab to it, operate it, leave it. Automation is necessary and not sufficient: it does not see a wrong focus order, a valid-but-wrong heading level, or an accessible name that reads as nonsense. A part is scored on both, and an automated-only pass is scored **partial**.

**Measuring contrast in a running browser — the six traps, and why they are silent.** Where the contrast check is met by measurement rather than by a token-pair table — a themed console, a hovered state, any pairing that only exists at runtime — the measurement is written against a live page, and the obvious implementation is wrong on a modern token theme in ways that produce specific, plausible, false numbers. The direction of the error is what makes this a rule rather than a tip: each trap is as capable of a false pass as a false failure, and a PR body carries a table of numbers either way, so nothing downstream can tell them apart. Six traps, all observed, all in the same sixty lines:

1. **Resolve every colour through a canvas, never a regex.** A utility such as `bg-primary/90` compiles to `color-mix(in oklab, var(--color-primary) 90%, transparent)`, and `getComputedStyle` returns that string verbatim; a parser matching `rgb(` / `rgba(` matches nothing, returns "no colour", and silently measures the element *behind* the one under test. Set `ctx.fillStyle` on a 1×1 canvas and read the pixel back — the browser resolves any CSS colour syntax.
2. **Detect a refused parse.** After `ctx.fillStyle = value`, confirm the browser did not fall back to `#000000` on a value it could not read, and report a refusal as an error rather than measuring it. Without that check the failure is the same silent one — a plausible number for a colour nobody resolved.
3. **Start the ground walk at the element itself when the element paints its own background.** Walking up from `parentElement` unconditionally measures a filled button's label against the surface behind the button.
4. **Plant the probe inside the element the tokens are scoped to.** Where a theme scopes its token layer to a container rather than `:root`, a probe appended to `body` reads a different set of values under the same names — and reads them as a legitimate, passing pairing.
5. **Choose the negative control, and prove it fails.** A control is only a control if it is known to fail and observed failing. Two ways it silently does nothing: it draws a token that is not the kind of token you assumed (a surface where you wanted text), or it is injected as a *class* the source scan has never seen, so the class compiles to nothing and the probe inherits the foreground. Write it as an inline token reference, and **assert the control's two theme readings differ** — a control drawn from the theme's own tokens cannot measure the same in both themes, so identical figures mean the probe is reading a fixed ground rather than the theme.
6. **Wait out the transition after a theme switch.** Where the theme is a class on the root element and the surface carries `transition-colors`, a probe that toggles the class itself is not covered by the theme library's transition suppression, and sampling immediately reads mid-interpolation colours. Set the motion duration to zero for the probe, or settle past it before sampling.

Where a repo meets this bar on more than one ticket, the measurement is **committed to that repo as a script and run**, not re-implemented per change: a fix to the maths then fixes it for every future run, which re-implementation cannot. (Worked case: `assoc-one/firmup-operator`, 2026-09-08 — six console leaves against one contrast bar, four legs each hand-building the harness, and four separate rounds lost to correcting the instrument before the measurement could be believed. A1-385 reported a hovered button at 1.22:1 dark / 1.14:1 light against a true 4.71 / 6.41; A1-386 round 2 reported a negative control passing at 6.86:1 in both themes; A1-387 reported light-mode muted text at 2.19:1 against a true 5.16:1. Every one of those was the instrument, not the build — APP-1195.)

**Where a bar cannot be met**, record it as a named exception carrying a reason and a ticket. An unmet criterion that is written down is a finding; one that is not is a defect the system will keep reproducing into every part built after it.

**Scope.** This is the design-system half — the parts and the contracts they carry. Accessibility at the product level (whole journeys, content and alt-text authoring, service copy) crosses product design, build, review and content, has no canon home, and is not closed by this section.

## Scoring an existing system

When ingesting a system Pixel did not create (the read mode in `skill-brand-system`), score each bar above **pass / partial / fail**, and tag each finding **known** (observed directly), **assumed** (inferred), or **unknown** (not inspectable from the available surface). Two bars score at a finer grain: **naming** is scored on alignment and on the two properties that follow from it, never on the lexical form; **accessibility** is scored per check, so a system passing contrast and failing focus reads as two findings rather than one blurred verdict. The scored inventory doubles as the craft-level input Sonar's Design System audit consumes — produce it once, let the strategy layer read it.

## Guardrails

* The bar judges the system, not the taste — aesthetic judgement lives in `convention-aesthetic` and the venture's `convention-brand` instance.
* Score what is observable; name what is not (unknown), never fill it.
* Score naming on alignment, not on the lexical scheme — and treat the Meirion scheme as a trial instance until its verdict lands.
* Accessibility is a bar, not a preference, and is not traded against an aesthetic choice: where the two collide, the palette or the treatment moves.
* A failed bar is a finding to raise, not a licence to unilaterally restructure a client's or venture's live system.
