Evidence sample: 11 deep components across 10 archetypes; 1 flow(s); 2 pattern(s); intentional design dispositions data-table=compound, form-control=compound, page-layout=code-only; governance scope public.
Public-evidence scope: Missing public evidence remains unknown.
How to read the score: Scores show how closely the captured evidence meets each rubric anchor. A criterion score of 3 maps to 75 and means the configured anchor is comprehensive, operational, or tested as defined by that criterion.
Foundationscontinue, 100% reviewed75
Bindingspartial, 100% reviewed68.75
Componentscontinue, 100% reviewed75
Structurecontinue, 100% reviewed75
Governancecontinue, 100% reviewed75
Diagnostic area
Foundations findings
Token taxonomy and semanticsPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
The Primer Web Figma file was captured at full-file variable scope (880 variables, 8 collections, 1,507 alias references), and the taxonomy_prefixes breakdown shows a genuine multi-tier alias structure: base primitives (base/size, base/typography, base/text/weight) feed functional/pattern-layer semantic tokens (button/*, control/*, text/*, bgColor/*, borderColor/*, fgColor/*, spinner/*, avatar/*). The primer-primitives source repo documents the alias mechanism directly (a token's $value can reference another token, e.g. '{base.color.blue.3}') and enforces token-name validity via an automated lint (lint:tokens). This is comprehensive and operational, not merely documented-basic, but the packet contains no automated cross-consumer validation or measurement of the graph itself (only source-repo-level lint/build/contrast checks), so anchor 4 is not met. Note a disclosed conflict: primer-figma-guidance's prose states only color and size tokens are supported as Figma variables ('we still provide text and shadow tokens using styles'), but the observed full-file capture shows live 'typography' (34 vars) and 'base/typography' (4 vars) variable collections: per evidence precedence the observed implementation is preferred over this stale claim.
Selected anchor: Semantic layers and aliases are comprehensive and operational
Next anchor: The token graph is validated, measured, and maintained across consumers
Evidence needed: Anchor 4 requires the token graph to be validated, measured, and maintained across consumers; the packet shows only primitives-source-level checks (lint:tokens, a11y-contrast CI) with no automated validation or measurement of token-graph correctness across the React/CSS/ViewComponents/Figma consumer matrix.
Scope: foundation; system-wide. Full-file coverage supports system-wide claims about the token graph's existence and structure; it does not by itself prove every individual token is bug-free or that downstream consumers implement the graph identically.
- What broke
- No confirmed break in the taxonomy itself, but current documentation (primer-figma-guidance) understates current capability: it claims typography still relies on styles-only, while the observed full-file Figma capture shows live typography variable collections.
- Impact
- Consumers relying on the published guidance page may wrongly assume typography tokens are not available as Figma variables and continue using styles, missing the more consistent variable-based workflow the current library actually supports.
- Why
- primer-figma-guidance (content.txt lines 387-390) is normative documentation that has not been updated to reflect the current 'typography' and 'base/typography' variable collections visible in the full-file Figma export.
- What to change
- Update the Figma guidance page to reflect that typography is now (partially) variable-backed, and clarify the current scope of styles-vs-variables usage.
Evidence citations
dfc83d5c91c7186f6315/figma-summary.json, original source: lines 303-305dfc83d5c91c7186f6315/figma-variable-summary.json, original source: lines 190-1995aeca3d83424d7cde1d4/files/README.md, original source: lines 83-1175aeca3d83424d7cde1d4/files/package.json, original source: line 45f478be6e00ccc7aec688/content.txt, original source: lines 387-390dfc83d5c91c7186f6315/figma-variable-summary.json, original source: lines 122-136
Primitive completenessPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Primitive breadth is comprehensive: 880 resolved variables span color (599), float/size (274), and string (7) types across scopes covering fills/strokes (663+124), text (167), gap (51), corner radius (6), typography metrics (38), and effects (132). The @primer/primitives npm package ships this as compiled, documented CSS for size, typography, borders, breakpoints, viewport, motion, and spacing. Consumption is directly evidenced rather than inferred from naming: 10 of 11 sampled deep components (Button, TextInput, ToggleSwitch, Dialog, ActionMenu, InlineMessage, FormControl family, Blankslate, Spinner, DataTable family) resolve with dozens of shared bound_variable_ids each (for example, Button binds roughly 90, Dialog roughly 37), showing the same primitive set consumed across distinct archetypes (action, input, overlay, navigation, feedback, content, status, data-display) without manual overrides. This meets 'comprehensive and consistently consumed' but adoption/coverage measurement across the full ~2,103-component library is not evidenced in the packet, so anchor 4 is not reached.
Selected anchor: Primitives are comprehensive and consistently consumed
Next anchor: Primitive coverage and adoption are measured and governed
Evidence needed: Anchor 4 requires primitive coverage and adoption to be measured and governed (for example, adoption metrics across the full component library); the coverage profile explicitly excludes adoption dashboards from this public evaluation, and the packet only demonstrates consumption within the 10-11 sampled deep components.
Scope: mixed; sample-only. Primitive definition breadth (880 variables, full-file) is system-wide, but the 'consistently consumed' claim rests on direct binding evidence from only the 10-11 sampled deep components; the remaining ~2,092 components in the library are not directly verified consumers.
- What broke
- No confirmed break in primitive completeness for the sampled components.
- Impact
- None observed within the sample; broader adoption across the full 2,103-component library cannot be confirmed or denied from public evidence.
- Why
- The evidence packet caps deep-component sampling at 11 named components and explicitly excludes adoption dashboards from public evaluation scope.
- What to change
- No change needed for the sampled scope; a system-wide adoption claim would require either a cited system-wide consumption mechanism or dashboard-level evidence, which is out of scope here.
Evidence citations
5aeca3d83424d7cde1d4/files/README.md, original source: lines 30-56dfc83d5c91c7186f6315/figma-variable-summary.json, original source: lines 139-154dfc83d5c91c7186f6315/figma-sample-details.json, original source: lines 1108-1206dfc83d5c91c7186f6315/figma-sample-details.json, original source: lines 358-397
Modes and adaptationPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Color/theme modes are defined at the variable level, not as manual overrides: the Figma 'mode' collection carries 699 variables each with an explicit value for all 9 modes (light, dark, dark dimmed, light/dark high contrast, light/dark protanopia deuteranopia, light/dark tritanopia), and @primer/primitives ships matching versioned CSS theme files for the same set, confirming delivery to code consumers, not just design. Sampled deep components (for example, Button) bind directly to mode-collection variable IDs, so switching Figma's mode dropdown propagates systematically rather than through per-instance overrides. A second mode axis (density: condensed/normal/spacious) is a dedicated variable collection consumed by the DataTable family sample. A third axis (responsive viewport ranges narrow/regular/wide plus six numeric breakpoints) is documented normatively with concrete per-breakpoint padding values and named component-level responsive behaviors (split-into-pages, bottom-sheet, stack-vertically). This satisfies 'systematic across tokens and components' for the evidenced sample. Anchor 4 is not reached because no packet evidence shows continuous, automated validation of mode compatibility across the supported consumer matrix (only an a11y-contrast CI check on the primitives source itself).
Selected anchor: Modes are systematic across tokens and components
Next anchor: Mode compatibility is continuously validated across supported consumers
Evidence needed: Anchor 4 requires mode compatibility to be continuously validated across the supported consumer matrix; the only automated check evidenced is an a11y-contrast CI workflow scoped to the primitives source repo, not a cross-consumer (React/CSS/ViewComponents/Figma) mode-parity test.
Scope: mixed; sample-only. The mode/breakpoint token definitions themselves are full-file/system-wide, but direct evidence that components consume them systematically (via bound_variable_ids) is limited to the resolved deep-component sample; responsive/accessibility-preference modes beyond color and viewport are documented as intent only, not implementation-verified.
- What broke
- No confirmed break; responsive user-preference modes (prefers-color-scheme, prefers-reduced-motion, forced-colors, prefers-contrast, inverted-colors) are stated as requirements in prose ('GitHub must respect these preferences') but the packet does not contain direct implementation evidence for them, unlike the color/density/viewport modes.
- Impact
- Claims about accessibility-preference mode support beyond color/contrast should not be generalized past documented intent, since no binding or CI evidence for those specific media features is in the packet.
- Why
- primer-responsive-guidance documents these as design requirements rather than pointing to implementation or test artifacts.
- What to change
- No change required for the scored color/density/viewport modes; a future evaluation could request evidence (code or tests) demonstrating prefers-reduced-motion/forced-colors handling specifically.
Evidence citations
dfc83d5c91c7186f6315/figma-variable-summary.json, original source: lines 38-80dfc83d5c91c7186f6315/figma-variable-summary.json, original source: lines 81-995aeca3d83424d7cde1d4/files/README.md, original source: lines 46-5590e111436b688c05c92e/content.txt, original source: lines 391-431dfc83d5c91c7186f6315/figma-sample-details.json, original source: lines 1169-1198
Consumable delivery and versioningPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
@primer/primitives is published to npm under semantic version 11.10.0 and installed via a standard package manager command. A single build pipeline compiles the same source token data into multiple consumer-specific, versioned artifacts released together: CSS variables (build:tokens/build:fallbacks), a Figma-compatible export with explicit $extensions.org.primer.figma metadata (build:figma), TypeScript types (build:types), and an LLM-oriented artifact (build:llm). Releases are governed by changesets, and CONTRIBUTING.md documents that PRs automatically produce a canary build for pre-merge testing. CI enforces lint/format/test/build gates and a dedicated a11y-contrast workflow, both visible as status badges in the published README. This meets 'multiple consumers receive consistent versioned artifacts.' It falls short of anchor 4 because, while a check:removed-tokens script exists to flag token removals, the packet contains no measured migration-impact evidence (for example, quantified downstream breakage tied to a specific release) and no direct evidence of automated contract validation between the delivered artifact formats themselves.
Selected anchor: Multiple consumers receive consistent versioned artifacts
Next anchor: Delivery contracts are validated and migration impact is measured
Evidence needed: Anchor 4 requires delivery contracts to be validated and migration impact measured; the repo has a check:removed-tokens script that can flag removed tokens, but no measured migration-impact results (for example, quantified downstream consumer breakage per release) are present in the packet.
Scope: foundation; system-wide. Evidence covers the primitives package's own build/release pipeline; it does not include evidence of how downstream consumer packages (primer/react, primer/css, primer/view_components) version-pin or adopt specific primitives releases, which is out of scope for this foundations-only evidence set.
- What broke
- No confirmed break in the delivery mechanism itself.
- Impact
- Without measured migration-impact evidence, it is unknown (not disproven) whether removed/changed tokens are tracked for downstream consumer impact beyond a build-time flag.
- Why
- check:removed-tokens (package.json scripts) exists as a detection mechanism, but the packet has no evidence of it producing measured, published impact reports.
- What to change
- To support anchor 4 in a future evaluation, publish or reference migration-impact measurements (for example, a changelog entry quantifying affected consumer surface) alongside removed/breaking token changes.
Evidence citations
5aeca3d83424d7cde1d4/files/package.json, original source: lines 2-45aeca3d83424d7cde1d4/files/package.json, original source: lines 31-535aeca3d83424d7cde1d4/files/README.md, original source: lines 10-165aeca3d83424d7cde1d4/files/CONTRIBUTING.md, original source: lines 38-465aeca3d83424d7cde1d4/files/README.md, original source: line 3

Diagnostic area
Bindings findings
Naming parityPasses: 3, 3, 2; approve3 / 4
Passes: 3, 3, 2
Primer publishes an explicit Figma naming policy requiring component and property names to mirror code ('reflected what is present in code whenever possible', PascalCase component names), and the observed Code Connect files for the sampled deep components show names matching almost exactly: Button, Dialog, ActionMenu, Blankslate, InlineMessage, Spinner, ToggleSwitch, TextInput, and FormControl.Label/Caption/Validation all use identical PascalCase identifiers in Figma node names and React exports, with prop names (variant, size, checked, open, align, disabled) matching documented React props. Known gaps are explicitly surfaced rather than hidden: Dialog.figma.tsx contains an auto-generated comment listing the Figma 'size' property as unmatched to any code prop. This satisfies systematic, policy-backed alignment with documented gaps rather than automated, measured drift detection.
Selected anchor: Names are systematically aligned across surfaces
Next anchor: Parity is automatically checked and drift is measured
Evidence needed: No evidence of an automated job that diffs Figma names against code exports and reports drift over time; reaching anchor 4 requires such a continuously-run parity/drift measurement, which is not present in the evidence.
Scope: deep-component; sample-only. Only the declared deep-component sample was inspected; PageLayout (code-only by design) and DataTable's compound sub-components were not confirmed to have Code Connect naming evidence, so parity cannot be generalized to the full 15-component-set Figma library.
- What broke
- No confirmed break at this score; the one documented exception is a labeled property gap (Dialog 'size'), not a naming mismatch.
- Impact
- Designers and engineers can reliably locate the code counterpart for a Figma component/prop by name for the sampled set, reducing translation errors during handoff.
- Why
- A published naming ADR/policy plus observed 1:1 identifier matches across nine sampled Code Connect files demonstrate systemic (not ad hoc) alignment.
- What to change
- Publish an automated name/property parity report (for example, comparing Code Connect prop keys to component prop tables in CI) to progress toward measured drift detection.
Evidence citations
65ea004a12641821fb26/content.txt, original source: lines 375-381172db810d1a67c7273e2/files/packages/react/src/Button/Button.figma.tsx, original source: lines 1-31172db810d1a67c7273e2/files/packages/react/src/Dialog/Dialog.figma.tsx, original source: lines 14-23172db810d1a67c7273e2/files/packages/react/src/FormControl/FormControl.figma.tsx, original source: lines 7-27
Design-to-code contract correspondencePasses: 2, 2, 2; approve2 / 4
Passes: 2, 2, 2
Across most of the sampled deep components, consumer-controlled design choices correspond to documented React APIs through Code Connect: Button's variant/size/disabled/leadingVisual/trailingVisual, FormControl's Label/Caption/Validation slots, ToggleSwitch's state-derived loading boolean, Blankslate's conditional secondary action, ActionMenu's trigger/open/align, and DataTable's Header/TextCell/LabelCell/RowActionsCell/ColumnHeaderCell subcomponents all map coherently to code composition and props. However, Dialog's Figma 'size' variant (including a 'full' option) has no code equivalent, and the Code Connect source itself states 'No matching props could be found' without an evidenced rationale or alternate mapping: code's width (small/medium/large/xlarge) and height (small/large/auto) maps do not include a 'full' value. This is a real, unexplained gap in a materially significant component, which keeps the sample below the 'surface-specific controls are explicitly classified' bar required for score 3.
Selected anchor: Core consumer choices correspond through documented APIs or runtime mechanisms, with justified surface-specific controls excluded
Next anchor: Supported choices, derived states, slots, and composition map coherently, and surface-specific controls are explicitly classified
Evidence needed: Dialog's Figma 'size' variant (including 'full') is flagged as unmapped without an evidence-backed rationale, so surface-specific controls are not consistently and coherently classified across the full sample as anchor 3 requires.
Scope: deep-component; sample-only. Findings apply only to the declared deep-component sample; correspondence quality for the remainder of the public catalog is not established by this evidence.
Binding mappings
- Button.variant:
derived_mapping; design: Figma enum variant: primary/secondary/danger/invisible; code: variant prop: primary/default/danger/invisible. Code Connect explicitly maps Figma 'secondary' to code 'default'; a documented one-to-one correspondence with a naming difference, not a gap. - Button.size/disabled/leadingVisual/trailingVisual:
shared_contract; design: Figma size/state enums, leadingVisual?/trailingVisual? booleans with instance swap; code: size, disabled, leadingVisual, trailingVisual props. Button.figma.tsx maps each property directly to a matching React prop. - Dialog.position:
shared_contract; design: Figma position enum: center/left/right/bottom; code: position prop. Dialog.figma.tsx maps the position variant directly to the code position prop. - Dialog.size:
genuine_gap; design: Figma size enum: small/medium/large/full/xlarge/small-portrait/medium-portrait; code: width prop (small/medium/large/xlarge) and height prop (small/large/auto). Code Connect explicitly states no matching code prop was found; code's width/height maps do not include a 'full' value and no rationale is evidenced for the omission. - FormControl.Label/Caption/Validation:
shared_contract; design: Figma Label/Caption/Validation components with textContent and variant properties; code: FormControl.Label, FormControl.Caption, FormControl.Validation subcomponents. FormControl.figma.tsx maps each Figma subcomponent 1:1 to its React slot subcomponent with matching content/variant. - ToggleSwitch.loading:
derived_mapping; design: Figma 'state' variant option 'loading' (alongside rest/active/hover); code: loading boolean prop. A Figma state-enum option is flattened into a standalone boolean code prop; the same meaning expressed through a different mechanism. - DataTable subcomponents:
shared_contract; design: Figma component sets DataTable/Header, DataTable/ColumnHeaderCell, DataTable/TextCell, DataTable/LabelCell, DataTable/RowActionsCell; code: DataTable/Header, DataTable/ColumnHeaderCell, DataTable/TextCell, DataTable/LabelCell, DataTable/RowActionsCell code subcomponents. Figma component set names and structure match the compound code contract declared for DataTable in the coverage profile. - PageLayout:
unknown; design: not applicable (no Figma design component); code: PageLayout.Header/Content/Pane/Sidebar/Footer React API. PageLayout is declared code-only with no Figma counterpart (no 'View in Figma' link on its docs page, unlike other sampled components); this is a documented scope boundary rather than a mapping to compare.
- What broke
- Dialog's Figma 'size' variant (small/medium/large/full/xlarge/small-portrait/medium-portrait) has no corresponding code prop; Code Connect explicitly notes 'No matching props could be found' with no rationale or alternate mapping, and code's width/height maps omit a 'full' value.
- Impact
- A designer selecting the 'full' Dialog size in Figma has no documented way to verify what code output it corresponds to, risking silent visual drift for that variant.
- Why
- Dialog.figma.tsx (primer-react-code-connect) comments out the size prop as unmatched, while Dialog.tsx's width (small/medium/large/xlarge) and height (small/large/auto) props do not expose an equivalent 'full' value; other sampled components (Button, FormControl, ToggleSwitch, Blankslate, ActionMenu, DataTable) show coherent, classified mappings by contrast.
- What to change
- Add an explicit mapping or documented rationale in Dialog.figma.tsx for the 'full' and portrait size variants, or extend the code width/height props to cover the missing values.
Evidence citations
172db810d1a67c7273e2/files/packages/react/src/Dialog/Dialog.figma.tsx, original source: lines 5-24672126c6f32db3e8dd30/content.txt, original source: lines 753-766172db810d1a67c7273e2/files/packages/react/src/FormControl/FormControl.figma.tsx, original source: lines 7-43172db810d1a67c7273e2/files/packages/react/src/ToggleSwitch/ToggleSwitch.figma.tsx, original source: lines 4-33
Token parityPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Primer Primitives compiles a single token source (src/tokens, via style-dictionary) into both a Figma-compatible export (npm run build:figma) and code-consumable CSS variables (dist/css), using $extensions.org.primer.figma metadata to record each token's Figma collection, mode, and scopes. The same light/dark/high-contrast/colorblind/tritanopia modes are shipped as both code CSS variable files and Figma variable modes. The Primer Web Figma file evidence confirms full-file variable coverage (880 resolved variables) with 175 bound variable references sampled across the deep-component set, showing tokens are actively bound to design assets, not just declared. This demonstrates token identity and modes aligning across design and code by construction (single source of truth), rather than merely matching visual values.
Selected anchor: Token identity and modes align across design and code
Next anchor: Cross-surface token parity is automatically validated
Evidence needed: No evidence of an automated CI check that continuously validates that shipped code tokens and published Figma variables remain in sync after each change (the primitives repo's a11y-contrast workflow validates contrast, not cross-surface token parity); reaching anchor 4 requires such a validation/measurement job.
Scope: foundation; system-wide. The build pipeline mechanism (single-source generation) is a system-wide guarantee, but the confirmed bound-variable evidence is limited to the sampled component nodes; full-file coverage was declared for the Figma file overall, not independently verified for every one of the 2103 components in the library.
- What broke
- No confirmed break; the deprecated separate Primer Primitives Figma file is explicitly excluded as an evidence source, so no stale-token conflict from that legacy path is scoped in.
- Impact
- Designers using bound Figma variables and engineers using the corresponding CSS variables can expect matching semantics and mode behavior for the sampled tokens, reducing visual drift risk.
- Why
- Token generation from one source (buildTokens.ts / buildFigma.ts) with shared collection/mode metadata structurally prevents naming or value divergence between the two build outputs.
- What to change
- Add a CI job that fetches the currently published Figma variable set and diffs it against the code token build output, to move from single-source-guaranteed parity to actively measured, automatically validated parity.
Evidence citations
5aeca3d83424d7cde1d4/files/README.md, original source: lines 93-1175aeca3d83424d7cde1d4/files/README.md, original source: lines 46-56dfc83d5c91c7186f6315/figma-sample-details.json, original source: /nodes/0/bound_variable_ids
Traceability and deprecationPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Code Connect mapping files exist for every sampled deep component and link each to a specific, versioned Figma node-id/URL within the Primer Web file; a CI workflow automatically republishes these mappings whenever .figma.tsx files change. Separately, Primer's public component status policy defines Experimental/Ready/Deprecated lifecycle states, requires a consumer-facing warning for deprecated components, and is applied consistently as a visible status badge on every sampled component's documentation page. A public migration guide index further cross-references specific deprecated components (experimental SelectPanel, Flash) to their replacements with dedicated upgrade guides. Together these constitute versioned mappings that expose both supported and deprecated APIs, though no automated enforcement of mapping accuracy or deprecation drift was evidenced.
Selected anchor: Versioned mappings expose supported and deprecated APIs
Next anchor: Traceability and deprecation drift are automatically enforced
Evidence needed: To reach automatically enforced traceability/deprecation drift (4), evidence would be needed that the Code Connect publish workflow or another CI gate validates mappings against the component's current prop types/status (failing the build on mismatch) rather than only publishing whatever mapping is committed.
Scope: mixed; sample-only. Automated publish/versioning is directly evidenced only for the sampled deep components with Code Connect files; the component status/deprecation policy and migration index are documented as system-wide mechanisms and observed applied consistently across all sampled component pages, but full-catalog application across all ~192 component sets was not individually verified.
- What broke
- No confirmed break in the traceability mechanism itself; the gap is the absence of automated enforcement of mapping accuracy or deprecation-status drift.
- Impact
- A Code Connect mapping or a component's deprecation status could drift from the live implementation between publishes without an automated signal, relying on manual review to catch inconsistencies.
- Why
- The publish workflow runs 'figma connect publish' on push to main for changed .figma.tsx files but includes no validation/test step comparing the mapping against the component's current prop types or status.
- What to change
- Add a CI validation step (for example, 'figma connect parse'/dry-run or a type-check against the mapped component's props) that fails the build when a Code Connect mapping references props or values no longer present in the component API.
Evidence citations
172db810d1a67c7273e2/files/.github/workflows/figma_connect_publish.yml, original source: lines 1-364066f07b7f661cc381e7/content.txt, original source: lines 361-378b0ecf70d6ccc6d2328dc/content.txt, original source: lines 361-374172db810d1a67c7273e2/files/packages/react/src/Button/Button.figma.tsx, original source: lines 33-36

Diagnostic area
Components findings
Core component coveragePasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
primer.style/product's Components navigation lists a large, multi-archetype catalog (ActionMenu, Blankslate, Button, DataTable, Dialog, FormControl, InlineMessage, PageLayout, Spinner, TextInput, ToggleSwitch, plus dozens more spanning navigation, overlay, feedback, layout, and data-display needs) each tagged with React/Rails readiness status (e.g. 'ready' for Button and Dialog, 'experimental' for DataTable and Blankslate). This shows documented, reusable components exist for the common product needs implied by the declared deep-component and composition scope. However, the packet contains no public evidence of aggregate gap-tracking, adoption metrics, or component-health measurement (the governance source list references a component-status policy but its content was not part of this evidence packet), so the coverage claim stops at comprehensiveness rather than measured health.
Selected anchor: Coverage is comprehensive for the stated scope
Next anchor: Gaps, adoption, and component health are measured
Evidence needed: No public evidence of adoption dashboards, gap analysis, or aggregate component-health measurement was present in the packet; only per-component readiness tags are shown, not system-wide gap/adoption tracking.
Scope: catalog; system-wide. Catalog listing proves breadth of published, documented components across archetypes but does not itself prove per-component state, accessibility, binding, or responsive quality: those are assessed under the other criteria using sampled deep-component evidence.
- What broke
- No confirmed break; the catalog is comprehensive but public gap/adoption/health measurement is absent from the evidence.
- Impact
- Consumers can see a component exists and its readiness tag, but cannot independently verify from public evidence which components are under-adopted, deprecated in practice, or degrading in health.
- Why
- The evidence packet's governance source_ids reference a component-status policy but its content was not fetched into this packet, leaving readiness tags as the only visible signal.
- What to change
- Publish (or make visible in this evidence scope) adoption/usage metrics and gap-tracking dashboards tied to the component catalog to support anchor-4 measurement claims.
Evidence citations
c7ead02fcec241d44d97/content.txt, original source: lines 113-238230a179f4b704e33f45b/content.txt, original source: line 368e3ee03b83355a8615d5c/content.txt, original source: line 368
Normal, recovery, and permission statesPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Across the declared deep-component sample, normal, loading, error/validation, success, and permission-like (inactive) states are both documented and evidenced in executable test stories. Button documents Loading, 'Loading with visuals', and Inactive ('visually disabled... intended when a system error such as an outage prevents the action') states. TextInput and FormControl document and test Error/Success validation and Loading states, and FormControl.test.ts includes dedicated 'With Success Validation' and 'With Error Validation' VRT stories. ActionMenu documents an inactive-item state explicitly tied to system outages (a permission/degraded scenario) and a loading-item state. InlineMessage provides critical/warning/success/unavailable tone variants directly supporting error and recovery messaging. The loading and degraded-experiences pattern docs describe a full lifecycle (initiated, in-progress, succeeded, failed) and recovery guidance (replace with error message, or Blankslate for larger areas), matching the declared loading-recovery and empty-state-recovery composition states. This constitutes comprehensive ownership and guidance across the sample, but there is no evidence that state-contract completeness itself (for example, which states exist per component, whether all declared composition states are covered) is measured or enforced as a gate.
Selected anchor: Normal and recovery states have comprehensive ownership and guidance
Next anchor: State contracts are executable, tested, and measured
Evidence needed: No evidence that state-contract coverage itself is measured or gated (for example, a report showing which components/states are missing tests); tests confirm individual states exist and render correctly but not that the full declared state matrix is continuously tracked for completeness.
Scope: mixed; sample-only. State evidence is drawn from the declared deep-component sample and named UI patterns; it cannot be generalized to catalog components outside this sample, and composition-level state evidence (e.g. a full create-configure-save flow test) was not directly observed: only the component-level state building blocks and pattern-level guidance were.
- What broke
- No confirmed break in the sampled state contracts; the gap is that state-coverage completeness is not itself measured.
- Impact
- Teams building on the declared create-configure-save and loading-recovery flows have solid documented and partially-tested state contracts for the sampled components, but cannot verify from public evidence that all required states across the whole catalog are tracked or enforced.
- Why
- VRT/AAT tests exist per-component for specific named states (e.g. FormControl 'With Error Validation'), but no dashboard or CI gate enforcing a required-states checklist per component was found in the packet.
- What to change
- Publish or evidence a state-coverage checklist/gate (for example, CI check requiring loading/error/disabled/empty stories per new component) to support anchor-4 'executable, tested, and measured' state contracts.
Evidence citations
230a179f4b704e33f45b/content.txt, original source: lines 459-528add9c1967361d4f74f06/content.txt, original source: lines 441-54425d93d131a1dc240928e/files/e2e/components/FormControl.test.ts, original source: lines 62-69052102bad533067eee4a/content.txt, original source: lines 479-52180ebd2b43e8b45cb676d/content.txt, original source: lines 371-392ae0faabbdc1e23d922e0/content.txt, original source: lines 429-437
Responsive behaviorPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Primer publishes strong responsive foundations: a viewport-range model (narrow <768px single column, regular >=768px up to 2 columns, wide >=1400px up to 3 columns), a breakpoint scale (xsmall-xxlarge) with per-breakpoint padding rules, minimum viewport width (320px) and minimum target size (24px AA / 44px AAA) requirements, and user-preference media feature support (prefers-color-scheme, prefers-contrast, prefers-reduced-motion, forced-colors, inverted-colors). Within the declared flow scope, Dialog documents explicit narrow/regular responsive positioning (`position={narrow: 'bottom', regular: 'center'}`) and PageLayout.Sidebar's `responsiveVariant="fullscreen"` explicitly expands to a full-viewport overlay below 768px (narrow) versus staying inline at regular. Blankslate.test.ts executes VRT screenshots at explicit narrow/regular-adjacent breakpoints (`primer.breakpoint.xs`, `primer.breakpoint.sm`). This shows components adapt consistently across the supported narrow/regular conditions for the sampled components, but the packet does not show equivalent viewport-specific VRT coverage for TextInput, FormControl, ToggleSwitch, ActionMenu, InlineMessage, or Spinner, so matrix-wide continuous testing is not established.
Selected anchor: Components adapt consistently across supported conditions
Next anchor: Responsive behavior is continuously tested across the matrix
Evidence needed: Viewport-specific VRT/AAT execution was only confirmed for Blankslate (and responsive props for Dialog/PageLayout) in this packet; TextInput, FormControl, ToggleSwitch, ActionMenu, InlineMessage, and Spinner show no direct narrow/regular test evidence, so a continuously-tested full responsive matrix is not established.
Scope: mixed; sample-only. Foundational viewport-range and breakpoint guidance is system-wide, but executable responsive test evidence (viewport-specific VRT) is confirmed only for Blankslate in this packet, with responsive code contracts (not test execution) confirmed for Dialog and PageLayout; this cannot be generalized to the rest of the declared sample or catalog without further evidence.
- What broke
- No confirmed responsive failure; the gap is that continuous viewport-matrix testing is only evidenced for part of the declared sample.
- Impact
- Consumers building the create-configure-save flow (PageLayout, FormControl, TextInput, ToggleSwitch, Button, InlineMessage, Dialog, Spinner) at narrow and regular viewports get strong foundational guidance and confirmed adaptive behavior for Dialog/PageLayout/Blankslate, but cannot verify from public evidence that TextInput, FormControl, ToggleSwitch, ActionMenu, InlineMessage, and Spinner are regression-tested at those same viewports.
- Why
- The e2e test files provided for TextInput, FormControl, ToggleSwitch, ActionMenu, and InlineMessage run VRT/AAT checks without setting narrow/regular-specific viewport sizes, unlike Blankslate.test.ts which explicitly does.
- What to change
- Extend explicit narrow/regular viewport VRT coverage to the remaining sampled components (TextInput, FormControl, ToggleSwitch, ActionMenu, InlineMessage, Spinner) to support anchor-4 continuous matrix testing.
Evidence citations
90e111436b688c05c92e/content.txt, original source: lines 391-406672126c6f32db3e8dd30/content.txt, original source: lines 601-6583817f5910fb4bfa7835e/content.txt, original source: lines 1636-171925d93d131a1dc240928e/files/e2e/components/Blankslate.test.ts, original source: lines 70-8504e4c4a823862a0c2a6d/content.txt, original source: lines 383-406
Accessibility behaviorPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Component-level accessibility semantics are documented and implemented across the sample: ToggleSwitch requires `aria-labelledby`; ActionMenu single/multi-select patterns use `role="menuitemradio"`/`role="menuitemcheckbox"` with `aria-checked`; Dialog documents `returnFocusRef`/`initialFocusRef`/`role` for focus management; Button's loading state 'sets aria-disabled and preserves focus automatically'; the loading pattern doc gives detailed AT guidance (aria-labelledby for indicators, aria-busy on live regions, avoiding over-announcement). Beyond documentation, primer/react runs an automated axe-based Accessibility Acceptance Test (AAT) suite (Axe.test.ts) that iterates over nearly all Storybook stories (excluding a small, explicitly-named skip list) and asserts `toHaveNoViolations()`, executed via a sharded CI workflow (aat-reports.yml) that is wired into the merge-gating `reports.yml` workflow triggered on push to main and on merge-queue `checks_requested`. This is a system-wide, continuously-executed mechanism (not limited to the declared sample), giving direct evidence that accessible behavior is comprehensive and tested. It does not, however, demonstrate assistive-technology (for example, screen-reader) interaction coverage or regression measurement: axe checks are automated DOM/ARIA rule validation, not simulated or manual AT testing: so the top anchor's 'assistive-technology coverage' clause is not met.
Selected anchor: Accessible behavior is comprehensive and tested
Next anchor: Assistive-technology coverage and regressions are continuously measured
Evidence needed: No evidence of assistive-technology interaction coverage (for example, screen-reader announcement correctness, simulated AT regression tracking) beyond automated axe/DOM rule checks; anchor 4 requires that AT coverage specifically, not just axe rule compliance, be continuously measured.
Scope: mixed; system-wide. The AAT/axe mechanism is a cited system-wide mechanism (it iterates over nearly all Storybook stories, not just the declared sample), justifying broader generalization for automated rule-based accessibility testing specifically; it does not extend to proving assistive-technology interaction coverage, which remains unevidenced at any scope.
- What broke
- No confirmed critical accessibility failure; the gap is between automated axe coverage and demonstrated assistive-technology interaction coverage.
- Impact
- Programmatic accessibility regressions (missing labels, invalid ARIA, contrast rule violations covered by axe) are caught continuously in CI across nearly the whole catalog, but real assistive-technology usage regressions (screen-reader announcement wording/timing, keyboard-only task completion) are not demonstrably measured from public evidence.
- Why
- Axe.test.ts uses axe-core's `toHaveNoViolations()`, a static/DOM ruleset, and the CI workflows execute it broadly and continuously, but no manual or simulated AT (for example, screen-reader) regression suite was present in the evidence.
- What to change
- Publish or add evidence of assistive-technology interaction testing (manual AT audits, simulated screen-reader test suites) tracked continuously alongside the existing axe AAT suite to support anchor-4 AT coverage claims.
Evidence citations
25d93d131a1dc240928e/files/.github/workflows/aat-reports.yml, original source: lines 1-9325d93d131a1dc240928e/files/.github/workflows/reports.yml, original source: lines 1-2125d93d131a1dc240928e/files/e2e/components/Axe.test.ts, original source: lines 33-61853c59dd3ecf59e28121/content.txt, original source: lines 462-473052102bad533067eee4a/content.txt, original source: lines 558-620ae0faabbdc1e23d922e0/content.txt, original source: lines 473-501

Diagnostic area
Structure findings
Layout primitivesPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
PageLayout is documented as code-only (no Figma design component) per its declared disposition, which the packet treats as an evidenced scope boundary rather than a failure. Its React API provides a real, working container/shell contract (Header, Content, Pane, Sidebar, Footer) with padding, divider, gap, sticky, resizable, and per-viewport 'hidden'/'responsiveVariant' props, demonstrated in normative code examples rather than asserted only in prose. This operational evidence is paired with foundation-level rules (viewport ranges narrow/regular/wide, breakpoint sizes, per-breakpoint padding) that are declared full-file foundation coverage. Together these cover the stated product scope for the declared flow's narrow/regular viewports, satisfying anchor 3, but there is no evidence of automated enforcement or usage measurement needed for anchor 4.
Selected anchor: Layout primitives cover the stated product scope
Next anchor: Layout use and exceptions are validated and measured
Evidence needed: No automated validation, linting, or usage/exception measurement of PageLayout adoption was found in the evidence; anchor 4 requires such enforcement or measurement evidence, which is absent.
Scope: mixed; sample-only. Evidence demonstrates PageLayout's own region/responsive contract and the site-wide viewport-range/breakpoint foundation, but does not establish that other layout surfaces (for example, SplitPageLayout, Stack, CSS Grid/Flexbox utilities) implement or enforce the same primitives. PageLayout itself carries a code-only design disposition, so no Figma-side layout-primitive evidence exists.
- What broke
- No confirmed break: layout primitive coverage is well-evidenced for the declared page-level scope.
- Impact
- Teams building the create-configure-save flow can rely on one documented container/shell contract (PageLayout) with responsive region behavior across narrow and regular viewports, reducing the risk of ad hoc page structure.
- Why
- PageLayout's props (padding, divider, hidden-by-viewport, Sidebar responsiveVariant) directly implement the viewport-range and breakpoint rules published in the Layout and Responsive foundation docs.
- What to change
- Publish evidence of automated enforcement or measurement of PageLayout usage and exceptions (for example, a lint rule flagging bespoke page shells, or adoption telemetry) to support a future anchor-4 claim.
Evidence citations
3817f5910fb4bfa7835e/content.txt, original source: lines 379-3973817f5910fb4bfa7835e/content.txt, original source: lines 1834-187890e111436b688c05c92e/content.txt, original source: lines 391-40590e111436b688c05c92e/content.txt, original source: lines 555-57704e4c4a823862a0c2a6d/content.txt, original source: lines 381-389
Composition guidancePasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
The Forms, Loading, Empty states, and Degraded experiences pattern pages provide explicit, non-isolated composition rules directly tied to the declared flow/pattern component sets. Forms pattern anatomy (Label/Input/Caption/Validation), structure ('default to vertically stacked FormControls'), and validation-on-submit focus/ARIA rules cover the create-configure-save flow's validation-error state. Loading pattern's lifecycle (initiated/in-progress/complete/failed) and scoping guidance cover the loading state and loading-recovery pattern. Degraded/empty-state pattern pages give explicit Blankslate+Button+InlineMessage assembly rules (leading visual, primary/secondary text and action, error copy) covering empty-state-recovery. This is documented compositional guidance beyond component inventories, satisfying 'composition contracts cover supported product assemblies,' but nothing indicates the rules are executable or automatically validated.
Selected anchor: Composition contracts cover supported product assemblies
Next anchor: Composition rules are executable or automatically validated
Evidence needed: No evidence that composition rules are executable or automatically validated (for example, lint rules enforcing FormControl+Validation pairing, Storybook interaction tests, or CI composition checks): required for anchor 4.
Scope: composition; sample-only. Only the three declared compositions/patterns are evidenced; guidance for other listed UI patterns (navigation, notification messaging, progressive disclosure, saving, feature onboarding) is not in this evidence packet and is not assumed
- What broke
- No confirmed break; composition guidance exists in documentation but is not shown to be enforced automatically
- Impact
- Engineers building the declared flow have clear documented rules for assembling FormControl/TextInput/Dialog/Spinner/InlineMessage correctly, but nothing prevents a non-compliant assembly from shipping since enforcement is not evidenced
- Why
- Evidence is limited to normative prose and code examples from pattern and component doc pages; no lint, CI, or Storybook interaction-test artifacts are in the packet
- What to change
- Not applicable to public-evidence scoring; a higher score would require publicly observable automated enforcement of composition rules
Evidence citations
53523b7aa25b8b2612cf/content.txt, original source: lines 429-43353523b7aa25b8b2612cf/content.txt, original source: lines 499-515ae0faabbdc1e23d922e0/content.txt, original source: lines 429-437
Patterns and templatesPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
The three declared compositions each map to a dedicated, normative Primer UI-pattern page (Forms, Loading, Degraded experiences, Empty states) addressing recurring product tasks (data entry/validation, async waiting, outage handling, first-use/no-data states). Each pattern page explicitly cross-references the components that implement it: for example, the Empty states page ties directly to Blankslate, the Degraded experiences page links to Blankslate/Dialog/Tooltip/Loading/Messaging, and the Loading page links to DataTable/SelectPanel/TreeView/Spinner/SkeletonLoaders/ProgressBar: satisfying anchor 3's 'remain linked to components' clause for important tasks. No evidence shows pattern adoption, outcomes, or lifecycle being tracked, so anchor 4 is not met.
Selected anchor: Patterns cover important tasks and remain linked to components
Next anchor: Pattern use, outcomes, and lifecycle are measured
Evidence needed: No evidence of pattern usage being tracked or of pattern lifecycle (for example, deprecation, adoption metrics, outcome measurement): anchor 4 requires pattern use, outcomes, and lifecycle to be measured, which is absent.
Scope: composition; sample-only. Only the three declared compositions/patterns and their directly linked components were evaluated; Primer's broader 'Scenario Patterns' set (Copy, Create, Delegate, Delete, Filter, Search, View) referenced in site navigation is outside this evidence's scope and was not assessed.
- What broke
- No confirmed break: the sampled patterns are documented and explicitly linked to their implementing components.
- Impact
- Teams solving forms, loading/degraded-recovery, or empty-state tasks have a canonical pattern reference tied to specific components rather than having to reverse-engineer conventions from isolated examples.
- Why
- Each pattern doc (Forms, Loading, Degraded experiences, Empty states) contains dedicated related-links or inline cross-references to the exact components (Blankslate, Dialog, Spinner, etc.) used in the declared compositions.
- What to change
- Publish pattern-adoption or outcome metrics (for example, which teams use the documented forms pattern, error-recovery success rates) to progress toward anchor 4; none is present in current evidence.
Evidence citations
53523b7aa25b8b2612cf/content.txt, original source: lines 561-573ae0faabbdc1e23d922e0/content.txt, original source: lines 365-373ae0faabbdc1e23d922e0/content.txt, original source: lines 533-5449802e3b6b3f4bf5f5ac2/content.txt, original source: lines 359-3657cbf28d142607f99b54f/content.txt, original source: lines 576-589
Navigation, focus, and state ownershipPasses: 3, 3, 2; approve3 / 4
Passes: 3, 3, 2
Dialog's code contract explicitly assigns focus ownership (returnFocusRef, initialFocusRef, onClose gesture argument) and responsive positioning (position={{narrow:'bottom', regular:'center'}}). The Loading pattern's dedicated Focus management section specifies who owns focus during and after async state changes (Dialog auto-returns focus on close, move focus to first invalid field on failure, move focus into newly loaded content, aria-live/aria-busy rules to prevent premature announcement). The Forms pattern's validation-on-submit section specifies the exact ownership contract for error state (aria-invalid, aria-describedby wiring, focus to interactive summary or first invalid field). This is comprehensive, code- and documentation-level ownership coverage across the declared flow's dialog, validation-error, loading, and responsive-transition scenarios, satisfying 'navigation, focus, and state contracts are comprehensive.' No evidence shows these contracts are executable or continuously verified (for example, automated focus-trap or ARIA tests), so anchor 4 is unreached.
Selected anchor: Navigation, focus, and state contracts are comprehensive
Next anchor: Cross-composition contracts are executable and continuously verified
Evidence needed: No evidence that these ownership contracts are executable or continuously verified (for example, automated focus-trap tests, axe-core CI gating, or interaction test suites tied to these specific contracts): required for anchor 4. The primer/react package.json lists @github/axe-github and @playwright/test as devDependencies, but a tool declaration alone does not prove these contracts are actually tested or gated.
Scope: composition; sample-only. Ownership evidence is drawn from Dialog, Forms, Loading, and Degraded-experiences documentation for the declared flow/patterns; it does not establish ownership contracts for other overlays, routes, or nested compositions outside this sample (for example, Popover, Overlay, TreeView) beyond incidental mentions
- What broke
- No confirmed break; ownership is documented comprehensively but not shown to be automatically or continuously verified
- Impact
- Implementers of the create-configure-save flow have clear, comprehensive guidance for who owns focus/state during dialogs, validation errors, loading, and responsive transitions, but nothing confirms these contracts are enforced or regression-tested in practice
- Why
- Evidence includes explicit component props and detailed pattern-level prose, but the packet contains no test results, CI gate output, or execution logs tied to these specific ownership contracts
- What to change
- Not applicable to public-evidence scoring; a higher score would require publicly observable automated/continuous verification of these ownership contracts
Evidence citations
672126c6f32db3e8dd30/content.txt, original source: lines 768-772ae0faabbdc1e23d922e0/content.txt, original source: lines 513-52153523b7aa25b8b2612cf/content.txt, original source: lines 503-513

Diagnostic area
Governance findings
Named ownershipPasses: 2, 3, 2; revise3 / 4
Passes: 2, 3, 2
Operational ownership is evidenced across the public surfaces without requiring one specific mechanism. Primer Web requires named DRI approval before maintainer merge and identifies the Design Infrastructure support boundary; Primer Primitives has repository-wide engineering review ownership; Primer React documents Primer-team review and merge decision rights, a response expectation, and weekly proposal triage. The absence of a public Primer React CODEOWNERS file is an unknown implementation detail, not evidence that these documented decision rights are non-operational.
Selected anchor: Decision rights and support boundaries are operational
Next anchor: Ownership health and service expectations are measured
Evidence needed: Additional evidence must satisfy the next anchor: Ownership health and service expectations are measured
Scope: governance; sample-only. The primer/react code repository's own CODEOWNERS file was not retrievable in this snapshot (marked as a missing optional path), so named/operational code-review ownership for the React implementation specifically is unconfirmed beyond the generic 'a contributor of Primer React will review' language; conclusions generalize confidently to primitives and Figma but not to the React code repo specifically.
- What broke
- Ownership health (for example, response-time or backlog-aging expectations) is not publicly measured or reported.
- Impact
- Consumers can identify who owns a component/library and how to escalate, but cannot verify from public evidence how well or how quickly that ownership function actually performs over time.
- Why
- Public evidence documents named DRIs, an enforced CODEOWNERS team, and support channels, but includes no publicly available dashboard or service-level metric; governance known_limitations explicitly place internal adoption/service dashboards out of scope.
- What to change
- Publish periodic public reporting on ownership responsiveness (for example, median PR review time for CODEOWNERS-gated repos, DRI review backlog) to move from operational to measured ownership.
Evidence citations
afbb7a337d452b0d0a14/files/contributor-docs/CONTRIBUTING.md, original source: line 42afbb7a337d452b0d0a14/files/contributor-docs/CONTRIBUTING.md, original source: lines 259-27265ea004a12641821fb26/content.txt, original source: lines 365-36765ea004a12641821fb26/content.txt, original source: lines 519-525f478be6e00ccc7aec688/content.txt, original source: lines 456-4585aeca3d83424d7cde1d4/files/.github/CODEOWNERS, original source: line 1
Contribution and reviewPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
primer/react's contributor docs publish an operational, cross-discipline review process: a documented 'What we look for in reviews' checklist (code style, theme values, API design, type definitions, documentation, tests, bundle size, CI checks), a stated review turnaround ('within a day or two'), and a documented release cadence (weekly minor/patch, biannual major). This is enforced mechanically by the check_for_changeset.yml CI gate, which blocks PRs lacking a changeset unless an explicit skip-label override is applied. In parallel, Primer Web Figma has an operational design-review path: contributors branch, request review from the file's DRI, and a maintainer merges once approved, backed by an explicit contribution checklist (including accessibility). Together these show cross-discipline (engineering + design) review criteria and decisions operating, satisfying anchor 3. No public throughput, rejection-rate, or review-quality metrics exist, so anchor 4 is not reached.
Selected anchor: Cross-discipline review criteria and decisions are operational
Next anchor: Review quality, throughput, and outcomes are measured
Evidence needed: No public review-outcome measurement (for example, time-to-merge distributions, rejection rates, changeset-compliance rate) is published; internal Slack and GitHub-internal review artifacts are explicitly out of the governance evidence scope, so throughput/quality of review cannot be measured.
Scope: governance; sample-only. Evidence covers primer/react and Primer Web Figma only; other Primer repos (CSS, ViewComponents, Octicons) and internal review artifacts are not evidenced.
- What broke
- Review throughput, quality, and outcome metrics are not published, so the effectiveness of the documented review process cannot be measured from public evidence.
- Impact
- It cannot be confirmed from public evidence whether the documented checklist and DRI review model translate into consistently fast or high-quality outcomes at scale, versus being aspirational policy.
- Why
- contributor-docs/CONTRIBUTING.md and the Figma contribution guide document review criteria and process, but per the governance known_limitations, private Slack discussions and GitHub-internal review artifacts are outside evidence scope, and no aggregate review-outcome data is published.
- What to change
- Publish periodic review-health metrics (for example, median PR review time, changeset-compliance rate, Figma branch approval latency) to demonstrate the documented process operates at the claimed level.
Evidence citations
afbb7a337d452b0d0a14/files/contributor-docs/CONTRIBUTING.md, original source: lines 257-272afbb7a337d452b0d0a14/files/.github/workflows/check_for_changeset.yml, original source: lines 1-4465ea004a12641821fb26/content.txt, original source: lines 364-36965ea004a12641821fb26/content.txt, original source: lines 494-51851ca77b7640ae2ad9db1/content.txt, original source: lines 378-379
Releases, migration, and deprecationPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Deprecation and migration have explicit operational contracts, satisfying anchor 3. The normative component-status policy defines a lifecycle (Experimental/Ready/Deprecated) with concrete, binding consequences: for 'Ready' components, 'breaking changes... will result in a major version bump... and Primer will provide a migration path,' and 'Deprecated' components must have deprecation documentation and 'a warning is shown to the consumer' at use-time. This contract is backed by two published, code-level migration guides (experimental SelectPanel -> stable SelectPanel, with detailed prop-mapping tables and before/after code; Flash -> Banner) and by changeset-driven automated versioning and changelog generation (evidenced by extensive, PR-linked CHANGELOG.md entries on both primer/react and primer/primitives) enforced by the CI changeset gate. RELEASING.md documents a release-candidate testing and publish process. This is current, normative documentation plus directly observed automation (changesets config, CHANGELOG structure), meeting the evidence-precedence bar. Anchor 4 (validated/measured compatibility and migration success) is not evidenced: release-candidate testing is explicitly scoped '(GitHub staff only)' with no public test results, and there is no published migration-success or compatibility-regression measurement.
Selected anchor: Deprecation and migration have explicit operational contracts
Next anchor: Compatibility and migration success are validated and measured
Evidence needed: No public evidence that compatibility or migration success is validated or measured (for example, automated codemod verification, migration completion tracking, or compatibility test results); release-candidate testing is explicitly gated as staff-only with no published outcomes.
Scope: governance; system-wide. The component-status policy and changeset versioning apply to the whole public component catalog, not just the sampled components, supporting system-wide generalization; however, only two migration guides are directly evidenced and staff-only release-candidate testing/private migration outcomes remain unknown.
- What broke
- No confirmed break at anchor 3; the gap is that migration/compatibility success is asserted in policy but not measured.
- Impact
- Consumers get a documented, binding process for breaking changes and concrete migration steps, but cannot verify from public evidence that migrations actually succeed or that release candidates pass their internal tests, since that testing is GitHub-staff-only.
- Why
- The RELEASING.md release-candidate testing step is explicitly scoped to internal staff, and no public compatibility/migration-success metrics are published alongside the changelog or migration guides.
- What to change
- Publish aggregate or anonymized migration-success/compatibility validation results (for example, percentage of consumers migrated off deprecated SelectPanel, automated compat-test pass rates) to close the anchor-4 gap.
Evidence citations
4066f07b7f661cc381e7/content.txt, original source: lines 369-379b0ecf70d6ccc6d2328dc/content.txt, original source: lines 361-37466d075d0ac541a7fc467/content.txt, original source: lines 382-391cf82bbb745b7e997d38f/files/CHANGELOG.md, original source: lines 1-168c1871123ec8eee1b57e/files/RELEASING.md, original source: lines 9-21
Quality enforcement and feedbackPasses: 3, 3, 3; approve3 / 4
Passes: 3, 3, 3
Primer's public repos show concrete, executing quality gates rather than mere tool declarations: an axe-based accessibility test (Axe.test.ts) iterates over nearly all Storybook stories and asserts `toHaveNoViolations()`, with an explicit, commented exception list (SKIPPED_TESTS, including open TODOs for known contrast issues) run via a sharded CI workflow (aat-reports.yml); Playwright visual regression tests exist for each sampled deep component (Button, Dialog, FormControl, DataTable, ToggleSwitch, TextInput, InlineMessage, PageLayout, ActionMenu, Blankslate); required CI jobs enforce lint, format, unit tests (React 18/19 matrix), and type-checking on every push/PR/merge-queue event; CodeQL SAST scanning runs on push, PR, and a weekly schedule; and primer/primitives runs an automated a11y-contrast check. This direct code and workflow evidence, spanning the full sampled component set with documented exceptions, satisfies anchor 3's requirement that quality gates cover the supported lifecycle and exceptions with comprehensive operational behavior. No public evidence shows measured, acted-upon quality/regression/adoption outcomes over time (anchor 4): report artifacts and a Datadog code-metrics pipeline are configured, but their result contents are not in evidence.
Selected anchor: Quality gates cover the supported lifecycle and exceptions
Next anchor: Quality, adoption, exceptions, and regressions are measured and acted on
Evidence needed: Anchor 4 requires that quality, adoption, exceptions, and regressions be measured and acted on (for example, published pass-rate trends, tracked regression counts). CI report artifacts (blob-report, playwright-report) and a Datadog code-metrics pipeline (codescan.yml) are configured to produce measurements, but their actual result contents, trends, or evidence of follow-up action are not present in the supplied evidence.
Scope: governance; sample-only. Evidence confirms gates are wired to execute across nearly all stories/components (a system-wide mechanism via Axe.test.ts iterating the full Storybook story set), but per-component outcome confirmation is limited to the sampled deep components; actual pass/fail results, trend history, and confirmation that flagged issues (for example, open color-contrast TODOs) are eventually resolved rather than perpetually skipped are not in evidence.
- What broke
- Quality and regression outcomes are not publicly measured or reported over time, and exception entries (for example, color-contrast TODOs in SKIPPED_TESTS) have no visible resolution tracking.
- Impact
- Strong automated gates exist and demonstrably execute across nearly the full sampled component set, but the public evidence does not show whether gate failures, waived checks, or regressions are tracked to closure and acted upon systematically.
- Why
- CI workflow definitions and test source files prove the gates execute and cover exceptions explicitly, but report artifacts and the Datadog code-metrics pipeline are configured without any visible published results in the evidence, so outcome measurement and follow-up cannot be confirmed.
- What to change
- Publish periodic public summaries of AAT/VRT pass rates, CodeQL finding trends, or the aging/resolution of SKIPPED_TESTS-style exceptions so quality outcomes: not just gate existence: become independently verifiable.
Evidence citations
25d93d131a1dc240928e/files/e2e/components/Axe.test.ts, original source: lines 1-6225d93d131a1dc240928e/files/.github/workflows/aat-reports.yml, original source: lines 1-93afbb7a337d452b0d0a14/files/.github/workflows/ci.yml, original source: lines 58-91afbb7a337d452b0d0a14/files/.github/workflows/codeql.yml, original source: lines 1-71
