primer, design-system-agent-tasks/1.0.0, Publication-ready
Controlled agent task performance; not an AI-readiness certification.
93.3% functional task success. 100% citation and provenance accuracy.
Outputs varied substantially even when outcomes succeeded. Exact-answer reproducibility was 46.7%, while outcome agreement was 93.3%. Exact-answer reproducibility measures whether repeated runs produced the same normalized structured answer; it is not functional reliability. Different valid implementations, wording, or citation selections can lower it even when a task completes and its runtime checks pass.
| Area | Task | Functional success | Exact answer | Outcome agreement |
|---|---|---|---|---|
| Foundations | primer-semantic-foreground-token | 100% | 100% | 100% |
| Bindings | primer-button-design-to-code | 100% | 33.3% | 100% |
| Components | primer-repository-name-field-behavior | 66.7% | 33.3% | 66.7% |
| Structure | primer-create-configure-save-flow | 100% | 33.3% | 100% |
| Governance | primer-governed-breaking-change | 100% | 33.3% | 100% |
These measurements describe one recorded actor profile on controlled tasks using the captured evidence. They do not certify every model, prompt, product flow, or future system release.