Agent task validation

primer, design-system-agent-tasks/1.0.0, Publication-ready

Controlled agent task performance; not an AI-readiness certification.

93.3%functional task success
100%task completion
100%citation accuracy
100%provenance accuracy
46.7%exact-answer reproducibility
0human interventions
67.99smean duration
733740input tokens
$6.07scored-run model cost

93.3% functional task success. 100% citation and provenance accuracy.

Outputs varied substantially even when outcomes succeeded. Exact-answer reproducibility was 46.7%, while outcome agreement was 93.3%. Exact-answer reproducibility measures whether repeated runs produced the same normalized structured answer; it is not functional reliability. Different valid implementations, wording, or citation selections can lower it even when a task completes and its runtime checks pass.

Task results

AreaTaskFunctional successExact answerOutcome agreement
Foundationsprimer-semantic-foreground-token100%100%100%
Bindingsprimer-button-design-to-code100%33.3%100%
Componentsprimer-repository-name-field-behavior66.7%33.3%66.7%
Structureprimer-create-configure-save-flow100%33.3%100%
Governanceprimer-governed-breaking-change100%33.3%100%

Interpretation

These measurements describe one recorded actor profile on controlled tasks using the captured evidence. They do not certify every model, prompt, product flow, or future system release.

Supporting data: JSON, CSV, Markdown