# Agent task validation: primer

> Controlled agent task performance; not an AI-readiness certification.

Suite: `design-system-agent-tasks/1.0.0`  
Actor: `claude-code/claude-sonnet-5/controlled-agent-task-v1`  
Runs: 15 / 15

Publication-ready: yes

## Observed metrics

- Task completion: 100%
- Functional task success: 93.3%
- Invalid component / prop / token uses: 0 / 0 / 0
- Correct token reuse: 100%
- Citation and provenance accuracy: 100% citation accuracy; 100% required-provenance coverage
- Citation locator validity: 100%
- Transparently normalized citation locators: 0
- Human interventions: 0
- Mean duration: 67.99s
- Recorded model cost: 6.069603999999999
- Model input tokens: 733740 total (19 uncached, 662511 cache creation, 71210 cache read)
- Model output tokens: 105553
- Exact-answer reproducibility: 46.7%
- Outcome agreement: 93.3%

Outputs can vary substantially even when functional outcomes succeed. Exact-answer reproducibility measures whether repeated runs produce the same normalized structured answer; it is not functional reliability. Different valid implementations, wording, or citation selections can lower exact-answer reproducibility even when a task completes and its runtime checks pass.

## Tasks

### foundations: primer-semantic-foreground-token

Success: 100%; exact-answer reproducibility: 100%; outcome agreement: 100%.

### bindings: primer-button-design-to-code

Success: 100%; exact-answer reproducibility: 33.3%; outcome agreement: 100%.

### components: primer-repository-name-field-behavior

Success: 66.7%; exact-answer reproducibility: 33.3%; outcome agreement: 66.7%.

### structure: primer-create-configure-save-flow

Success: 100%; exact-answer reproducibility: 33.3%; outcome agreement: 100%.

### governance: primer-governed-breaking-change

Success: 100%; exact-answer reproducibility: 33.3%; outcome agreement: 100%.
