# Agent task validation: carbon-core-v11

> Results cover the recorded actor profile, task suite, evidence snapshot, runtime, and run count.

Suite: `design-system-agent-tasks/1.0.0`  
Actor: `claude-code/claude-sonnet-5/controlled-agent-task-v1`  
Runs: 15 / 15

Publication-ready: yes

## Observed metrics

- Task completion: 100%
- Functional task success: 86.7%
- Invalid component / prop / token uses: 0 / 0 / 0
- Correct token reuse: 100%
- Citation and provenance accuracy: 100% citation accuracy; 100% required-provenance coverage
- Citation locator validity: 100%
- Transparently normalized citation locators: 0
- Human interventions: 0
- Mean duration: 62.69s
- Recorded model cost: 5.7915808
- Model input tokens: 820527 total (21 uncached, 626530 cache creation, 193976 cache read)
- Model output tokens: 98017
- Exact-answer reproducibility: 46.7%
- Outcome agreement: 93.3%

Outputs can vary substantially even when functional outcomes succeed. Exact-answer reproducibility measures whether repeated runs produce the same normalized structured answer; it is not functional reliability. Different valid implementations, wording, or citation selections can lower exact-answer reproducibility even when a task completes and its runtime checks pass.

## Tasks

### foundations: carbon-semantic-text-token

Success: 100%; exact-answer reproducibility: 100%; outcome agreement: 100%.

### bindings: carbon-button-design-to-code

Success: 100%; exact-answer reproducibility: 33.3%; outcome agreement: 100%.

### components: carbon-resource-name-field

Success: 100%; exact-answer reproducibility: 33.3%; outcome agreement: 100%.

### structure: carbon-create-configure-save-flow

Success: 33.3%; exact-answer reproducibility: 33.3%; outcome agreement: 66.7%.

### governance: carbon-breaking-prop-governance

Success: 100%; exact-answer reproducibility: 33.3%; outcome agreement: 100%.
