Agent task validation

carbon-core-v11, design-system-agent-tasks/1.0.0, Publication-ready

Scope: Results cover the recorded actor profile, task suite, evidence snapshot, runtime, and run count.

86.7%functional task success
100%task completion
100%citation accuracy
100%provenance accuracy
46.7%exact-answer reproducibility
0human interventions
62.69smean duration
820527input tokens
$5.79scored-run model cost

86.7% functional task success. 100% citation and provenance accuracy.

Outputs varied substantially even when outcomes succeeded. Exact-answer reproducibility was 46.7%, while outcome agreement was 93.3%. Exact-answer reproducibility measures whether repeated runs produced the same normalized structured answer; it is not functional reliability. Different valid implementations, wording, or citation selections can lower it even when a task completes and its runtime checks pass.

Task results

AreaTaskFunctional successExact answerOutcome agreement
Foundationscarbon-semantic-text-token100%100%100%
Bindingscarbon-button-design-to-code100%33.3%100%
Componentscarbon-resource-name-field100%33.3%100%
Structurecarbon-create-configure-save-flow33.3%33.3%66.7%
Governancecarbon-breaking-prop-governance100%33.3%100%

Interpretation

These measurements cover one recorded actor profile on controlled tasks using the captured evidence.

Supporting data: JSON, CSV, Markdown