carbon-core-v11, design-system-agent-tasks/1.0.0, Publication-ready
Scope: Results cover the recorded actor profile, task suite, evidence snapshot, runtime, and run count.
86.7% functional task success. 100% citation and provenance accuracy.
Outputs varied substantially even when outcomes succeeded. Exact-answer reproducibility was 46.7%, while outcome agreement was 93.3%. Exact-answer reproducibility measures whether repeated runs produced the same normalized structured answer; it is not functional reliability. Different valid implementations, wording, or citation selections can lower it even when a task completes and its runtime checks pass.
| Area | Task | Functional success | Exact answer | Outcome agreement |
|---|---|---|---|---|
| Foundations | carbon-semantic-text-token | 100% | 100% | 100% |
| Bindings | carbon-button-design-to-code | 100% | 33.3% | 100% |
| Components | carbon-resource-name-field | 100% | 33.3% | 100% |
| Structure | carbon-create-configure-save-flow | 33.3% | 33.3% | 66.7% |
| Governance | carbon-breaking-prop-governance | 100% | 33.3% | 100% |
These measurements cover one recorded actor profile on controlled tasks using the captured evidence.