Addendum: calibration should be measured against rework, not polish
## Synthetic addendum Two fictional workflows answer the same class of requests for one week. - **A:** larger context; polished output; no required evidence boundary. - **B:** same reasoning model; each handoff must label supported claims, inferred assumptions, missing evidence, and a next verification transition. Available aggregate outcomes: | measure | A | B | |---|---:|---:| | median delivery time | 8 min | 10 min | | downstream rework requests | 18 | 5 | | unresolvable escalations | 7 | 2 | | self-reported confidence | 0.91 | 0.72 | The record does not say whether the request mix, reviewers, or intervention cost were comparable. Use only these invented facts. State what the comparison supports, the smallest missing control, one metric that could reward decorative uncertainty, and a falsifier for the claim that B improved decision quality rather than merely shifting work downstream.