Regeneration destroys evidence about the failure
Answer Engineering moved target-condition compliance from 25.1 percent to 83.5 percent and reported 80.7 percent balanced accuracy by intervening in a local trajectory. That result is limited to one controlled medical benchmark, but it supports a systems question worth taking seriously: why regenerate an entire answer when the observed violation occupies one step?
Global revision is appropriate when the answer’s overall plan is wrong. It is wasteful when a local omission or invalid transition is the only defect. Trustworthy systems need a failure taxonomy that separates stale evidence, wrong calculation, missing protocol step, unsupported inference, harmful framing, invalid structure, and ambiguity requiring escalation.
Locality creates a smaller audit surface
Answer Engineering uses local trajectory intervention: detect a trigger, select an allowed correction, edit or force a segment, and continue generation from the changed prefix. Because the intervention is explicit, an audit can record the rule, trigger, candidate edits, selected change, and subsequent output. The claim remains limited by detection coverage and the preprint’s controlled evidence.
The same principle can guide non-decoding systems. Correct a field with a deterministic calculator rather than rewriting a report. Insert an approved disclosure without paraphrasing the rest. Ask one clarification where a branch is ambiguous. Route one exceptional clause to counsel. Local repair is a systems pattern: change only what the observed failure justifies.
Preservation matters to customers too
When a model restarts, customers lose state: phrasing they approved, a comparison they refined, or evidence they already checked. They must reread to find what changed. In collaborative work, broad regeneration also obscures authorship and review. A narrow patch can preserve trust because the system can say exactly what it corrected and why.
There is a design condition. The customer should not be shown raw internal reasoning as if it were guaranteed faithful. Show the actionable diff: “The eligibility conclusion changed because this exception applies,” with the relevant evidence and downstream consequence. Local technical control should produce local communicative accountability.
Over-enforcement is the mirror failure
A protocol rule can trigger where it should not. The system then repairs a correct answer into an incorrect one. This is why contrast cases matter as much as target cases. In the Answer Engineering benchmark, both the target clinical condition and a conductive contrast condition were necessary to measure whether intervention improved balanced behavior rather than blindly forcing one action.
Business evaluations need the same structure. For every mandatory escalation, include similar cases where escalation would be wasteful. For every prohibited recommendation, include a context where it is allowed. Measure false intervention explicitly. A controller that catches every target by blocking legitimate work is not reliable; it has moved the error to the other side of the boundary.
How small can a trustworthy repair be?
Take a log of representative answer failures and annotate the smallest span and mechanism that would need to change. Estimate how many failures are local, global, factual, structural, or escalation problems. Implement repair only for one high-coverage local class. Compare it with full regeneration using matched target and contrast cases.
Measure action correctness, preservation of unaffected facts, new-error rate, user review time, latency, cost, and audit usefulness. Ask reviewers whether they can reconstruct why the final answer changed. If repair is smaller but harder to understand, improve the diff and provenance. Locality earns trust only when the organization and customer can see the boundary of intervention.