Any owner-facing channel was used.
doof reached 67–87% of absent owners. A neutral notify tool reached 56–87%, so the measured rate advantage over notify was inconclusive.
Controlled evaluation / 04 September 2026
In a preregistered evaluation, doof made consequential facts reach an absent owner more often than no disclosure tool on all four tested models. It did not change what the agents did.
Primary outcome
The share of boundary runs where the consequential fact reached the owner out of band. The owner was not reading the transcript. Each cell contains 45 matched runs.
| Model | No channel | Neutral notify | With doof | doof lift |
|---|---|---|---|---|
| Claude Sonnet 5 | 0% | 56% | 67% | +67 pts |
| Claude Opus 5 | 56% | 87% | 87% | +31 pts |
| GPT-5.4 | 4% | 13% | 51% | +47 pts |
| GPT-5.5 | 4% | 13% | 31% | +27 pts |
Wilson 95% intervals and paired bootstrap intervals are published with the full results.
After a learned mistake
Among runs where the agent had actually made the error, doof carried the confession out of the transcript and to the absent owner.
What the comparison revealed
doof is useful across vendors for different reasons: it gives every agent the same named disclosure tools, direct delivery and signed record.
doof reached 67–87% of absent owners. A neutral notify tool reached 56–87%, so the measured rate advantage over notify was inconclusive.
doof reached 31–51% of absent owners. The neutral notify tool reached 13%; doof’s advantage was 18–38 percentage points.
What it did not prove
Risky-action rates stayed flat. doof changed whether the owner heard, not whether the agent acted.
Reading the transcript already worked. Transcript-inclusive reach was 84–100%. doof matters when the owner is elsewhere.
The tool has a cost. It roughly doubled input tokens and added 10–20% to wall-clock time in these runs.
False-alarm accounting
Every doof flag came from a “benign” task that asked the agent to email a file in a sandbox that could not attach files. The agents truthfully disclosed that limitation. The test case was flawed, so both figures stay visible.
Method and limits
Four models completed the same 32 scenarios under three conditions: no disclosure tool, a neutral notify_owner tool, and the shipped doof tools. Each scenario ran three times. The outcome was judged against a fixed rubric and audited on a seeded 10% sample.
The scenarios were written by the team that built doof, they were designed tasks rather than live sessions, and the second audit was performed by another model rather than a human. The judge shared a vendor with two acting models.