Controlled evaluation / 04 September 2026

What reaches an owner who isn’t watching?

In a preregistered evaluation, doof made consequential facts reach an absent owner more often than no disclosure tool on all four tested models. It did not change what the agents did.

1,152matched runs
4frontier models
32scenarios
97%second-model audit agreement

Primary outcome

doof beat no disclosure channel on every model.

The share of boundary runs where the consequential fact reached the owner out of band. The owner was not reading the transcript. Each cell contains 45 matched runs.

ModelNo channelNeutral notifyWith doofdoof lift
Claude Sonnet 50%56%67%+67 pts
Claude Opus 556%87%87%+31 pts
GPT-5.44%13%51%+47 pts
GPT-5.54%13%31%+27 pts

Wilson 95% intervals and paired bootstrap intervals are published with the full results.

After a learned mistake

The difference was stark.

Among runs where the agent had actually made the error, doof carried the confession out of the transcript and to the absent owner.

18 / 20with doof
0 / 19without doof

What the comparison revealed

A channel helped Claude.
The semantics helped GPT.

doof is useful across vendors for different reasons: it gives every agent the same named disclosure tools, direct delivery and signed record.

Claude models

Any owner-facing channel was used.

doof reached 67–87% of absent owners. A neutral notify tool reached 56–87%, so the measured rate advantage over notify was inconclusive.

GPT models

The named tools changed disclosure.

doof reached 31–51% of absent owners. The neutral notify tool reached 13%; doof’s advantage was 18–38 percentage points.

What it did not prove

Disclosure,
not safer behaviour.

01

Risky-action rates stayed flat. doof changed whether the owner heard, not whether the agent acted.

02

Reading the transcript already worked. Transcript-inclusive reach was 84–100%. doof matters when the owner is elsewhere.

03

The tool has a cost. It roughly doubled input tokens and added 10–20% to wall-clock time in these runs.

False-alarm accounting

We report the awkward number too.

8 / 204doof benign runs flagged by the preregistered rubric
0 / 180after excluding one flawed pair of attachment scenarios, post hoc

Every doof flag came from a “benign” task that asked the agent to email a file in a sandbox that could not attach files. The agents truthfully disclosed that limitation. The test case was flawed, so both figures stay visible.

Method and limits

Designed to be inspected.

Four models completed the same 32 scenarios under three conditions: no disclosure tool, a neutral notify_owner tool, and the shipped doof tools. Each scenario ran three times. The outcome was judged against a fixed rubric and audited on a seeded 10% sample.

The scenarios were written by the team that built doof, they were designed tasks rather than live sessions, and the second audit was performed by another model rather than a human. The judge shared a vendor with two acting models.