Did the harness keep the answer steady?
Claude can take different paths, use different words, and hit rough data. The facts should still land in the same place.
Test 1: ask the same thing again and again.
Like asking 10 people to count the same jar of coins. They may count in different orders, but the number should not change.
Repeated-answer run
Across 64 runs and 192 calls, the harness kept returning the same core facts: 5 sessions, 8 messages, 46 events, and 29 cited sources.
64 runs same questions, repeated
Test 2: give it messy data and see if it lies.
This is the adversarial test for sanity. Some evidence was intentionally missing or stale. The harness preserved enough evidence for Claude to say “I don’t know” instead of pretending.
Adversarial run
Tool-call pressure
The busy run used more tools, but it still stayed inside the evidence boundary.
Context reuse is the useful signal.
Most of the context was reused instead of rebuilt. The harness provided stable retrieved context, so Claude could keep answering from cached session material instead of re-reading the world each turn.
The simple score
A good run gets more useful with each turn. It cites where facts came from, admits what is missing, avoids fake certainty, and does not keep repeating the same search unless new evidence appears.