C runs a deterministic root rule. It is not the same model plus Minority Prophet, so this page does not measure MP lift.
METHOD COMPARISON
Same packet.
Different methods.
Models, tool-using agents, conventional voting, and a deterministic evidence rule faced the same frozen packet.
Keep every comparison inspectable.
- Freeze the protocol first.
- Share the same packet.
- Preserve failures and raw scores.
- Keep cost and time attached to each run.
- Repeat before generalizing.
READ THIS FIRST
Eight cases.
Not 128 trials.
Each case contains 16 related decisions. The result shows rule execution, not 128 independent replications.
This tests reasoning over declared lineage—not discovery of hidden copying or proof of external truth.
01 / WHY THIS JUNCTION MATTERS
Agents will talk
faster than humans can check.
Claims, receipts, delegated tasks, and proposed actions can cross agent boundaries continuously. This exercise asks whether a known lineage rule belongs in transparent code instead of being re-inferred each time.
A deterministic check handles the known invariant. Judgment stays reserved for what remains uncertain.
The methods ran through different execution paths: local deterministic code versus subscription-backed model CLI calls. The ratio describes this run only. It must not be linearly projected to production throughput, latency, or cost.
Concurrency, hardware, networks, provider behavior, and packet shape can all change the comparison.
Use expensive probabilistic intelligence where judgment is needed. Use deterministic code where the invariant is already known. At scale, that separation keeps the agent network fast without turning uncertainty into permission.
02 / OBSERVED TELEMETRY BY LANE
No combined score.
Every run stands alone.
Each amount below belongs to one contestant lane running the full eight-case packet. Cost, accuracy, and time stay attached to that exact run; they do not establish stable provider rankings.
Minority Prophet
Canonical v1128/128protocol score
18.7 mswall time
$0 modelmodel cost
GPT-5.6 Terra
OpenAI / Codex116/128protocol score
6.1 minwall time
≈ $0.89run estimate
68/128protocol score
4.1 minwall time
≈ $1.02run estimate
Claude Opus 5
Anthropic / Claude Code106/128protocol score
8.0 minwall time
≈ $3.25run estimate
0/12896/128 raw
6.1 minwall time
≈ $4.32run estimate
GPT-5.6 Sol
OpenAI / Codex102/128protocol score
8.5 minwall time
≈ $2.24run estimate
69/128protocol score
5.9 minwall time
≈ $2.42run estimate
Cluster vote
Conventional baseline96/128protocol score
9.5 mswall time
$0 modelmodel cost
Claude Sonnet 5
Anthropic / Claude Code23/128protocol score
13.5 minwall time
≈ $3.47run estimate
32/12874/128 raw
26.8 minwall time
≈ $6.45run estimate
GPT-5.6 Luna
OpenAI / Codex16/128protocol score
4.9 minwall time
≈ $0.37run estimate
35/128protocol score
8.3 minwall time
≈ $0.81run estimate
Claude Haiku 4.5
Anthropic / Claude Code0/128protocol score
17.6 minwall time
≈ $1.15run estimate
0/12810/128 raw
35.9 minwall time
≈ $1.92run estimate
Costs are per run, never combined. GPT uses API list-price proxies; Claude uses CLI-reported estimates. Haiku B excludes one timed-out attempt with no completed usage record.
03 / THE TEST
One input.
Different capabilities.
AI reasons alone
The model receives the complete raw packet, including immediate parent links. Shell, files, web, retrieval, and every other tool are disabled.
The same AI may use tools
The same model receives the identical complete-lineage packet. It may choose shell, scripts, packages, or web tools. It is not told to use Minority Prophet.
Canonical Minority Prophet
The raw packet enters standalone deterministic code—not the same model plus MP output. It follows parent links, counts distinct roots, and abstains on exact ties.
DESCRIPTIVE RESULT TABLE
Observed scores, not stable ranks
One clean replicate · eight cases · correlated within-case dispositions
| Contestant | Lane | Protocol score | Raw answers | Exact cases | Invalid trials | Wall time | Tools | Input / output tokens | Cost estimate |
|---|---|---|---|---|---|---|---|---|---|
| Minority ProphetCanonical v1 | Lane C | 128/128 | 128/128 | 8/8 | 0/8 | 18.7 ms | 0 | 0 / 0 | $0 model cost |
| GPT-5.6 TerraOpenAI / Codex | Lane A | 116/128 | 116/128 | 5/8 | 0/8 | 6.1 min | 0 | 300,475 / 16,746 | $0.890 |
| Claude Opus 5Anthropic / Claude Code | Lane A | 106/128 | 106/128 | 6/8 | 0/8 | 8.0 min | 0 | 326,569 / 41,681 | $3.253 |
| GPT-5.6 SolOpenAI / Codex | Lane A | 102/128 | 102/128 | 5/8 | 0/8 | 8.5 min | 0 | 300,694 / 25,952 | $2.237 |
| Cluster voteConventional baseline | Standard | 96/128 | 96/128 | 6/8 | 0/8 | 9.5 ms | 0 | 0 / 0 | $0 model cost |
| GPT-5.6 SolOpenAI / Codex | Lane B | 69/128 | 69/128 | 4/8 | 0/8 | 5.9 min | 19 | 1,101,609 / 12,559 | $2.418 |
| GPT-5.6 TerraOpenAI / Codex | Lane B | 68/128 | 68/128 | 4/8 | 0/8 | 4.1 min | 18 | 1,097,326 / 9,151 | $1.021 |
| GPT-5.6 LunaOpenAI / Codex | Lane B | 35/128 | 35/128 | 2/8 | 0/8 | 8.3 min | 41 | 1,933,985 / 20,412 | $0.809 |
| Claude Sonnet 5Anthropic / Claude Code | Lane B | 32/128 | 74/128 | 2/8 (4 raw) | 5/8 | 26.8 min | 27 | 2,431,595 / 199,228 | $6.454 |
| Claude Sonnet 5Anthropic / Claude Code | Lane A | 23/128 | 23/128 | 0/8 | 0/8 | 13.5 min | 0 | 639,869 / 82,258 | $3.466 |
| GPT-5.6 LunaOpenAI / Codex | Lane A | 16/128 | 16/128 | 1/8 | 0/8 | 4.9 min | 0 | 289,024 / 13,697 | $0.371 |
| Claude Opus 5Anthropic / Claude Code | Lane B | 0/128 | 96/128 | 0/8 (6 raw) | 8/8 | 6.1 min | 35 | 1,489,495 / 19,965 | $4.316 |
| Claude Haiku 4.5Anthropic / Claude Code | Lane A | 0/128 | 0/128 | 0/8 | 0/8 | 17.6 min | 0 | 218,476 / 98,985 | $1.146 |
| Claude Haiku 4.5Anthropic / Claude Code | Lane B | 0/128 | 10/128 | 0/8 | 8/8 | 35.9 min | 59 | 3,145,023 / 145,142 | $1.917 |
Lane B means tools were available, not necessarily used. Protocol score penalizes failed or boundary-violating runs; raw answers preserve answer accuracy before that penalty. With one replicate, differences remain descriptive.
SPEED COMPARISON
How long each run took
Wall time · shorter is faster · logarithmic bars
Descriptive subscription-CLI wall time, including provider and harness overhead. This is not a controlled API-serving latency benchmark.
Reasoning only
GPT-5.6 TerraOpenAI / Codex
116/1286.1 min · 0 toolsClaude Opus 5Anthropic / Claude Code
106/1288.0 min · 0 toolsGPT-5.6 SolOpenAI / Codex
102/1288.5 min · 0 toolsClaude Sonnet 5Anthropic / Claude Code
23/12813.5 min · 0 toolsGPT-5.6 LunaOpenAI / Codex
16/1284.9 min · 0 toolsClaude Haiku 4.5Anthropic / Claude Code
0/12817.6 min · 0 toolsTools available
GPT-5.6 SolOpenAI / Codex
69/1285.9 min · 19 toolsGPT-5.6 TerraOpenAI / Codex
68/1284.1 min · 18 toolsGPT-5.6 LunaOpenAI / Codex
35/1288.3 min · 41 toolsClaude Sonnet 5Anthropic / Claude Code
32/12826.8 min · 27 tools · 5 invalidClaude Opus 5Anthropic / Claude Code
0/1286.1 min · 35 tools · 8 invalidClaude Haiku 4.5Anthropic / Claude Code
0/12835.9 min · 59 tools · 8 invalidDeterministic root vote
Minority ProphetCanonical v1
128/12818.7 ms · 0 toolsEvery lane traced its own roots.
All lanes received raw records and immediate parent links. None received an answer key, root map, root count, or precomputed root IDs.
This measures method conformance.
It tests a known distinct-origin rule under complete lineage. It does not prove real-world roots are honest, independent, current, authorized, or true.
For same-model MP lift, open the lift study →