I expected the orchestrator to win, or at worst to cost more for nothing. It lost on both tasks that separated the two arms, and I can point to the sentence behind each loss.

Both arms ran the same model, Opus 5.5, with the same tools: Bash, Edit, Read, Write, Grep and Glob. The orchestrator had one extra tool, which let it spawn subagents and hand them a piece of the task. Hidden tests judged every task. No model graded another model’s work.

The mechanism

A subagent never sees the original spec. It sees the orchestrator’s summary of it, a message of prose and Bash commands written under time pressure by the same kind of model that will carry it out. The subagent follows that summary, mistakes and all.

Three moments in the transcripts show it, each specific enough to rule out noise.

Moment 1

An invented rule

The task was to implement a method that pulls one field out of a structured pandas column. The spec is precise about naming. For a string, bytes or integer selector, the result takes the selected field’s name. The spec says nothing about decoding.

Orchestrator to worker, the message as sent (excerpt)

Name rules: for str/bytes/int name is the selected field's name
(for int, pa_type.field(i).name; for str, the str; for bytes, decode)

That parenthetical, “for bytes, decode,” appears nowhere in the spec. The worker built it anyway:

The worker’s patch

if isinstance(name, bytes):
    name = name.decode()

ResultThe hidden test test_struct_accessor_field_expanded[string_col-string_col] expects the raw bytes value b"string_col" back. The run fails 11 of 12. Both solo runs on this task left bytes alone, matched the spec and passed 12 of 12.

Moment 2

A paraphrased algorithm

Same task, another run. This time the orchestrator skipped the rule and wrote the implementation out as pseudocode for the worker to transcribe:

Orchestrator to worker, the message as sent (excerpt)

elif is_list_like: reverse the list, iterate, popping from the reversed
list: name=get_name(popped, data); data = pc.struct_field(data,
[popped_index_or_name]) ... i.e. walk nested levels, returning the
final field name

In the branch that handles a list of nested selectors, nothing ever sets name to the final field’s real name. The orchestrator dropped that step somewhere between the upstream logic and this prose. The worker copied it line for line, bug included.

ResultFails test_struct_accessor_field_expanded[indices5-string_col], 11 of 12. Both solo runs and the reference implementation compute the name from the resolved field object, and all three pass.

Two reps on the same task produced two different paraphrase bugs. Solo won this task 2 of 2. Orchestration won it 0 of 2.

Moment 3

The handoff that never returned

A different task failed a different way. The orchestrator asked its worker to verify the change by “running the existing test suites” for the area it touched. The worker escalated on its own, from a few relevant test files to the whole pandas test suite:

The worker’s own commands, in order

pytest pandas/tests/groupby pandas/tests/test_sorting.py pandas/tests/resample -q -x -n 8
...
timeout 3000 python -m pytest pandas/tests -q -n 8 -p no:cacheprovider -m "not slow..."
timeout 3000 python -m pytest pandas/tests -q -n 8 -p no:cacheprovider -m "not slow..."   (again)
timeout 3000 python -m pytest pandas/tests -q -n 8 -p no:cacheprovider -m "not slow..."   (again)
timeout 600  python -m pytest pandas/tests -q -n 8 -p no:cacheprovider -m "not slow..."
timeout 900  python -m pytest pandas/tests --collect-only -q
for d in $(ls -d pandas/tests/*/ ) pandas/tests/test_*.py; do
  timeout 1500 python -m pytest $d -q -n 8 ... ; done

That makes ten escalating full-suite runs under emulation, each eating minutes. The run hit its 45-minute cap with the worker still mid-loop, and the orchestrator never got control back to report anything.

This run also exposed a bug in my scoring. The harness recorded it as a pass, because it salvaged a code patch from disk after the timeout. My rules already counted a timeout as a fail, so I corrected the record. The orchestrator’s score on this task drops from an apparent 2 of 2 to 1 of 2.

What didn’t move

On 6 of the 10 tasks, solo and orchestrator landed on the same pass or fail: the same missing code, the same under-scoped fix, however the work was split. Two of those tasks were unfair to both arms. One had a masked spec that also stripped unrelated functions the fix depended on. The other’s description named the wrong class for the missing piece, and every run in both arms hit that wall.

Struct field selector
Solo 2/2 · orchestrator 0/2 · lost to delegation, moments 1 and 2
Groupby quantile
Solo 2/2 · orchestrator 1/2 · lost to delegation, moment 3
6 other tasks
Same outcome in both arms · a shared cause
2 remaining tasks
Mixed in both arms · inconclusive at two reps

Hidden-test passes per arm, two reps each. Ten tasks in all.

Token cost came out about even. The orchestrators mostly delegated once per task, so they duplicated little context. On the clock, orchestration ran slower on most tasks. Nothing on the efficiency side makes up for the lost passes.

How far this reaches

This is a small run. Two of the ten tasks separated the arms, and both came from the same codebase. I won’t put a percentage on it, and “orchestration loses by X%” is no part of the claim.

The claim is the mechanism. It held up when I read the transcripts, as well as on the scoreboard.

Splitting a task adds a retelling of the spec, and each retelling gives the instruction a chance to drift from what was asked. Nobody sees the drift until a test or a person catches it.

The hidden tests come from FeatureBench. Both arms ran Opus 5.5 with the same tool access. The three failures come from two tasks, so read them as a mechanism and draw no statistic from them.