Anatomy of Autonomous Coding Agents: What 7 System Prompts Reveal About the Industry
Evaluate what seven public coding-agent prompt surfaces reveal about harness structure without mistaking visible prompts for complete vendor architectures.
Evaluate what seven public coding-agent prompt surfaces reveal about harness structure without mistaking visible prompts for complete vendor architectures.
- Edition
- Comparative edition 1.0
- Published
- Feb 22, 2026
- Preview updated
- Aug 27, 2026
- Reading time
- 14 minutes
- Full edition
- 3,200 words
- Status
- historical
Literature Synthesis
- Research question or engineering problem
- What common harness pattern appears across seven public coding-agent prompt snapshots?
- Principal finding
- The visible prompt surfaces converge on tool-rich single-agent loops more often than on enforceable multi-role authority separation.
- Evidence type
- Public prompt corpus review.
- Method summary
- Compare public prompt and tool snapshots using a common harness rubric, then separate observed text from architectural inference.
- Scope
- Seven public snapshots available in February 2026, not vendor internals or current product behavior.
- Limitations
- Public prompt collections may be incomplete, stale, or decontextualized.
- No claim is made about undisclosed vendor systems.
- Public source or reproduction note
- Public source note
- Published
- 2026-02-22
- Last verified
- 2026-08-27
- Status
- historical
Who this is for
- Agent-tool builders comparing planning, execution, and review boundaries.
- Technical evaluators who need a reproducible comparison rubric.
- Readers studying the difference between one-loop and separated-role harnesses.
Not for
- Readers seeking undisclosed vendor internals, a current product ranking, or a claim that prompt snapshots reveal every runtime behavior.
Detailed contents
- 01
Public corpus and selection boundary
- 02
The common harness template
- 03
Tool-by-tool comparison
- 04
The mono-brain problem
- 05
Separated roles and review
- 06
Comparison matrix
- 07
Observed patterns and inferences
- 08
Limitations and dated conclusions
Named artifacts
Public source ledger
The dated prompt corpus and source-note boundary used for comparison.
Harness comparison rubric
Identity, tools, authority, state, review, and recovery fields.
Seven-system comparison matrix
A normalized view of the public surfaces examined.
Authority-boundary map
Where one loop combines duties and where roles are separated.
Prompt snapshots reveal harness shape, not the whole product
A public system prompt can reveal the identity assigned to a model, the tools it can call, visible behavioral constraints, output rules, and some task-management structure. That is enough to compare parts of a harness. It is not enough to infer every hidden runtime service, evaluation loop, policy layer, or product behavior.
The comparison therefore scores only visible structure. Does the surface separate planning from execution? Can review be independent from the action loop? Is durable state named? Are tool permissions explicit? Is recovery a first-class phase or merely another instruction to the same context window? The answers describe the inspected snapshot, not an eternal vendor ranking.
The principal inference is bounded: the sampled public surfaces converged on a common skeleton of one model, one working context, tools, constraints, and self-review. Greyforge's response was to test whether conflicting duties should live behind separate authority contracts. That response is a position drawn from the comparison, not proof that every multi-role system is more reliable.
Harness comparison rubric
- Identity
- What role and operating objective does the visible prompt declare?
- Tools
- Which read, write, terminal, browser, and external actions are exposed?
- Authority
- Are planning, implementation, review, and publication separated?
- State
- How does the visible harness carry task and durable context?
- Verification
- Can evidence block completion, or does the same loop judge itself?
- Recovery
- What forces a new diagnosis after repeated failure?
Evidence and method
literature synthesis: Dated comparison of seven visible public prompt surfaces using a common harness rubric. Facts from the source corpus are separated from Greyforge inferences and positions.
Limitations
- Public prompt snapshots do not expose every hidden vendor runtime or policy layer.
- The corpus is dated and does not establish a current product ranking.
- Architectural implications are Greyforge inferences, not measured reliability results.
Access and updates
Purchase includes lifetime read access to this edition, email-based recovery, and revisions published to the same edition.
Public companion: Open the public comparison source note