Skip to main content
Paid Chronicle · Public Preview

Anatomy of Autonomous Coding Agents: What 7 System Prompts Reveal About the Industry

Evaluate what seven public coding-agent prompt surfaces reveal about harness structure without mistaking visible prompts for complete vendor architectures.

Outcome

Evaluate what seven public coding-agent prompt surfaces reveal about harness structure without mistaking visible prompts for complete vendor architectures.

Edition
Comparative edition 1.0
Published
Feb 22, 2026
Preview updated
Aug 27, 2026
Reading time
14 minutes
Full edition
3,200 words
Status
historical
Research disclosureinferredhistorical

Literature Synthesis

Research question or engineering problem
What common harness pattern appears across seven public coding-agent prompt snapshots?
Principal finding
The visible prompt surfaces converge on tool-rich single-agent loops more often than on enforceable multi-role authority separation.
Evidence type
Public prompt corpus review.
Method summary
Compare public prompt and tool snapshots using a common harness rubric, then separate observed text from architectural inference.
Scope
Seven public snapshots available in February 2026, not vendor internals or current product behavior.
Limitations
  • Public prompt collections may be incomplete, stale, or decontextualized.
  • No claim is made about undisclosed vendor systems.
Public source or reproduction note
Public source note
Published
2026-02-22
Last verified
2026-08-27
Status
historical

Who this is for

  • Agent-tool builders comparing planning, execution, and review boundaries.
  • Technical evaluators who need a reproducible comparison rubric.
  • Readers studying the difference between one-loop and separated-role harnesses.

Not for

  • Readers seeking undisclosed vendor internals, a current product ranking, or a claim that prompt snapshots reveal every runtime behavior.

Detailed contents

  1. 01

    Public corpus and selection boundary

  2. 02

    The common harness template

  3. 03

    Tool-by-tool comparison

  4. 04

    The mono-brain problem

  5. 05

    Separated roles and review

  6. 06

    Comparison matrix

  7. 07

    Observed patterns and inferences

  8. 08

    Limitations and dated conclusions

Named artifacts

  • Public source ledger

    The dated prompt corpus and source-note boundary used for comparison.

  • Harness comparison rubric

    Identity, tools, authority, state, review, and recovery fields.

  • Seven-system comparison matrix

    A normalized view of the public surfaces examined.

  • Authority-boundary map

    Where one loop combines duties and where roles are separated.

Substantive sample · Complete section

Prompt snapshots reveal harness shape, not the whole product

A public system prompt can reveal the identity assigned to a model, the tools it can call, visible behavioral constraints, output rules, and some task-management structure. That is enough to compare parts of a harness. It is not enough to infer every hidden runtime service, evaluation loop, policy layer, or product behavior.

The comparison therefore scores only visible structure. Does the surface separate planning from execution? Can review be independent from the action loop? Is durable state named? Are tool permissions explicit? Is recovery a first-class phase or merely another instruction to the same context window? The answers describe the inspected snapshot, not an eternal vendor ranking.

The principal inference is bounded: the sampled public surfaces converged on a common skeleton of one model, one working context, tools, constraints, and self-review. Greyforge's response was to test whether conflicting duties should live behind separate authority contracts. That response is a position drawn from the comparison, not proof that every multi-role system is more reliable.

Harness comparison rubric

Identity
What role and operating objective does the visible prompt declare?
Tools
Which read, write, terminal, browser, and external actions are exposed?
Authority
Are planning, implementation, review, and publication separated?
State
How does the visible harness carry task and durable context?
Verification
Can evidence block completion, or does the same loop judge itself?
Recovery
What forces a new diagnosis after repeated failure?

Evidence and method

literature synthesis: Dated comparison of seven visible public prompt surfaces using a common harness rubric. Facts from the source corpus are separated from Greyforge inferences and positions.

Limitations

  • Public prompt snapshots do not expose every hidden vendor runtime or policy layer.
  • The corpus is dated and does not establish a current product ranking.
  • Architectural implications are Greyforge inferences, not measured reliability results.

Access and updates

Purchase includes lifetime read access to this edition, email-based recovery, and revisions published to the same edition.

Public companion: Open the public comparison source note