Branch

WRITER.md

Human Duke

Everyone agrees that LLMs are pretty terrible at Essay writing. They use a repetitive and easy to spot style. They don't usually say anything interesting either.

We've seen that agent harnesses like Claude Code and Codex have given software engineers super powers, but there's nothing like this yet for essay writers (or writers in general).

What sort of tooling would give writers super powers?

How does it look different from the work software engineers do?

How can writers retain their unique style and voice? Or how can the tooling help them develop one?

AI · author-edited OpenAI

The equivalent of Claude Code for writers is probably not “a model that writes the essay.” It is a research, argument, and revision environment in which the writer remains the source of judgment.

Coding agents became useful when they gained a harness: access to the repository, search, tests, version control, execution, and a loop for checking their work. A writing agent needs analogous machinery—but prose has no compiler, and “good” is not a single measurable state.

What a writer’s harness would contain

1. A living model of the project

Not just a chat transcript or a folder of documents. The system should maintain:

  • the intended reader
  • the central question
  • candidate theses
  • claims and supporting evidence
  • objections and unresolved tensions
  • sources, quotations, and provenance
  • definitions of important terms
  • examples, scenes, and anecdotes
  • structural outlines
  • stylistic constraints
  • decisions the writer has made and rejected

This would be closer to a graph than a document. An essay is merely one rendering of that graph. WRITER.md You are here

This branch begins here WRITER.md
Human Duke

This would function kind of like AGENTS.md/CLAUDE.md for writing projects. What would an example of that look like?

AI OpenAI

A writing project could have a WRITER.md: part creative brief, part argument map, part style guide, and part operating manual for agents.

Crucially, it should not be a prompt saying “write beautifully.” It should record specific decisions, constraints, and preferences that help an agent collaborate without flattening the author’s voice.

Here’s an example for an essay about AI tools for writers:

# WRITER.md

## Project

Working title: The Writer’s Harness

Form: Argumentative essay with reported examples
Target length: 2,500–3,500 words
Likely venue: General-interest technology magazine
Status: Research and thesis development

## Governing Question

Coding agents became useful when models were placed inside a harness with
repository access, tools, tests, and feedback loops.

What would an equivalent harness for essayists look like, given that prose
has no compiler and quality is contested?

## Reader

The primary reader:

- knows what ChatGPT and Claude are
- may have used AI for coding or writing
- is skeptical of AI-generated prose
- values human originality but is open to useful tools
- does not need transformer mechanics explained

Assume intelligence, not prior agreement.

## Current Thesis

The useful AI writing system will not be an automatic essay generator.
It will be an environment that helps a writer gather evidence, develop
arguments, test structure, and revise deliberately while leaving aesthetic
and intellectual judgment with the writer.

This thesis is provisional. Challenge it if the evidence points elsewhere.

## What the Essay Is Not Arguing

- That writing and programming are the same activity
- That every writer should use AI
- That style can be reduced to measurable rules
- That faster production necessarily produces better writing
- That current models can reliably judge originality
- That AI should replace editors, sources, or fact-checkers

Do not quietly strengthen the thesis into any of these claims.

## Stakes

The essay should explain why current AI writing disappoints:

1. Chat interfaces encourage one-shot generation.
2. Models optimize for plausible continuation, producing conventional prose.
3. The system lacks durable knowledge of the writer’s project.
4. There is no strong equivalent to tests or program execution.
5. Delegating composition too early also delegates discovery and judgment.

The larger question is whether tools can augment thought without standardizing it.

## Argument Map

### Claim 1: Coding agents succeeded because of the harness

Evidence needed:

- examples of repository search, test execution, linting, and version control
- evidence that agent performance improves with environmental feedback
- comments from engineers about how their workflow changed

Possible objection:

- coding models also simply became more capable

Response:

- acknowledge capability improvements
- argue that capability became practically useful through tool access and feedback

### Claim 2: Writing requires a different harness

Important differences:

- there is rarely a single correct output
- goals change during composition
- writing often discovers its subject
- factual validity and aesthetic success are separate
- reader response cannot be simulated reliably
- disagreement among evaluators can be meaningful rather than noise

### Claim 3: The best tools operate around the act of composition

Candidate capabilities:

- source ingestion and provenance tracking
- claim-to-evidence mapping
- contradiction and objection discovery
- outline comparison
- retrieval from the writer’s archive
- sentence-level revision with explicit rationales
- voice-drift detection
- reader simulations, clearly labeled as speculative
- reversible edits and version comparison

### Claim 4: Voice is preserved through authorship of decisions

Voice is not just sentence rhythm. It includes:

- what the writer notices
- which examples they choose
- where they qualify a claim
- what they find funny or offensive
- which tensions they leave unresolved
- how much certainty they claim
- the sequence in which ideas are revealed

The tool should learn these preferences without treating them as a formula.

## Open Questions

- What are the writing equivalents of unit tests?
- Which judgments can be automated, and which should remain visibly subjective?
- Does “voice detection” preserve voice or cause self-parody?
- How should the system represent uncertainty and disagreement between sources?
- When does assistance become ghostwriting?
- Can a harness help a writer become more surprising rather than more consistent?

Do not resolve these prematurely.

## Source Policy

For every factual claim, retain:

- source title
- author
- publication
- URL or stable identifier
- publication date
- date accessed
- exact supporting passage
- location within the source
- notes distinguishing quotation, paraphrase, and inference

Never invent bibliographic details or quotations.

Mark unsupported claims as `[SOURCE NEEDED]`.

Prefer primary sources, direct observation, and named interviews over summaries.
Do not cite an LLM-generated answer as evidence.

## Evidence Ledger

| Claim | Evidence | Status | Notes |
|---|---|---:|---|
| Harnesses improved coding-agent utility | Benchmark and practitioner evidence | Needed | Separate model gains from harness gains |
| Writers dislike generic AI prose | Existing criticism plus examples | Partial | Avoid “everyone agrees” |
| Writing lacks objective tests | Argument, not simple fact | Drafted | Discuss fact-checking as a partial test |
| Voice exceeds surface style | Scholarship/interviews | Needed | Look at rhetoric and literary studies |

## Counterarguments to Take Seriously

1. A sufficiently capable model may not need a complex harness.
2. Human editors already provide most of these functions.
3. Constraints and voice profiles may make writers more repetitive.
4. AI research assistance can contaminate a project with fabricated facts.
5. Many writers want competent generic prose, not literary distinction.
6. “Keeping the human in control” may be branding rather than a real boundary.

Represent these in their strongest form. Do not include objections merely to
dismiss them in one sentence.

## Structure Under Consideration

Do not treat this as fixed.

1. Open with the contrast: coding agents feel transformative; generated essays do not.
2. Explain that the difference is partly the harness, not just the model.
3. Show why a direct translation from coding fails.
4. Walk through the components of a writer’s harness.
5. Examine voice, judgment, and the danger of standardization.
6. End with the idea that the target is not frictionless writing but more
   productive friction.

Avoid beginning with a broad history of writing technology.

## Voice

Desired qualities:

- curious but not breathless
- precise without sounding academic
- skeptical without becoming cynical
- concrete before abstract
- willing to make a claim, then define its limits
- occasional dry humor
- varied sentence length
- first person only when it records a real observation or judgment

The prose should sound like one person thinking carefully, not an institution
issuing a report.

## Voice Examples

Canonical samples are in:

- `voice/sample-01.md`
- `voice/sample-02.md`
- `voice/sample-03.md`

When proposing prose, retrieve comparable passages from these samples.
Identify relevant traits, but do not copy phrases.

Prefer the samples over generic style advice in this file.

## Anti-Style

Avoid:

- “In today’s rapidly evolving landscape”
- “It’s important to note”
- “This isn’t just X; it’s Y”
- repetitive triads
- excessive em dashes
- fake quotations or imaginary scenes
- rhetorical questions used as transitions
- announcing that something is “profound”
- calling every change a “paradigm shift”
- symmetrical “on the one hand/on the other hand” treatment
- concluding by repeating the introduction
- headings for every minor point

Do not use “delve,” “tapestry,” “multifaceted,” “leverage,” or “game-changer”
unless quoting someone.

No paragraph should exist solely to summarize the preceding paragraph.

## Collaboration Rules

The writer owns:

- the thesis
- selection of examples
- moral and aesthetic judgments
- final wording
- decisions about unresolved ambiguity

The agent may:

- search project materials
- propose competing theses
- expose assumptions
- locate missing evidence
- produce structural alternatives
- identify repetition or voice drift
- suggest local edits
- ask questions that force a decision

The agent should not draft complete sections unless explicitly asked.

When suggesting a substantial change:

1. identify the problem
2. quote or point to the affected passage
3. offer two or more options when reasonable
4. explain the tradeoff
5. preserve the original in version history

Distinguish clearly among:

- factual correction
- logical criticism
- structural suggestion
- stylistic preference
- speculative reader reaction

Never present the last two as objective defects.

## Default Workflow

When asked to “help with the essay”:

1. Read this file and the current draft.
2. Inspect `decisions.md` and `open-questions.md`.
3. Summarize the draft’s actual argument, not its intended argument.
4. Identify the highest-leverage unresolved issue.
5. Ask whether the writer wants research, argument, structure, or prose help.
6. Make the smallest intervention that addresses that issue.
7. Record accepted decisions; do not record rejected suggestions as policy.

Do not respond by rewriting the whole essay.

## Revision Passes

Keep these passes separate:

### Factual pass
Check names, dates, quotations, citations, and claim support.

### Argument pass
Check inference, assumptions, contradictions, counterarguments, and scope.

### Structural pass
Check sequence, pacing, section function, and delayed context.

### Prose pass
Check clarity, rhythm, abstraction, repetition, and unnecessary qualifiers.

### Voice pass
Compare against canonical samples and flag deviations with evidence.

Do not polish sentences that may be removed for structural reasons.

## Quality Checks

Before calling a draft complete, ask:

- Can the thesis be stated without jargon?
- Does each major section alter or advance the argument?
- Is every strong factual claim supported?
- Is the strongest counterargument represented fairly?
- Are examples doing intellectual work rather than decorating the prose?
- Does the ending produce a consequence rather than a recap?
- Are there passages that could have appeared in almost any AI essay?
- Has fluency concealed uncertainty?
- Did the writer make the important decisions?

## Current Decisions

- Use “harness” as the central metaphor, but acknowledge its limits.
- Focus on essayists rather than fiction writers.
- Include at least one working prototype or concrete interface.
- Avoid claiming that objective evaluation is entirely absent in writing.
- Do not frame the essay as a product pitch.

## Parking Lot

Potential ideas not yet earned by the argument:

- writing as an IDE
- an “argument debugger”
- reader personas
- taste as a test suite
- editors as integration tests
- personal archives as context repositories

Do not insert these merely because they are memorable.

Better as a small hierarchy

Like software instructions, this would work best when split by scope rather than crammed into one enormous file:

WRITER.md                 # Durable collaboration and project rules
project/
  brief.md                # Reader, form, stakes, thesis
  argument-map.md         # Claims, evidence, objections
  decisions.md            # Accepted decisions and their reasons
  open-questions.md       # Live uncertainties
  sources/
    ledger.md
    papers/
    interviews/
  voice/
    README.md              # Specific voice observations
    sample-01.md           # Author-approved exemplars
    sample-02.md
  drafts/
    001.md
    002.md
  feedback/
    editor-a.md
    reader-notes.md

The distinction between files matters:

  • WRITER.md tells the agent how to behave.
  • The brief says what this particular work is trying to accomplish.
  • The argument map represents its intellectual state.
  • The decision log prevents the agent from repeatedly reopening settled questions.
  • Voice samples provide stronger evidence than adjectives like “witty” or “incisive.”
  • Draft history keeps experimentation reversible.

The closest equivalent to CLAUDE.md is therefore not merely a style guide. It is a compact constitution for the collaboration: what the project believes, what remains unsettled, what the agent is authorized to do, and which judgments must remain with the writer.

Explore conversation