---
name: deep-research-stored
description: >-
  Run deep research so that it is stored rather than only reported: the folder is the deliverable, a red team argues against the answer before anything is synthesized, and the report is written out of the folder afterward. Use when research findings need to outlive the session, when a reader will need to check a single claim without re-running the study, when a stale figure should be fixable by re-fetching one known page, or when a decision will rest on the answer and it has to survive somebody arguing against it. Also use when asked for research that is saved, filed, or kept rather than delivered as prose.
---

# Deep research, stored

Ordinary research treats the report as the deliverable, so whatever files exist are a
byproduct. The evidence dies with the session and the report has to be trusted whole.

**This inverts it.** The structure is built before anything runs, the agents write into it,
and the report is derived from it afterward. A reader can then check one claim without
repeating the work, and updating a stale figure means re-fetching one known page rather
than redoing the study.

## The topic is whatever you were given

If this was invoked with no subject, ask for one rather than guessing. Everything below is
about how rather than what.

## The steps

Give these to whoever does the work. **When you fan out to subagents, put this whole block
in each agent's prompt rather than summarizing it**, because an agent working from your
summary of the method is the exact failure the method exists to prevent.

0. **No file you author may exceed 5,000 characters.** This covers every note, table,
   synthesis and index. It does not cover saved sources, which are primary material and are
   kept whole. When a note outgrows the limit, split it at a heading boundary into
   `<name>--01-<slug>.md`, `<name>--02-<slug>.md` and so on, never mid-sentence, and leave
   the parent as a short pointer listing the parts with one line each.

   **The reason is retrieval rather than tidiness.** Anything that searches a folder and
   returns whole files works against a fixed context budget. One 110,000-character report
   consumes most of that budget by itself and crowds out everything else that matched, so a
   folder of large files retrieves badly however well written it is.

1. **Write down the decision the research serves, then build the directory structure, both
   before you run anything.** A short brief at the top of the folder says what decision the
   answer will feed and what tests the answer will be judged by. Below it go one numbered
   subdirectory per part of the question, a `sources/` tree mirroring them, and a
   `_synthesis/` directory. Every note then says which part of the question it bears on and
   whether it SUPPORTS or UNDERMINES the working answer. The red team in step 6 attacks the
   answer to that decision, so without a stated decision it has nothing to attack.

2. **Save the full text of every substantive source**, one file per source, named
   `<slug>--<domain>--<date>.md`, with a header giving the URL, the title, the date fetched
   and one line on who wrote it. **Save it before you write about it.** Assume you will
   remember nothing.

3. **Fan out by discovery method, not only by topic.** Where a question could be answered
   by several different routes, run one agent per route rather than one per subject. Each
   route has a different blind spot, and the overlap between them is itself evidence. A
   fan-out organized by topic gives every agent the same blind spot applied to a different
   subject.

   **Tell every route to hunt for evidence against the working answer as well as for it**,
   aiming part of its searches there and logging which searches were aimed that way. An
   agent told only what to find will find it and nothing else, and the red team in step 6
   can only argue from what the routes brought back.

4. **Mark every figure VERIFIED, SELF-REPORTED or ESTIMATED**, and carry the marking as a
   field in the data rather than as a word in a sentence, so it survives being summarized.
   Every figure carries its URL and the date checked.

5. **Have a separate agent, one that did not do the finding, open the source of every
   numeric claim and of every claim a conclusion rests on, and try to refute it.** Anything
   unconfirmed gets corrected or dropped rather than repeated. **Choose the claims by what a
   conclusion rests on, never by what is easiest to check**, because prices and counts are
   the easiest things to verify and are often the ones no decision depends on.

   **Grade the interpretation separately from the words.** One verdict says whether the
   source says what the claim says, as CONFIRMED, CORRECTED, REFUTED or UNVERIFIABLE. A
   second says whether the source supports what the finder concluded from it, as WARRANTED,
   OVERREACH or UNSUPPORTED, because a quotation can be exact and still not mean what the
   finder says it means.

   **Give each verifier a small slice, roughly fifteen claims rather than sixty**, and have
   it write verdicts to a file as well as returning them. Both halves are about visibility
   rather than accuracy: an agent emits nothing until its whole structured object lands, so
   a large slice produces minutes of silence indistinguishable from a dead agent.

6. **Red team the answer before anybody synthesizes it.** After verification, run agents
   that did none of the finding.

   - **One red team per route.** Each is given the working answer, taken from the brief or
     from what the study is leaning towards, together with that route's notes and verdicts,
     which it reads whole. It builds the strongest honest case that the working answer is
     wrong, using only evidence in those files, citing the file behind each point, and
     saying what evidence would overturn each point. It writes
     `_synthesis/redteam--<route>.md`. **One short merge** then reads the per-route files
     and writes `_synthesis/redteam.md` with the strongest point first.
   - **A completeness critic** reads the brief, the verdict files and the per-route red-team
     files, and says what is missing rather than what is wrong. It reads from three or four
     named perspectives chosen for the question, such as the person who will act on the
     answer, the person who has to approve it, and somebody with money at stake, and for
     each gap it says what evidence would settle it. It writes `_synthesis/completeness.md`.

   **The red team is split by route because one agent cannot read a whole study.** A large
   study runs to hundreds of notes and hundreds of thousands of characters outside
   `sources/`, so a single red team told to read everything would sample, and a case built
   from a sample is exactly the partial read this method forbids.

   None of these agents may invent anything, so a point with no file behind it is dropped.
   **This is the step that catches a well-evidenced answer to a nearby question.** A study
   can save hundreds of sources and pass every checked claim while still answering the
   wrong thing, because every agent before this one was told to establish facts and none was
   told to want the answer to be wrong.

7. **The synthesizing agent reads the files, not the other agents' summaries**, and is told
   that where a note and a summary disagree the note wins, because the note was written next
   to the source. **Pass it every verdict, the passes as well as the failures.** Forward only
   the failures and the synthesizer will mark every independently confirmed claim as never
   checked, because an agent cannot report a confirmation it was never shown.

   **Pass it the red team's and the critic's points too, and make it answer every red-team
   point**, accepting or rebutting each one with evidence from the files. If your
   environment will not let a subagent write the report file, have it return the report as
   text and write the file yourself.

   **The synthesizer has the same size limit as the red team.** When the notes will not fit
   in one agent's reading, add one agent per route that reads that route's notes and
   verdicts whole and writes `_synthesis/route--<route>.md`, answering the question from
   that route's evidence, both for and against. The synthesizer then reads those files, the
   red-team files, the critic and the verdicts, opens whole any note or source that a
   conclusion rests on, and says what it did not read.

8. **Produce three files, each under the limit in step 0**: the report, an evidence table of
   every quantitative claim with source URL, date and confidence, and a gaps file with two
   non-empty sections, being what could not be established and what was not covered. Build
   the annotated contents last, from an actual directory listing rather than from memory,
   **because an index cannot tell you what it left out, so the only check that works is
   comparing it against `ls`.**

   **Do not name that file `INDEX.md`.** Use a content name such as
   `00-annotated-contents.md`. Index generators overwrite a file of that name from a bare
   listing, keep no copy, and report success while doing it. **The general form is that a
   filename some tool generates is not yours to write into**, so check what else writes the
   name before you put work into it.

9. **Vet it yourself before anybody else sees it.** Whoever launched the research reads the
   synthesis, the red team and the critic in full, re-opens the sources that the main
   conclusions rest on, and asks whether the answer addresses the question the reader will
   actually ask rather than a nearby one. A checked fact is not a checked answer, and an
   answer built entirely from facts that passed their checks can still be the wrong answer.

## Which steps actually change the output

**Three of them do the work and the rest is bookkeeping.**

**Step 3**, because routes have different blind spots and topics do not.

**Step 5**, because it is the only thing standing between a confidently wrong number and
your name on it. The finder has already decided the claim is true, so asking the finder to
check it is asking somebody to disagree with themselves.

**Step 6**, because steps 3 and 5 make each fact more likely to be right while neither tests
whether the answer built from those facts is right. The red team is the only agent in the
method that is told to want the answer to be wrong.

**Step 2 is what makes the others honest.** An agent that writes notes without saving the
source has produced a summary, and everything downstream then works from that. Saving first
also makes a truncated fetch visible later as a short file, rather than invisible as a
confident paragraph.

## Mechanics

Use a parallel fan-out rather than sequential agent calls, because running the
investigation routes at the same time is the point. Give each agent its own notes directory
and sources directory, and **prefix every filename with that agent's key** when several
agents share a directory, or they will overwrite each other.

Each route's red team can start as soon as that route's checker returns, so finding,
checking and red-teaming run as a pipeline, route by route. The merge, the critic and the
synthesizer need every route, so they wait for all of them.

## What this does not replace

Never work from a summary. A truncated fetch counts as a summary, so read the end of every
one. **Absence is not failure to find**, which means "I could not establish this, and here
is where I looked" is a correct answer, while an empty gaps file is a false claim of
complete coverage.
