← All resources

Agent skill

Appendix — derive it from the paper, then prove every line of it

Use when writing or auditing the appendix / online appendix / supplementary materials of an empirical paper or thesis — variable-definition tables, sample-construction details, method and identification descriptions, robustness tables, and reproducibility notes. Derives what the appendix must contain FROM the main text (which main-text claims need deferred support), then traces a three-link chain — main-text claim → appendix item → evidence in the project directory (data, code, output, logs, codebook) and literature — checking coverage (does the appendix back every main-text claim that needs it), grounding (does every appendix statement trace to real evidence), cross-artifact consistency (do the numbers match across text, appendix, and data), and internal self-consistency (does the main text agree with itself and the appendix agree with itself — a document that contradicts itself is flagged even when neither copy is anchored in the data). Decomposes every statement into a typed claim (definitional / factual-quantitative / methodological / citational / result-robustness / procedural) and classifies each as SUPPORTED / UNSUPPORTED / MISMATCH. Make sure to use this skill whenever the user mentions an appendix, online appendix, or supplementary materials, asks to "check my appendix", "does my appendix match the paper/code/data", "what should go in my appendix", wants a data or variable-definitions appendix drafted, or wants appendix claims verified — even if they don't say the word "verify". Never silently rewrites a number; flags mismatches for the user to resolve.


Appendix — derive it from the paper, then prove every line of it

An appendix is not a free-standing document — it exists to support the main text. So its contents are not arbitrary: they are determined by what the article claims. Every appendix item should answer a demand the main text creates (“see Appendix A for the full definition”, a headline number whose construction is deferred, a method whose assumptions are relegated, a promised robustness check), and every such demand should have a matching appendix item. That is the chain this skill builds and checks:

main-text claim  →  appendix item  →  evidence (project directory + literature)

Three things break along that chain and a fourth breaks within a node — the skill checks all four:

The unit of work, throughout, is borrowed from literature-review’s claim-decomposition: a sentence is not the unit of truth. A single sentence — in the main text or the appendix — bundles a definition, a count, a method, and a citation, each true or false on its own terms and verifiable in a different place. So you split prose into atomic, typed statements and chase each to ground.

This skill does two jobs against the chain — audit an existing appendix end to end, and author a new one by deriving it from the main text so it’s anchored by construction. Authoring is “audit, run forward.”

When to use

Mode A — Audit an appendix

Step 1 — Read the main text; derive the required-support set

The appendix is judged against the article, so start there. Read the main text and decompose its claims into atomic typed statements (same six types as below), then keep only the ones that create a demand for deferred support — i.e., the appendix should contain something for them. A main-text statement needs an appendix anchor when it:

The output is the required-support set: a list of main-text claims that demand appendix backing, each tagged with the appendix section that should hold it, and with the specific value/claim to later check for consistency (e.g. “abstract says N = 4,213”). Explicit “see Appendix X” pointers are the easy ones; the valuable catches are the implicit demands — a method named in §3 with no assumption discussion anywhere, a number in a main-text table never reconstructed.

While reading the main text, also record every place a single quantity or definition recurs within it — the sample N in the abstract, §3, and Table 1; a construct defined in §2 and re-used in §5 — with each location and value. This is the raw material for the internal-consistency check in Step 4: the main text must agree with itself before the appendix can be asked to match it. (references/empirical-exemplars.md shows, per method family, the categories of support a top-journal main text implies.)

Step 2 — Survey the directory: build the evidence map

You cannot verify against a project you haven’t looked at. Before reading the appendix closely, inventory where evidence could live, so each later claim has somewhere to go.

# data
find . -maxdepth 3 -regextype posix-extended -iregex '.*\.(csv|dta|parquet|rds|xlsx|feather|json)$' 2>/dev/null
# analysis code
find . -maxdepth 3 -regextype posix-extended -iregex '.*\.(R|do|py|jl|ipynb|sql)$' 2>/dev/null
# output & logs
find . -maxdepth 3 -regextype posix-extended -iregex '.*\.(tex|log|out|txt)$' 2>/dev/null | head -50
# documentation
find . -maxdepth 3 -iname 'readme*' -o -iname '*codebook*' -o -iname '*dictionary*' -o -iname 'requirements*' -o -iname '*.lock' -o -iname 'renv.lock' -o -iname 'environment.y*ml' 2>/dev/null

Skim names and headers (not whole files yet). Produce a short evidence map: where do data live, which script builds the sample, which produces each table, is there a codebook, is there an environment/seed record. This map tells you, for each statement type, the first place to look. (Adapt find to PowerShell Get-ChildItem -Recurse -Include on Windows if the Bash tool isn’t available; the project’s code/, output/, docs/ layout is described in CLAUDE.md.)

Step 3 — Decompose the appendix into typed statements

Read the appendix and split it into atomic statements, tagging each with exactly one of six types. Tag at the level of the smallest thing that could independently be true or false. The types, their evidence homes, and detection cues are in references/statement-taxonomy.md — read it before your first decomposition; it has the cue patterns and a worked example.

Type Asserts Evidence home
Definitional meaning/scope of a construct, variable, sample bound codebook, construction code
Factual-quantitative a count, range, share, moment the data, summary-stat output, the script
Methodological an estimator, design, specification, procedure estimation code + canonical citation
Citational attribution to a source Crossref
Result-robustness a finding or its stability the table/figure/log produced
Procedural-reproducibility software, versions, seeds, run order, access env files, README, scripts

Keep the statements in a working list (a table) — each row is one statement, its type, the parent sentence, and (to be filled) evidence + verdict.

Don’t skip the table and figure notes. Running prose is the easy part; the small print under each appendix table/figure is the densest claim source and the one most often missed — a clean-reading body sits above a note with the wrong cluster level or a stale star cutoff. Decompose every notes block the same way, expecting several types in a few lines: the SE/clustering key (“clustered at the firm level”) → methodological (verify against the estimation call); the significance legend (“*** p<0.01…”) → methodological/definitional (must match both the journal convention and the actual test); the sample/N line → factual-quantitative (re-derive); “includes firm & year FE” → methodological; “Source: …” / “see Table 3” → citational / consistency. The tables skill documents the canonical bottom-anchored notes order (../tables/SKILL.md → “The Notes block”); use it for coverage too — a results table whose note lacks the SE/clustering key or the legend is a COVERAGE GAP, not just a style nit.

Step 4 — Trace the chain: coverage, grounding, consistency (cross-artifact + internal)

Now connect the three lists — the required-support set (Step 1), the evidence map (Step 2), and the appendix statements (Step 3) — along the chain.

Internal consistency (each document against itself). Do this first, because it is cheap and needs no evidence. Take the recurring quantities and definitions collected in Step 1 (main text) and Step 3 (appendix) and compare each occurrence to the others within the same document: the abstract’s N against §3’s N against the main-text table’s N; the appendix’s own §A.2 sample count against its §C.1 note; a construct defined twice. Two statements of one quantity that disagree are an INTERNAL MISMATCH — and this verdict is independent of the external evidence. It stands even when both occurrences are UNSUPPORTED (nothing anchors either 4,213 or 4,198, yet they cannot both be the sample size), and it is the error the chain checks structurally miss: grounding each copy against evidence only surfaces the conflict if both copies happen to trace cleanly, so a self-inconsistent document with a missing anchor sails through as “two UNSUPPORTED rows” unless you compare the copies to each other. Run the check on the main text too, not only the appendix: a main text that disagrees with itself has no single value for the appendix (or the data) to be checked against, so resolve it before the cross-artifact pass can even be stated.

Coverage (main-text → appendix). Walk the required-support set and match each item to an appendix statement that backs it. An unmatched item is a COVERAGE GAP: the main text needs or promises support the appendix doesn’t provide (the classic “see Appendix B for the first stage” with no first stage). The reverse — an appendix item that supports no main-text claim — is an ORPHAN (usually harmless leftover, occasionally a sign the main text dropped a result; note it, lightly).

Grounding (appendix → evidence) and consistency. For each appendix statement, find its evidence and assign one of three verdicts — applied the way verify-citations classifies and the way analysis-cleanup refuses to silently change a number. Where a statement also appears in the required-support set with a value (e.g. the sample N), check that the chain matches in precedence order: main text == appendix == the displayed figure/table (all reader-facing, must agree to the printed digit) == the results-dir output (CSV/log) within rounding == raw data on re-derivation. The tolerance widens as you move upstream — see the factual-quantitative rule below.

The non-negotiable rule: never edit the appendix to match the evidence without surfacing the discrepancy first. A mismatch can mean the prose is stale or you found the wrong artifact — only the author knows which. Report it; let them decide. This is the same discipline as analysis-cleanup’s “never silently changes a number.”

Per-type evidence-finding (full detail in references/statement-taxonomy.md):

Parallelize the grounding pass (optional, for a large appendix). Once Step 2 (the evidence map) and Step 3 (the typed statement list) exist, the grounding verdicts are independent across statements — a textbook fan-out, the way paper-review runs its six reviewers. In a single message, launch one general-purpose subagent per statement type (or per appendix section when one type dominates), each read-only and each handed only what it needs: (a) its slice of the statement list, (b) the evidence map, and (c) the relevant reference — references/methodology-standards.md for methodological statements, verify-citations / Crossref for citational ones. Each subagent returns only its rows, in the verdict-table format — appendix location | statement | type | SUPPORTED/UNSUPPORTED/MISMATCH | evidence (file:line / table cell / DOI) — and you, the coordinator, merge them.

Three things stay single-threaded. Coverage (matching the required-support set to appendix items, the next sub-step) needs the whole picture, not a per-type slice. Internal consistency (comparing every recurrence of a quantity or definition within a document) is cross-statement by nature, so no per-statement subagent — each holding only its own slice — can see the conflict. And the final merge catches what no single subagent can see — the same value surfacing under two types (a sample N is both definitional and factual), or two slices disagreeing. As in paper-review, if a subagent returns nothing or malformed output, insert a placeholder row (<type> — agent returned no output) and surface it in the report; a missing slice must be visible, never silently dropped and never mistaken for “all SUPPORTED.” Skip the fan-out entirely for a short appendix — the dispatch overhead isn’t worth it.

Step 5 — Report the audit

Lead with the two tables that mirror the chain: a coverage table (main-text demand → appendix item) and a grounding table (appendix statement → evidence), each ordered so the author can walk their own document. Use the report skill’s Quick Template framing for the summary line.

# Appendix audit — <paper/section>

**Main-text claims needing support:** R  |  **Covered:** C  |  **Coverage gaps:** G
**Appendix statements checked:** N  |  **Supported:** S  |  **Unsupported:** U  |  **Mismatch:** M
**Internal conflicts:** main text: Im  |  appendix: Ia
Evidence map: <data dir> · <analysis scripts> · <output dir> · <codebook?> · <env record?>

## Coverage — does the appendix back what the main text needs? (main text → appendix)
| # | Main-text claim (location) | Needs | Appendix item | Status |
|---|---|---|---|---|
| 1 | abstract: "N = 4,213 firms" | data construction | §A.2 sample waterfall | COVERED |
| 2 | §3: "we instrument with judge leniency" | exclusion/first stage | — | COVERAGE GAP |
| 3 | §4: "robust to winsorization (App. C)" | robustness table | §C.1 (claim only) | COVERED (see grounding) |

## Grounding — does each appendix statement trace to evidence? (appendix → evidence)
| # | Appendix location | Statement (atomic) | Type | Verdict | Evidence |
|---|---|---|---|---|---|
| 1 | §A.2 | panel has 4,213 firms | factual | MISMATCH | data has 4,198 distinct firm_id (build_panel.R:88) |
| 2 | §B.1 | Callaway–Sant'Anna estimator | method | SUPPORTED | csdid call, estimate.do:42; cite ✓ |
| 3 | §B.1 | clustered at firm level | method | MISMATCH | code clusters at industry (estimate.do:44) |
| 4 | §A.3 | "active user" = ≥1 login / 28 days | definitional | SUPPORTED | active flag, clean.py:120 |
| 5 | §C.1 | robust to 1% winsorization | result | UNSUPPORTED | no winsorized table found in output/ |

## Internal consistency — does each document agree with itself? (main text ↔ itself, appendix ↔ itself)
| # | Document | Quantity / definition | Occurrence A (location → value) | Occurrence B (location → value) | Verdict |
|---|---|---|---|---|---|
| 1 | main text | sample N | abstract → 4,213 | §3.1 Table 1 → 4,198 | INTERNAL MISMATCH |
| 2 | main↔appendix | "active user" | §2 → ≥1 login / 28 days | §A.3 → ≥1 login / 30 days | INTERNAL MISMATCH |
| 3 | appendix | winsorization level | §A.4 → 1% | §C.1 note → 5% | INTERNAL MISMATCH |

## Coverage gaps (main text promises/needs support the appendix lacks)
- §3 instruments with judge leniency but no appendix first stage / exclusion discussion (see methodology-standards.md: judge-IV standards).

## Mismatches to resolve (author decides)
- Main text internal: abstract says N = 4,213; §3.1 Table 1 says 4,198 — resolve before the appendix can be checked against either.
- §A.2: prose says 4,213; recomputed 4,198 — stale draft, or different sample than build_panel.R?
- §B.1: prose says firm-level clustering; code uses industry — which is intended?

## Grounding gaps & orphans
- §C.1 robustness claim has no backing artifact — produce the table or soften the claim.
- §D (orphan): describes a dataset never referenced in the main text — leftover, or a dropped result?

Win condition: every required-support item COVERED, every appendix statement SUPPORTED, each document self-consistent, all values consistent across the chain — empty coverage gaps, internal conflicts, mismatches, and grounding gaps.

Mode B — Author an appendix

Authoring is the audit run forward: derive the appendix from the main text, then write each statement already knowing its evidence anchor, so the audit passes by construction.

  1. Derive the required-support set from the main text (Step 1 above). This is the appendix outline — one appendix item per main-text claim that needs deferred support. Don’t start from a blank template; start from what the article actually demands.
  2. Map onto a venue skeleton — slot the required-support items into the sections econ / marketing-IS / CS / general-interest reviewers expect (references/appendix-conventions.md, default Data → Methods → Robustness → Reproducibility). If an expected section has no corresponding main-text demand, ask the user whether the main text is missing something (e.g. a named method with no stated assumption) rather than padding the appendix.
  3. Calibrate depth against exemplars. For each method the paper uses, check references/empirical-exemplars.md for the categories of support a top-journal appendix in that field carries (construction + assumption/validation + robustness + materials).
  4. Survey the directory (Step 2 above) so you write from real artifacts, not memory.
  5. Anchor as you write — beside each statement, note the file:line / table / DOI it rests on (a parallel column or comment), then verify it the moment it’s written using the Step 4 machinery. Delete the scaffolding once green. When you re-state a value or definition the main text already gives, copy it from the anchored source rather than re-typing it — that’s how the internal-consistency check passes by construction instead of by luck.
  6. Cite methods to current standards from references/methodology-standards.md; run any new citation through verify-citations before it lands.
  7. Format — tables for definitions/parameters/sample-waterfalls (tables skill), self-contained captions, explicit cross-refs to the main-text equation/table each section supports. Prose register follows the paper-writing skill. Write it as publishable, reader-facing material — strip internal-report artifacts (edit logs, process narration, redundancy), keep it brief, one paragraph per topic; see references/appendix-conventions.md § Register.

Composition with sibling skills

This skill is a coordinator; it leans on the repo’s existing skills rather than re-implementing them:

Agent process notes