nonprofit-grant-finder — measured, then run live
Agent skill execution · replay

A skill that measures itself, then rewrites itself

The auto-research skill was pointed at nonprofit-grant-finder v1.0 and told to improve it. Not by rewriting it, but by running it five times, scoring all thirty outputs against binary evals, and changing exactly one thing. This is that run, replayed from its own logs. The skill has since been substantially rewritten as v2.1 — try the current version on the second tab.

28/30
Baseline · 93.3% pass
method Karpathy autoresearch loop
target nonprofit-grant-finder v1.0
evaluated 2026-03-19
grid 5 runs × 6 evals

Configure the loop

Auto-research refuses to start until five things are pinned down. Vague inputs make unreliable scores, so the skill blocks on this.

Target skill

nonprofit-grant-finder v1.0 — read in full, including every referenced file, before anything is changed.

Test inputs

Five prompts spanning different use cases. Variety is deliberate: a single scenario would let the skill overfit.

Eval criteria

Six binary yes/no checks. No 1–7 scales — scales compound variability and give unreliable results.

Runs per experiment

Five. More runs is more reliable but slower; five is the stated sweet spot.

Max score

6 evals × 5 runs = 30. That is the ceiling every experiment is measured against.

phase 01 / 08
The plate — every run scored against every eval
30 wells · awaiting scoring
Run · prompt E1Lundstrum Specificity E2Local Relevance E3Funding Mix E4Grounded Details E5Prioritization E6Reusable Strategy Value
P112 high-fit opportunities
P2Balanced ranked prospect list
P3Funder-fit analysis
P4Local & regional funders
P5Reusable grant language

Select any scored well to read the grader’s verbatim rationale.

On this data

Every score, rationale, and root cause here is transcribed from the run’s own baseline-evaluation.json, evals.md, and changelog.md. Experiment 1 has not been formally re-scored, so no score is claimed for it — only the structural change observed in its output files. Grader rationales are quoted with the client’s registration number, street address, and staff names removed.