The auto-research skill was pointed at nonprofit-grant-finder v1.0 and told to improve it. Not by rewriting it, but by running it five times, scoring all thirty outputs against binary evals, and changing exactly one thing. This is that run, replayed from its own logs. The skill has since been substantially rewritten as v2.1 — try the current version on the second tab.
Auto-research refuses to start until five things are pinned down. Vague inputs make unreliable scores, so the skill blocks on this.
nonprofit-grant-finder v1.0 — read in full, including every referenced file, before anything is changed.
Five prompts spanning different use cases. Variety is deliberate: a single scenario would let the skill overfit.
Six binary yes/no checks. No 1–7 scales — scales compound variability and give unreliable results.
Five. More runs is more reliable but slower; five is the stated sweet spot.
6 evals × 5 runs = 30. That is the ceiling every experiment is measured against.
| Run · prompt | E1Lundstrum Specificity | E2Local Relevance | E3Funding Mix | E4Grounded Details | E5Prioritization | E6Reusable Strategy Value |
|---|---|---|---|---|---|---|
| P112 high-fit opportunities | ||||||
| P2Balanced ranked prospect list | ||||||
| P3Funder-fit analysis | ||||||
| P4Local & regional funders | ||||||
| P5Reusable grant language |
Select any scored well to read the grader’s verbatim rationale.
nonprofit-grant-finder v2.1 is built around one rule: never state a grant, deadline, or eligibility fact that wasn’t verified against the funder’s own page this session. This page has no network access, so it cannot verify a funder — which means it cannot honestly produce a grant list. What it can do faithfully: run intake, build your exact research queries, score real candidates you found elsewhere against the documented rubric, and draft from the real templates.
The matching key for everything downstream. mission area, geography, population, funding need, and target amount are the skill's own stated blocking minimum — everything else sharpens matching but isn't required.
The replay on tab one covers a single v1.0 experiment from a demo archive. Separately, a real autoresearch loop ran against v2.0 across three experiments and four unrelated nonprofit missions — and its two kept mutations are what the current SKILL.md actually ships. This is that run.
Run separately from the v1.0 demo on tab one — same autoresearch method, this time against the v2.0 skill that actually shipped, with a harder eval suite built around fabrication rather than output shape.
Haiku — executes the skill and does live web search. Grading and mutation decisions are made by the main model, kept separate from the runner.
Four unrelated missions: Lundstrum Performing Arts (MN youth arts), Prairie Roots Food Bank (rural MN), Riverside Community Health Center (Toledo OH), and Clearwater Watershed Alliance (NC) — rotated in for experiment 2 specifically as a generalization check.
Five binary checks, 3 runs per experiment · max score 15.
Two experiments run after baseline, both kept. Progression: 60.0% → 86.7% → 93.3%.