Drop a loopy.yml into any repo. It says how often to run, what it may spend, where ideas come from, what counts as better and what nobody but you may touch.
# loopy.yml: the whole loop in one file
cadence:
mode: scheduled # always-on | scheduled | manual
every: 1w
budget:
per_experiment: { tokens: 400k, minutes: 20 }
per_cycle: { tokens: 20M, usd: 150 }
on_exhausted: report
sources:
- { id: community, adapter: intake }
- { id: research, adapter: arxiv, with: { query: "agent memory" } }
- { id: market, adapter: github-releases }
targets:
editable: [src/**, docs/**]
frozen: [bench/**, tests/**, loopy.yml]
evaluators:
- { id: bench-dev, run: make bench, metric: score, direction: up }
- { id: tests, run: make test, required: true }
- { id: bench-sealed, run: make bench-sealed, sealed: true }
gates:
keep_if: required_pass and primary_improves
human_signoff: [spec/**]
runner:
adapter: gh-aw # or local, claude-code, codex
cadenceAlways on, hourly, daily, every few days or weekly. Cycles start on the clock; budgets decide when they stop.
budgetHard caps in tokens and dollars per experiment, per cycle and per day. When a cap is hit, the loop writes its report and waits.
sourcesResearch papers, market moves, product feedback and public requests. Anything that can write a record can be a source.
evaluatorsThe commands that define "better". You write them, agents can't touch them, and a sealed set keeps the loop honest.
gatesKeep or revert, decided in code. Anything on a protected path waits for a person.