Every loop, a leap.

Loopy is an open protocol for autonomous research loops. Agents read the frontier, run the experiments and ship whatever beats the last best, day and night, on a budget you set. Built to outgrow itself.

Give this to your agent

Read github.com/technodotventures/loopy and set up a loop that improves this project.

Star on GitHub
Simulated loop · runningcycle 1 · day 1 of 7
0experiments
0kept
0.40best score
$0of $150 this cycle
Runs onNo vendor required · Apache-2.0
  • git + cronThe floor. Any repo, any scheduler, nothing else.
  • GitHub Agentic WorkflowsThe default binding. Loops run as Actions.
  • Claude CodeAgent adapter for proposals and experiments.
  • CodexAgent adapter. Swap in any other the same way.

how it works

Hypothesis in. Progress out.

Each lap takes in signals, turns them into hypotheses, tests them against evaluators you lock, and ships only what beats the record. Then it logs what it learned and goes again.

Listensignals inProposeagents draft RFCsProveexperiments, draft PRDecidegates, in codeShipkeep, or ask a humanReportscores and lessonsone cyclealways on, hourly, daily, every few days or weekly

Agents run every stage but one. Decide is plain code, because nothing should grade its own work.

the manifest

One file. The whole experiment.

Drop a loopy.yml into any repo. It says how often to run, what it may spend, where ideas come from, what counts as better and what nobody but you may touch.

# loopy.yml: the whole loop in one file cadence: mode: scheduled # always-on | scheduled | manual every: 1w budget: per_experiment: { tokens: 400k, minutes: 20 } per_cycle: { tokens: 20M, usd: 150 } on_exhausted: report sources: - { id: community, adapter: intake } - { id: research, adapter: arxiv, with: { query: "agent memory" } } - { id: market, adapter: github-releases } targets: editable: [src/**, docs/**] frozen: [bench/**, tests/**, loopy.yml] evaluators: - { id: bench-dev, run: make bench, metric: score, direction: up } - { id: tests, run: make test, required: true } - { id: bench-sealed, run: make bench-sealed, sealed: true } gates: keep_if: required_pass and primary_improves human_signoff: [spec/**] runner: adapter: gh-aw # or local, claude-code, codex
cadence

Always on, hourly, daily, every few days or weekly. Cycles start on the clock; budgets decide when they stop.

budget

Hard caps in tokens and dollars per experiment, per cycle and per day. When a cap is hit, the loop writes its report and waits.

sources

Research papers, market moves, product feedback and public requests. Anything that can write a record can be a source.

evaluators

The commands that define "better". You write them, agents can't touch them, and a sealed set keeps the loop honest.

gates

Keep or revert, decided in code. Anything on a protected path waits for a person.

cadence

Set the clock speed.

Some projects want a loop that never sleeps. Others want a careful lap once a week. Same protocol, same guarantees.

Always onEvery signal triggers triage; experiments run until the day's budget is spent.
HourlyFast code with cheap tests.
DailyThe overnight run: wake up to a report.
Every 3 daysTime for long benchmarks and human sign-off.
WeeklySpecs, protocols and anything slow to evaluate. Where Smartware and expresso run.

guardrails

Autonomous, not unsupervised.

A loop that improves itself needs rules it can't rewrite. These hold even if Loopy itself has a bug, because branch protection and code owners enforce them.

Humans own the scorer

Agents can change the code they're improving. Never the evaluators, the gates, the budget or the manifest.

Keep or revert

A change survives only if it beats the current version and passes every required check. Everything else is struck out.

Nothing grades itself

The verdict comes from deterministic code, not from the agent that wrote the change.

Sealed tests

Held-out evaluations run once per cycle on data no agent has seen, and retire before they go stale.

People sign the rules

Protected paths, like a spec, wait for a named approver. Releases are always a human action.

Every change has a reason

Each shipped change links the proposal, the runs and the signal that started it. Replay any loop from its ledger.

adapters

Bring any instrument.

Every part outside the core is an adapter, and an adapter is just a command that reads and writes JSON. Wrap any app, any API, any model.

KindWhat it doesBuilt in (v0.1)
SourceBrings signals inintake, files, webhook, rss, arxiv, github-issues, github-releases
EvaluatorScores a changecommand, pytest, unittest, benchmark-json
RunnerRuns the agent stages within budgetlocal, gh-aw, claude-code, codex
SinkPublishes progressfeed, files, github-projects, slack, discord

autonomously improving

Live on the bench.

Loopy runs in public on Techno Ventures' own work first. Each project brings its own evaluators and cadence; Loopy brings the loop, the budget and the guardrails.

smartwareAutonomously improving

The memory and context layer for software.

cadence
weekly
evaluators
LoCoMo recall, conformance suite, sealed test set
signals
memory research, the 15 systems it benchmarks against, community requests
Autonomously improving

The language for agents that act.

cadence
weekly
evaluators
conformance fixtures, verifier break attempts
signals
language research, design-partner feedback, community requests

Anyone can file a request or point either loop at a paper. A triage agent merges duplicates and adds it to the next cycle.

brand kit

Wear the loop.

Running Loopy on your project? Say so. Put a badge in your README or footer and it links back here, so others can follow the loop.

Autonomously improving with loopyThe standard. Use it wherever there's room.
In the loopThe compact cut, for tight spots like nav bars and package pages.
  • Autonomously improving with Loopy
    for light sites
  • In the loop with Loopy
    for light sites
  • Autonomously improving with Loopy
    for dark sites
  • In the loop with Loopy
    for dark sites

Copy the HTML for websites or the Markdown for READMEs. Keep the wording as is and don't recolour the loop. The full kit, with the wordmark, icon and colours, lives in the brand kit.

the spec

Four levels of trust.

Loopy is a protocol first. Run the reference tools or build your own; meet the requirements and say which level you run.

L1

Loop

Manifest, append-only ledger, an evaluator, keep or revert, hard budgets.

L2

Governed

Frozen paths, required checks, human sign-off, every change linked to its signals.

L3

Open

Public feed, public intake and a public board.

L4

Learning

Lessons carried between cycles, sealed evaluators, and a loop that tunes its own prompts.

Read the full spec · Agent setup guide

Close the loop. Open the frontier.

Give your project a cadence, a budget and a definition of better. Then let it run.

Give this to your agent

Read github.com/technodotventures/loopy and set up a loop that improves this project.

View on GitHub