Lab experiment

Self-Improving AI Agents

Exploring how agents improve across attempts through execution feedback, verification, memory, and evaluation.

Status
Exploring
Concepts
  • Verification
  • Feedback
  • Evaluation

This diagram is a working model, not a deployed architecture.

A bounded improvement loop under exploration: a task reaches an agent, the agent works through tools, code and search to produce a result, and a verification gate either completes the attempt or routes it into feedback. Feedback becomes retained lessons that shape the next attempt, with human review bounding the loop rather than letting it run indefinitely.
01

observations

The system around the model

The question is whether coding and research agents get more reliable by treating their own output as something to execute, inspect, evaluate and revise, rather than returning the first plausible result.

The interesting work sits around the model rather than inside it: better verification, better retry decisions, execution feedback, memory worth keeping, search and branching, and lessons that survive a failure.

02

workflow

Attempt, verify, improve

  1. 01Task

    A bounded goal with its constraints attached.

  2. 02Attempt

    The agent works through tools, code, and search.

  3. 03Verify

    Execution and tests decide, not a second opinion.

  4. 04Improve

    A failure becomes feedback rather than a retry.

  5. 05Retain

    Lessons worth keeping go to memory; the rest does not.

03

next questions

Open questions

  • What feedback should survive between attempts, and what should not?
  • What belongs in persistent memory rather than temporary task context?
  • When is a deterministic verifier better than another model reviewing the work?
  • How should an agent decide whether to retry, branch, search, or stop?
  • How do you measure improvement without over-fitting to a weak evaluator?
  • How much can improve without touching model weights?
  • Where does human review belong in the loop?
04

observations

Current scope

The current scope is agent-system design around feedback, verification, memory, evaluation, retry decisions, and human review.

Previous experiment
AGENTCATALOGIDENTITYCHECKOUTCOMMERCE PLATFORMMAPPING
Universal Commerce Protocol

Still exploring

Thinking about the same problem?

I'm always interested in what works, what fails, and how to test the difference.

Next experiment
RAWCLEANLABELTRAINVALTEST
ML & Datasets