# Coding assistant on your own repo

_Recipe · needs level 5_

A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.


A small open-source library has more open issues than its two maintainers have hours. Most are
small: a failing edge case, a typo in an error message, a function that needs a test it never got.
The job is an assistant that reads an issue, edits the repository, runs the tests and stops with a
branch ready for review, not one that decides what to merge.

## Example run

_The web page for this technique includes an interactive step-through of Level 5 · assembled for this recipe. The same steps are described in the sections below._

## Walkthrough

The agent reads the issue, and for a fix touching this repository's own conventions (adding a
test, say) calls `load_skill` for the relevant one rather than working from memory of how the
last repository it saw was organized. It proposes an edit, code applies it and runs the existing
tests, and the agent reads the result: if a test still fails, it tries again; once everything
passes, it asks to open a branch.

That request is where the scope check runs. The model is not judging its own change: code
compares the actual diff, and any commands the agent asked to run, against what this task was
scoped to touch. A request inside the issue's own files, running only the test suite, becomes a
branch. One that reaches outside (a workflow file, an unrelated dependency pin, a shell command
off the allowed list) is refused, with the reason handed back so the agent can try again.

Everything past the branch belongs to a person. Nothing merges its own pull request, and nothing
the agent proposes reaches the default branch or a push credential until a maintainer reads the
diff. The tests it ran are the ones already in the repository, not new coverage it invented.

Running this on every incoming issue takes a few more decisions. Nothing persists that the agent
controls: the maintainers' skills, the branches it opened, a log per issue. Every run starts from
the repository as committed, so a bad one leaves a branch nobody has to merge. That log is what a
maintainer reads the next morning: the issue, every edit attempted, every test result, every
refusal the check made, since a refusal is the most interesting line in it and the easiest to
lose. Stopping it means stopping the trigger: stop handing it issues and nothing new starts. The
issue text, the files it opens and the test output go to whichever model runs it, unremarkable for
a public library and a longer conversation for a private one. Cost is one agent loop per issue,
and what goes wrong is rarely a wrong fix; it is a scope-creeping one the check let through.

## What to measure

Coding agents don't answer the site's own document-QA questions, so this recipe needs its own
small test set: synthetic issues against a synthetic library, each with a known-good fix and its
test file, written to look like the issues this repository gets. Measure the share that reach a
passing suite and the attempts each took. Then run the finished fix against the *rest* of the
suite, not just the file the agent touched, to see whether an accepted change broke behavior the
shown tests never exercised. Measure the scope check the way any permission check should be: try
to get the agent to request something out of scope on purpose, in ways its author did not think
of, and confirm every attempt is refused. None of this has been run here.

## Variations

- Add [human approval](/gradient_ascent/techniques/human-in-the-loop/) as a pause before the
  branch is opened at all, for a repository where even a scoped, tested change should not reach a
  person's queue unannounced.
- Run several issues at once as [parallel calls](/gradient_ascent/techniques/parallelization/),
  one independent loop each, before reaching for anything that coordinates them.
- Let the skill set grow with the repository, the way a coding agent's instruction files are meant
  to, instead of fixing it once.
- Widen the scope check from paths and commands to the size of the diff, refusing a change large
  enough that no fixed suite is good evidence it did only what the issue asked.
## Design choices

### Why this level, and when to use another approach

[Coding agents](/gradient_ascent/techniques/coding-agents/) is the core: the model proposes an
edit, code runs the tests, and the model reads the result and decides whether to try again or
report what it did. [Skills](/gradient_ascent/techniques/skills/) holds the maintainers' own
conventions (a changelog format, a commit-message style, which test file pattern a given kind of
fix uses) as short descriptions the agent reads every time and longer bodies it loads only when a
task calls for one, rather than every convention sitting in the system prompt on every run. [Safety](/gradient_ascent/techniques/safety/) is a code-side check between what the agent proposes and
what happens to the repository: it reads the diff's paths and the commands the agent asked to run,
and refuses anything outside what this task was scoped to touch.

Two pieces of this sit below the agent and should stay there. The test suite is not a model at
all; it is the signal the whole loop turns on, the same suite a maintainer runs by hand. The scope
check is plain code reading a list, refusing by default anything nobody allowed in advance: a
check in code, not a line in a prompt asking the model to behave.

Level 5 is enough because one issue is one bounded loop: propose, test, read the result, decide
whether to try again. What would justify [a lead
agent and workers](/gradient_ascent/techniques/orchestrator-workers/) is not a longer backlog. That page's condition is that the split cannot
be written down before the work starts, and a queue of separate issues can be: ten known issues at
once is [parallel calls](/gradient_ascent/techniques/parallelization/), level 3, ten copies of
this loop with nothing to coordinate. The climb earns its cost when dividing one issue is itself a
judgment, or when a second agent reviewing the first one's diff catches what the tests cannot:
each costing another agent's worth of calls, spent on a conversation between two models about a
change a maintainer is about to read anyway.



Last reviewed 2026-09-18.
