Recipe

Coding assistant on your own repo

A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.

SourcedNeeds level 5

A small open-source library has more open issues than its two maintainers have hours. Most are small: a failing edge case, a typo in an error message, a function that needs a test it never got. The job is an assistant that reads an issue, edits the repository, runs the tests and stops with a branch ready for review, not one that decides what to merge.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Coding assistant on your own repo

Propose, test and load a skill in a loop; a code-side check gates the branch it opens.

Level 5 · Agent loops
IssueIssueMODELAgent plans and actsAgent plansand actsTOOLload_skill(name)load_skill(name)TOOLpropose_edit + run testspropose_edit+ run testsCheck the diff and commands are in scopeCheck the diff andcommands are in scopeBranch opened for reviewBranch openedfor review
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 11Your code chose

The issue arrives

"parse_duration() raises on '1h30m'; add a test
and fix it."
0 tokens · 0 ms

Walkthrough

The agent reads the issue, and for a fix touching this repository’s own conventions (adding a test, say) calls load_skill for the relevant one rather than working from memory of how the last repository it saw was organized. It proposes an edit, code applies it and runs the existing tests, and the agent reads the result: if a test still fails, it tries again; once everything passes, it asks to open a branch.

That request is where the scope check runs. The model is not judging its own change: code compares the actual diff, and any commands the agent asked to run, against what this task was scoped to touch. A request inside the issue’s own files, running only the test suite, becomes a branch. One that reaches outside (a workflow file, an unrelated dependency pin, a shell command off the allowed list) is refused, with the reason handed back so the agent can try again.

Everything past the branch belongs to a person. Nothing merges its own pull request, and nothing the agent proposes reaches the default branch or a push credential until a maintainer reads the diff. The tests it ran are the ones already in the repository, not new coverage it invented.

Running this on every incoming issue takes a few more decisions. Nothing persists that the agent controls: the maintainers’ skills, the branches it opened, a log per issue. Every run starts from the repository as committed, so a bad one leaves a branch nobody has to merge. That log is what a maintainer reads the next morning: the issue, every edit attempted, every test result, every refusal the check made, since a refusal is the most interesting line in it and the easiest to lose. Stopping it means stopping the trigger: stop handing it issues and nothing new starts. The issue text, the files it opens and the test output go to whichever model runs it, unremarkable for a public library and a longer conversation for a private one. Cost is one agent loop per issue, and what goes wrong is rarely a wrong fix; it is a scope-creeping one the check let through.

What to measure

Coding agents don’t answer the site’s own document-QA questions, so this recipe needs its own small test set: synthetic issues against a synthetic library, each with a known-good fix and its test file, written to look like the issues this repository gets. Measure the share that reach a passing suite and the attempts each took. Then run the finished fix against the rest of the suite, not just the file the agent touched, to see whether an accepted change broke behavior the shown tests never exercised. Measure the scope check the way any permission check should be: try to get the agent to request something out of scope on purpose, in ways its author did not think of, and confirm every attempt is refused. None of this has been run here.

Variations

  • Add human approval as a pause before the branch is opened at all, for a repository where even a scoped, tested change should not reach a person’s queue unannounced.
  • Run several issues at once as parallel calls, one independent loop each, before reaching for anything that coordinates them.
  • Let the skill set grow with the repository, the way a coding agent’s instruction files are meant to, instead of fixing it once.
  • Widen the scope check from paths and commands to the size of the diff, refusing a change large enough that no fixed suite is good evidence it did only what the issue asked.

Design choices

Why this level, and when to use another approach

Coding agents is the core: the model proposes an edit, code runs the tests, and the model reads the result and decides whether to try again or report what it did. Skills holds the maintainers’ own conventions (a changelog format, a commit-message style, which test file pattern a given kind of fix uses) as short descriptions the agent reads every time and longer bodies it loads only when a task calls for one, rather than every convention sitting in the system prompt on every run. Safety is a code-side check between what the agent proposes and what happens to the repository: it reads the diff’s paths and the commands the agent asked to run, and refuses anything outside what this task was scoped to touch.

Two pieces of this sit below the agent and should stay there. The test suite is not a model at all; it is the signal the whole loop turns on, the same suite a maintainer runs by hand. The scope check is plain code reading a list, refusing by default anything nobody allowed in advance: a check in code, not a line in a prompt asking the model to behave.

Level 5 is enough because one issue is one bounded loop: propose, test, read the result, decide whether to try again. What would justify a lead agent and workers is not a longer backlog. That page’s condition is that the split cannot be written down before the work starts, and a queue of separate issues can be: ten known issues at once is parallel calls, level 3, ten copies of this loop with nothing to coordinate. The climb earns its cost when dividing one issue is itself a judgment, or when a second agent reviewing the first one’s diff catches what the tests cannot: each costing another agent’s worth of calls, spent on a conversation between two models about a change a maintainer is about to read anyway.

Composition

Techniques this recipe uses

The highest level it needs is level 5.

Coding agents

Sourced

Agents that read, write, run and test code.

Skills

Sourced

Reusable instructions that an agent loads when it needs them.

Safety, privacy and governance

Sourced

Prompt injection, permissions, data handling and audit.

Same shape, other jobs

Carry out a multi-step task in software, where the steps depend on what it finds

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Fix a bug or add a feature in a repository
  • Work a bring-up problem with read-only queries to instruments, the log and the datasheet
  • Reproduce somebody else's measurement from their notebook and say where the two differ
  • Reconcile two systems when finding the matching record is itself the work, rather than a field-by-field comparison
  • Migrate configuration from one format to another
  • Reproduce a reported defect

Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page