Skip to content

Fit the engine to the chore, and let a fresh pair of eyes approve it

A common setup is one model for everything. The same one renames a variable, designs a database schema and untangles a bug that has run for a week. It is the simplest arrangement, and it spends effort in the wrong places.

Task routing is the alternative. Split the work into small jobs, give each to the one that suits it, and have something other than the author check the result.

For builders

~7 min read

One model for everything misses on both sides

Models come in tiers. Some are heavy and slow and strong at hard reasoning. Some are light and quick, and suited to well-defined, repetitive work. Treating them as interchangeable goes wrong in two directions.

Heavy work on a light model fails quietly. The output looks plausible and is wrong in a way you do not notice until later. A design decision made in a hurry is expensive to undo.

Light work on a heavy model wastes it. Renaming files, reformatting, writing the tenth near-identical test: all of it ties up the slowest, costliest tier for chores a quick one handles well.

The fix is to choose per job, every time.

The idea: short briefs, the right tier, a separate check

Three moves Each one is ordinary on its own

  1. Split the work into short briefs

    Each is one job with one outcome and a clear edge: what to change, where, and what done looks like. Small pieces pass on easily, and are simple to check and cheap to retry.

  2. Send each brief to the tier that fits

    Uncertain problems go to a heavy model, with time and room to think. Well-defined chores get a light one, fast. The match is a judgement, made fresh for each one.

  3. Have a separate judge check every return

    The model that wrote the code should not approve it. A second one reads the result against the brief, then passes it or sends it back for another round.

    That last move is the one most setups skip, and it is the one that makes the other two safe.

Why the check has to be independent

A model reviewing its own output tends to agree with itself. It made the choices, so they look reasonable to it. A fresh reader with a different brief, asked only whether the work meets the stated outcome, has no such attachment.

Independence is cheap here: one more read, of something already written. It also changes what you can claim afterwards. "The model reported it was done" is an assertion. "A separate judge checked it against the brief and passed it" is a result.

It also gives failed work somewhere to land. A return that does not pass goes back for another round, with the judge's reason attached, instead of arriving in your hands half right.

What it changes in practice

Three differences show up once routing is running.

You stop babysitting. With short briefs and a check on every return, you review the results at the end.

Heavy effort goes where it counts. The slow, careful tier spends its time on the decisions that deserve it. The fast one clears the routine.

Failures are contained. A bad return is caught at the judge, against one small brief, before anything is built on top of it.

What it needs is the discipline of the three moves, and tooling that holds them in place so they happen every time.

The tools we use, open source

Two of the three tools on our open source page belong to this work.

Router splits AI coding work into short briefs, sends each one to the model tier that fits it, and has a separate judge check every return. Anything that misses goes back for another round.

Harness is a packaged setup for Claude Code: hooks that fire at each step, skills it can load, and agents it can hand tasks to, all configured together. Router decides where each job goes; this is the environment it runs in.

The open source page describes both, with what each one does and who it is for.

How to start with what you have

You can try the idea this week, with the setup you already have.

  1. Write the next three jobs as briefs. One outcome each, with what done looks like.
  2. Decide the tier before you start. Hard and uncertain, or routine and defined.
  3. Check each return with a fresh read. A separate model, given only the brief and the result.
  4. Note what the check catches. Two weeks of that log will tell you whether it is earning its place.

If the check flags nothing, the briefs are probably too easy and the tier too heavy. If it catches a lot, you have found the work that needed a second pair of eyes all along.

If you build with AI for a living

The tools are public, so you can inspect them and decide for yourself. Fitting this approach into how your team ships is part of what we do.

hi [at] aeia.dev

Read next