← All posts
28 August 2026Gustforward Marketing Team

The approval queue is the interface

Human-in-the-loop fails as a design problem long before it fails as a safety one. If approving takes as long as doing, you've built nothing.

A vertical stack of cards on dark navy, the top one lifted and glowing yellow

Human-in-the-loop is usually described as a safety measure. It's really a design problem, and treating it as safety-only is how it fails.

Here's the failure, and it's almost universal. The agent works. Drafts pile into a queue. The reviewer opens it, sees sixty items, and starts clicking approve. Not reading — clicking. By item twelve they've stopped even scanning. The queue is now theatre: a person is nominally accountable for output nobody examined.

Worse, everyone believes there's a human in the loop. The safety story is intact on the org chart and gone in practice.

The cause is nearly always the same: approving costs almost as much attention as doing the work would have. If that's true, the system has no value, and the reviewer's behaviour is a rational response.

So the approval queue isn't an admin screen bolted on at the end. It's the primary interface of the whole system, and it deserves the design effort you'd give a checkout flow.

Show the decision, not the document

The bad queue shows you the draft. Just the draft, in a list, with two buttons.

To review it you have to reconstruct everything the agent knew — open the original message, check the client's history, verify the price. That's the actual work, done again, slower.

A good queue shows the draft and the basis for it, side by side: the incoming message, the retrieved facts with their sources, what the agent concluded, and what it's proposing. The reviewer's job becomes checking a chain, not rebuilding one.

The test: can someone approve correctly without opening another tab? If not, you've moved work rather than removed it.

Sort by risk, not by time

FIFO is the wrong order. It means the highest-stakes item in the queue is treated identically to the eighteenth routine confirmation, and it arrives when attention is lowest.

Sort by what's at stake. The agent already knows more than enough to score this — how confident retrieval was, whether the client is high-value, whether there's a number or a commitment in the message, whether this input looks like anything it's seen.

Put the risky items at the top, while the reviewer is fresh. Let the routine tail be reviewed fast, because reviewing it fast is correct.

Batch the boring, isolate the important

Not everything deserves the same interaction. Forty near-identical appointment confirmations should be reviewable as a batch — scan a list, approve all, with anything anomalous automatically pulled out of the batch and shown separately.

One draft response to a client threatening to leave should occupy the whole screen, with full history, and no approve-all button anywhere near it.

Treating those two identically is what breaks reviewer attention. Uniform interfaces train uniform behaviour, and the behaviour they train is clicking.

Make editing the cheapest path

A two-button queue — approve or reject — forces a lie. Most drafts are neither. They're right except for one sentence.

If editing means rejecting and rewriting from scratch elsewhere, reviewers will approve near-misses, because approving something 90% right is faster than redoing it. You have designed your quality bar downward.

Make the draft editable in place, with one action to send the edited version. And then capture the edit as a signal, because it's the most valuable data in the system: a human took a specific output and made it specifically better. That diff is a labelled correction, and it belongs in the eval set.

Approve-rate alone hides this entirely. Approve-with-edit rate is the number that tells you whether the agent is actually good.

Show the reviewer their own effect

Nobody sustains a review habit with no feedback. Show it: how many drafts you approved this week, how many hours that represents against writing them, where the agent has been improving, what your edits changed.

This isn't gamification. It's the difference between a queue that feels like a new chore and one that feels like leverage — and it determines whether the system is still in use in month three.

Design for the queue being wrong

Two failures need a designed response, because both will happen.

The queue gets too long. Never let it silently grow. Past a threshold, the system should act: pause the workflow generating the backlog, alert someone, or — if the workflow has earned it — surface the option to move it up an autonomy level rather than drown. A queue nobody can clear is a system that has already failed quietly.

Something wrong got approved. Assume it will. Every approved action needs a visible trail and, where possible, an undo. Recall the message. Cancel the booking. Flag the record. The reviewer needs to know that a mistaken click is recoverable, because a reviewer who believes every approval is irreversible reviews slowly, then stops reviewing.

The metric nobody tracks

Here's the one worth putting on a dashboard: time to review, versus time to do the task manually.

If reviewing takes 15 seconds and writing took 4 minutes, the system is working — you've turned four minutes of production into fifteen seconds of judgment.

If reviewing takes 2 minutes against 4, you've built a very expensive way to halve someone's throughput while adding a supervision burden nobody asked for. And you'll find out from adoption dropping, months later, rather than from the number.

Measure it early. It's the honest verdict on the whole design, and it's the number that tells you whether your human-in-the-loop is a safeguard or a formality.


Monday: an agent that reads compliance filings — what human-in-the-loop looks like in a regulated workflow.

Got a workflow worth automating?

Tell us about it in a 30-minute call. If there's a fit, you'll get a scoped agent plan within 48 hours.

Book a call →