← All Agent engineering lessons

Agent engineering / HANDS-ON LESSON

Stop a failing agent loop

Give a controller a fixed attempt budget and distinguish success, permanent failure, and exhaustion.

BeginnerAbout 15 min5 practice questionsReviewed 2026-10-05SaveMyToken editorial

WHAT YOU WILL BUILD

A finite execution trace that cannot retry forever.

The useful part of an agent loop is the feedback between a proposed action and an observed result. The controller also needs explicit terminal states. A sequence of unsuccessful calls should consume a finite budget rather than becoming an unbounded conversation.

Before you start

Use a text editor and Node.js 22 or newer ↗. Check your version with node --version. Save the downloaded file in an empty folder, then open a terminal in that folder.

All data is included. No packages, account, or API key are required. The outcomes are a deterministic test fixture. No model, service, network retry, or real delay is involved.

Download the lab (.mjs)

FOLLOW ALONG

Work through the example.

  1. 01

    Read the failure sequence

    The supplied sequence has three temporary failures followed by a success. With a maximum of three attempts, that fourth event must never be reached.

  2. 02

    Run the bounded controller

    The output is Attempts: 3 and Final state: exhausted. Follow the counter increment and the check at the top of the loop to understand the boundary.

  3. 03

    Exercise early termination

    Replace the second outcome with success and predict two attempts with complete status. Then use permanent-error and expect two attempts with failed status. Update the assertions for each fixture.

  4. 04

    Plan the production limits

    Alongside attempt limits, set a wall-clock deadline, per-call timeout, and spending limit for actual integrations. Decide which terminal states require human help and record the completed work before stopping.

THE COMPLETE LAB

Run it locally.

Run this command from the folder containing your downloaded file:

node agent-stop-conditions.mjs

View or copy the complete JavaScript
// SaveMyToken local lab. Run with Node.js 22 or newer.
import assert from "node:assert/strict";

// A deterministic controller, not an LLM agent or a live service.
const outcomes = ["transient-error", "transient-error", "transient-error", "success"];
const maxAttempts = 3;
let attempts = 0, status = "exhausted";
for (const outcome of outcomes) {
  if (attempts >= maxAttempts) break;
  attempts++;
  if (outcome === "success") { status = "complete"; break; }
  if (outcome === "permanent-error") { status = "failed"; break; }
}
console.log("Attempts:", attempts);
console.log("Final state:", status);
assert.equal(attempts, 3);
assert.equal(status, "exhausted");

Expected output for the unchanged example

Attempts: 3
Final state: exhausted

The assertions also check the baseline behavior. After changing an input, predict the result and update the relevant assertion.

NOW CHANGE ONE THING

Make the example your own.

Detect a repeated identical action and stop it with a distinct repeated-action state. Test that two different legitimate actions are not incorrectly treated as a loop.

You are done when…

You can demonstrate exhaustion, early success, and permanent failure without reading events beyond the stopping point.

If something goes wrong

A retry count cannot interrupt a request that never returns. Real asynchronous calls need cancellation or timeout handling as well as a controller budget.

EXPLAIN WHAT YOU LEARNED

Interview practice

Try answering aloud before opening the reference answer. These are original learning questions, not a record of any employer's interviews.

01What connects actions in an agent loop?

The next decision should use the observed result of the previous action and the remaining task state. The controller should not treat a proposed action as already completed.

Watch for: Planning text alone does not establish execution success.

02Which stop conditions should be explicit?

Task acceptance, permanent failure, attempt exhaustion, elapsed-time limits, spending limits, and cancellation. Choose the appropriate conditions for the operation being performed.

Watch for: One universal iteration count cannot protect against every failure.

03How do you handle a temporary service failure?

Classify it, check whether repetition is safe, and retry within bounded attempts and time. Respect service retry guidance and avoid synchronized retry bursts.

Watch for: A validation error generally needs corrected input rather than the same retry.

04How should a workflow report partial completion?

Record completed actions, evidence, unresolved work, and the reason for stopping. Make any external effects visible so the next attempt can resume safely.

Watch for: Calling the whole task failed can hide actions that already succeeded.

05How do you reduce repeated unsuccessful tool use?

Track attempted actions and their outcomes, detect unproductive repetition, and require new information or escalation before repeating the same failed approach.

Watch for: Changing the wording of the same action is not necessarily progress.

Practice more agent engineering questions

Sources and further reading

The explanation and local exercises were written for SaveMyToken. These references support the underlying concepts; the sample outputs describe only the supplied examples.

Anthropic: tool use ↗