Factory.ai

AI Coding Agents

Reliability

Retry safety for agent-generated application code

September 26, 2026 - 2 minute read

Retry safety is easy to describe and difficult to prove. A coding agent asked to “make this call reliable” may add another loop around an operation that already retries in the client or queue. The code can pass its happy-path tests while multiplying load, duplicating a side effect, or extending failure beyond the caller's deadline.

Amazon's guidance on timeouts, retries, and backoff with jitter explains why retries consume capacity on a service that may already be struggling. A safe change starts by identifying which layer owns the retry.

Map retry safety ownership and side effects

Give the agent the full call path from request boundary to dependency. Include client-library defaults, proxy behavior, queue delivery semantics, job-runner policy, and the caller's total time budget. Search configuration as well as source code because retry counts often live outside the changed function.

Classify the operation before editing it. Reads may still trigger metering or cache fills. Writes need an idempotency mechanism when the caller cannot distinguish a lost response from a lost request. Amazon's guidance on making retries safe with idempotent APIs describes using a caller-provided request identifier to recognize repeated intent.

Factory's Droid Exec can run a scoped investigation and save a report for CI. Ask it to name every retry layer and side effect before proposing code, then constrain edits to the selected owner.

Record the existing behavior before changing it. Attempt counts, timeout values, queue visibility settings, and idempotency keys form one contract even when they live in different systems.

Test retry safety under failure

Use a deterministic fake or fault-injection layer. Cover connection failure before send, timeout after the server commits, throttling, partial response, and permanent validation errors. Assert the number of attempts, elapsed budget, final error, and count of durable side effects.

Backoff should have a cap and jitter appropriate to the client population. The retry loop must also stop when the request deadline or cancellation signal expires. Do not retry authentication, authorization, malformed input, or another failure the next attempt cannot change.

Check observability without logging request bodies or credentials. Emit an operation name, attempt count, safe request identifier, reason category, and terminal outcome. Alerting should reflect final user-visible failure rather than every transient attempt.

Make the change reviewable

The pull request should state which layer owns retries, which errors qualify, the maximum attempt and time budgets, and how duplicate effects are prevented. Include the failure matrix and test output. Keep timeout changes separate when they alter the caller's latency contract.

Factory's automated code review can analyze the diff for error-handling and correctness problems. Add repository guidance for local retry conventions, then keep fault-injection tests as a required gate because they demonstrate timing and side-effect behavior.

Roll out with metrics for attempt count, exhausted retries, dependency errors, and operation latency. A lower visible error rate is insufficient if total dependency traffic or duplicate work rises. The evidence should show that recovery improved within the existing capacity and deadline constraints.

Further reading

Ready to build the software of the future?

Start building

Arrow Right Icon