Factory.ai

AI Coding Agents

Testing

Concurrency testing for coding agent changes

September 27, 2026 - 2 minute read

Concurrency testing gives reviewers evidence about behavior that ordinary unit tests often miss. A coding agent can make a local synchronization change that compiles, passes once, and still introduces a data race, deadlock, lost update, or stale read under a different schedule.

Dynamic detectors are a strong first gate. Go's official race detector instruments memory access when tests run with -race. Clang's ThreadSanitizer documentation describes a similar detector for C and C++. Both observe executed paths, so a clean report covers the tested workload rather than every possible execution order.

Scope concurrency testing around invariants

State the property that must remain true across workers. A counter must equal committed updates. A cache must never expose a partially initialized value. A job may have one active owner. A shutdown path must stop accepting work before shared resources close.

Map every shared value to its synchronization mechanism and ownership. Locks, channels, atomics, database constraints, and leases protect different boundaries. Ask the coding agent to explain which boundary changed and why the existing mechanism still covers every access.

Keep timing out of correctness assertions where possible. Sleeping for a fixed interval makes a test depend on machine speed. Barriers, latches, fake clocks, and controllable executors can place operations at the contested point and make the failure reproducible.

Run concurrency testing in layers

Start with focused tests under the language's race detector or sanitizer. Exercise repeated access, cancellation, error returns, and shutdown. Run enough iterations to explore schedules, but save the random seed and command so a failure can be repeated.

Add a controlled test for the reported or plausible execution order. Pause one worker after it reads state, let another update that state, then resume the first. Assert the business invariant, not the internal order, unless ordering is itself the contract.

Stress tests can reveal starvation, lock contention, and resource leaks that short tests miss. Record worker count, duration, CPU allocation, and detector settings. Treat a timeout as evidence to investigate rather than rerunning until green.

Race detectors do not establish freedom from deadlock. Add bounded completion tests for lock ordering, cancellation, and shutdown paths. When the language exposes blocked-thread or goroutine diagnostics, retain that output with a failure so the next run starts from evidence.

Make the result reviewable

The pull request should identify shared state, the synchronization decision, detector commands, controlled execution orders, and stress results. Include a regression test that fails without the fix when practical. Do not weaken assertions or add broad retries to hide nondeterminism.

Factory's remote delegation guidance recommends exact verification commands and acceptance criteria. That context helps a Droid reproduce the same concurrency test rather than choosing an easier substitute. Factory's Automated Code Review can inspect pull requests against repository-specific guidance, while runtime detectors provide the execution evidence.

Some failures remain environment-sensitive. Stop and report the limitation when a required architecture, kernel facility, sanitizer runtime, or production workload cannot be reproduced. A documented boundary is more useful than a confident result from an unrepresentative test.

Further reading

Ready to build the software of the future?

Start building

Arrow Right Icon