One of the broker’s features is preventing idle connections and slow handoffs. However, AI agents can sometimes take minutes (or longer) to reason, which may exceed a broker’s consumer timeout. There are three ways to handle it: raise the timeout and let the broker track the job for you, acknowledge the message immediately and track the job yourself somewhere else, or break the agent’s work into short steps that each fit inside the default window. This post walks through all three, what each one buys you, and what it costs.
The problem: a slow agent and a ticking clock
On the happy path, the worker leaves the request unacked while the agent works. When a good answer arrives, it publishes the response and acks the request. If the worker crashes, the broker redelivers the request to another worker. Nothing is lost.
The consumer timeout breaks this. Each delivery has a fixed window to be acked. If the agent is still thinking when the timeout window closes, the broker closes the channel and requeues the request, and the message will get processed again which might cause duplicates or other issues.
Why heartbeats don’t save you
Heartbeats prove the connection is alive. The timeout measures how long one delivery has gone unacked. A worker can heartbeat for hours and still time out. AMQP 0-9-1 also has no way to ask for more time on a delivery.
Why acking early isn’t the answer either
Acking on arrival removes the timeout and the safety net. If the worker dies mid-request, you can’t tell whether the agent was asked, answered, or the answer was lost. So who keeps track of a job while it runs?
Option 1: Let the broker keep track, with a longer timeout
Keep the happy path as it is. The request stays unacked while the agent works, and you raise the consumer timeout for that one queue to fit your slowest realistic request.
What you get:
- The unacked message is the job status. No extra table or service to keep in sync.
- Crash recovery for free. A dead worker’s request is redelivered automatically.
- Nothing new to operate. The broker you already run does the bookkeeping.
- Natural flow control. With a prefetch of one, queue depth shows exactly how much work is waiting.
- Readable code. Receive, ask, publish, ack.
To make it work well, scope the timeout to the slow queue, give the worker its own maximum runtime, keep heartbeats flowing while it waits on the agent, and make handling a request twice harmless.
Option 2: Ack early and track the job elsewhere
The worker acks right away and records the job in a database with an expiring lease, which it renews while the agent works. A separate process requeues requests whose leases expire. The broker wakes workers; the database is the source of truth.
What you get:
- Full visibility into what’s running, for how long, and where.
- Per-job control over cancelling, retrying and capping attempts.
- No timeout to tune, since deliveries are acked immediately.
The cost is building and operating a small job system. It pays off when you need that visibility.
Option 3: Break the work into steps
If the agent’s work has stages, like gathering context, drafting and reviewing, each can be its own message. A worker handles one stage, publishes the next, then acks.
What you get:
- Short deliveries that fit the default timeout.
- Progress survives crashes, losing only the current stage.
- Stages scale independently.
This only works when the work splits. A single long agent call can’t.
Choosing between them
All of them answer the question: who knows a request is in progress, and what happens if its worker disappears? Letting the broker hold the unacked request keeps the system small and the code obvious. An external job store buys visibility and control. Steps keep deliveries short, when the work allows it.
This post covers the happy path only. Handling retries because of agent failures, timeouts and bad answers is the topic of a follow-up post.