In answer-sheet generation, users ask for a batch while the backend handles many smaller operations. Most files may be ready while a few remain stuck. A batch-level percentage can hide the reason.
A recoverable document job
Separate acceptance from completion
Validate intent and return a stable batch identifier. Rendering happens asynchronously. The batch record describes expected work, finished items, and failures that need attention. Accepted does not mean completed; a queue acknowledgment does not mean a document exists.
Make one work item recoverable
Consider a worker that uploads a PDF, loses its connection before updating status, and receives the job again. A new storage key or unconditional completion insert can create duplicate effects.
job_id = batch_id + ":" + document_id
artifact_key = "batches/" + batch_id + "/" + document_id + ".pdf"
if completion_store.is_complete(job_id):
acknowledge()
else:
artifact = render_and_store(artifact_key)
completion_store.complete_once(job_id, artifact)
acknowledge()This is pseudocode, not a complete transaction. Storage and status cross a boundary. A real implementation needs an atomic claim or compare-and-set, safe replacement behavior, and reconciliation of orphaned artifacts. Standard SQS queues can redeliver messages; consumers must account for that.
Retries are additional load
Distinguish transient errors from invalid inputs. Delayed retries with jitter and a finite limit can reduce amplification. A dead-letter queue preserves failed work; it still needs an owner and a repair procedure. An immediate retry of expensive work can worsen overload.
A synthetic capacity exercise
Suppose 20 workers each complete one document per second. At 40 workers, database waits extend service time to three seconds. The idealized rates become 20 and about 13.3 documents per second. These are illustrative numbers, not production measurements. Workers divided by mean service time is only an approximation, before utilization and other bottlenecks.
Compare oldest pending work, useful completions, latency percentiles, retry rate, database waits, and memory over the same interval. Increase concurrency in small steps and stop when useful throughput stops improving.
Finish the last part of the batch
Can users download completed files? Are failed items repaired automatically? Does the batch become partially complete or stay in progress forever? A visible failed-item list is more actionable than a percentage.
What I carry into the next design
A bounded unit of work, idempotent effects, explicit partial outcomes, and a repeatable load test matter more than the queue library.
Related case: answer sheet generation
References and further reading
Primary sources for the technical concepts in this article. The examples and decisions above are my synthesis, not quotations from these sources.
- Amazon SQS documentationAt-least-once delivery
Why a consumer must handle repeated delivery.
- Amazon SQS documentationVisibility timeout
Processing windows and message redelivery.
- AWS Builders’ LibraryTimeouts, retries, and backoff with jitter
Retry load, bounded timeouts, and backoff.
