What is a Dead-Letter Queue (DLQ)?
Page topics
What is a Dead-Letter Queue?
A dead-letter queue (DLQ) is a special type of message queue that temporarily stores messages that a software system cannot process due to errors. Message queues are software components that support asynchronous communication in a distributed system. They let you send messages between software services at any volume and don’t require the message receiver to always be available. A dead-letter queue specifically stores erroneous messages that have no destination or that can’t be processed by the intended receiver.
Why are dead-letter queues important?
Dead-letter queues (DLQ) exist alongside regular message queues to manage message processing failures. They act as temporary storage for erroneous and failed messages. DLQs prevent the source queue from retaining and filling with messages that are not processed successfully.
For example, consider a software that has a regular message queue and a DLQ. The software uses the regular or main queue to hold messages it plans to send to a destination. If the receiver fails to respond or process the sent messages, the software moves them to the dead-letter queue.
There are two potential causes of delivery failures, where messages enter the DLQ pipeline: erroneous message content and changes in the receiver’s system.
Erroneous message content
DLQ isolates problematic messages that failed to be delivered successfully. Hardware, software, and network issues might corrupt the sent data. For example, hardware interference slightly changes some of the information during transmission, which compromises data integrity. The unexpected data corruption could cause the receiver to reject or ignore the message.
Changes in the receiver’s system
A message might also move to a DLQ if the receiving software has gone through changes that the sender is not aware of. For example, you could attempt to update a customer's information by sending a message for CUST_ID_005. However, the receiver could fail to process the incoming message because it has removed the customer from the system's database.
What are the benefits of a dead-letter queue?
Next, we talk about the benefits of dead-letter queues (DLQ).
Reduced communication costs
Regular or standard message queues keep processing messages until the retention period expires. This helps ensure continuous message processing and minimizes the chances of your queue being blocked.
However, if your system processes thousands of messages, a large number of error messages will increase communication overhead costs and burden the communication system. Instead of trying to process failing messages until they expire, it’s better to move them to a dead-letter queue after a few processing attempts.
Improved troubleshooting
If you move erroneous messages to the DLQ, this lets your team focus on identifying the causes of the errors. They can investigate why the receiver couldn't process the messages, apply the fixes, and perform new attempts to deliver the messages.
For example, a banking software might send thousands of credit card applications daily to its backend system for approval. From there, the backend system receives the applications, but cannot process all of them because of incomplete information. Instead of making endless attempts, the software moves the messages to the DLQ until the IT team resolves the problem. This allows the system to process and deliver the remaining messages without performance issues.
When should you use a dead-letter queue?
You can use a dead-letter queue (DLQ) if your system has the following issues.
Unordered queues
You can take advantage of DLQs when your applications don’t depend on ordering. While DLQs help you troubleshoot incorrect message transmission operations, you should continue to monitor your queues and resend failed messages.
FIFO queues
Message ordering is important in first-in, first-out (FIFO) queues. Every message must be processed before delivering the next message. You can use dead-letter queues with FIFO queues, but your DLQ implementation should be FIFO as well.
When should you not use a dead-letter queue?
You shouldn’t use a dead-letter queue (DLQ) with unordered queues when you want to be able to keep retrying the transmission of a message indefinitely. For example, don’t use a dead-letter queue if your program must wait for a dependent process to become active or available.
Similarly, you shouldn’t use a dead-letter queue with a first-in, first-out (FIFO) queue if you don’t want to break the exact order of messages or operations. For example, don’t use a dead-letter queue with instructions in an edit decision list (EDL) for a video editing suite. In this instance, by changing the order of edits, you change the context of subsequent edits, which impacts system reliability.
How does a dead-letter queue work?
For the most part, a dead-letter queue (DLQ) works like a regular message queue, except it's meant for error handling. It stores erroneous messages until you process them to investigate the reason for the error, which preserves message integrity. Developers or system administrators may examine the DLQ, leading to fixes. After fixing the errors, the DLQ messages can be manually routed to the original queues. Some modern messaging systems support automated DLQ redrive.
Next, we discuss the redrive policy for DLQs and how messages move in and out of DLQs.
Creating a redrive policy
Software moves messages to a dead-letter queue by referring to the redrive policy. The redrive policy consists of rules that determine when the software should move messages into the dead-letter queue. Mainly by defining the maximum retry count, the redrive policy regulates how the source queue and dead-letter queue interact with each other.
For example, if a developer sets the maximum retry count to one, the system moves all unsuccessful deliveries to the DLQ after a single attempt. Some failed deliveries may be caused by temporary network overload or software issues. This sends many undelivered messages to the DLQ. To get the right balance, developers optimize the maximum retry count to ensure the software performs enough retries before moving messages to the DLQ.
Moving messages into the dead-letter queue
Delivery attempts between the sender and receiver can fail for several reasons:
- The receiver fails to receive the message because it doesn't exist.
- The message contains errors.
- The message exceeds the queue or message length limits. For example, some receivers can't process messages that exceed a specific size.
- The message's time to live (TTL) has expired. TTL is a value that indicates how long a particular data packet is valid on the network.
- The sender exceeds the configured retry value.
Moving messages out of the dead-letter queue
When messages move into the dead-letter queue, you must inspect the erroneous messages to determine the causes. Messages in the DLQ might contain valuable insights to prevent future recurrences of similar issues. After you analyze and remediate the issues, the system moves the messages out of the DLQ and into the source queue. This allows the sender to continue processing the messages.
How can AWS support your dead-letter queue requirements?
Amazon Simple Queue Service (SQS) is a fully managed message queuing service for microservices, distributed systems, and serverless applications. Amazon SQS lets you send, store, and receive messages between software components at any volume, without losing messages or requiring other services to be available.
Here are other benefits of Amazon SQS:
- Standard queues with unlimited throughput, at least once delivery. And best-effort ordering
- FIFO queues with high throughput, exactly-once processing, and first-in-first-out delivery
- Unlimited queues and messages, up to 264KB of text in any format, batch message operations, long polling, fair queues, and more
Get started with message queues by creating a free AWS account today.
Browse all cloud computing concepts
Browse all cloud computing concepts content here:
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages