Foundations roadmap

Queues and Background Jobs

Log in to save this

Saving keeps this in your list across devices. It's a free account — no card.

When a user signs up, your server might need to save the account, send a welcome email, and resize their profile photo. If the request waits for all of that, the user stares at a spinner for several seconds, and a slow email provider can make signup fail entirely. A queue lets the request do only the part the user must wait for, and hands the rest to a separate process that works through it in the background. In the Queue Depot, every crate on the conveyor is a job someone promised to finish later.

Why slow work leaves the request

A good rule: the request should do what the user needs to see the result of, and nothing more. Saving the account must happen before you answer "welcome". Sending the email does not; nobody notices if it arrives two seconds later.

Slow work inside the request causes three problems. The response is slow. If the slow step fails (the email API is down), the whole request fails, even though the important part worked. And each waiting request holds server resources, so a burst of signups can tie up the server while it waits on someone else's API.

Producers, queues, and workers

A queue has three parts:

  • The producer is the code that adds a job, usually your request handler.
  • The queue stores jobs until someone takes them. It is often backed by Redis, a database table, or a hosted service.
  • The worker (the consumer) is a separate process that takes jobs off the queue and runs them.
// In the request handler (producer)
app.post('/signup', async (req, res) => {
  const user = await db.createUser(req.body);
  await queue.add('send-welcome-email', { userId: user.id });
  res.status(201).json({ id: user.id });
});

// In a separate worker process (consumer)
queue.process('send-welcome-email', async (job) => {
  const user = await db.findUser(job.data.userId);
  await email.send(user.email, 'Welcome!');
});

Notice the job carries an ID, not the whole user. The worker looks up fresh data when it runs, which might be minutes later.

Because the worker is its own process, you can run more workers when the queue gets long, and a crash in a worker does not take down the web server.

Jobs can run twice

Most queues promise at-least-once delivery. A worker takes a job, and the queue waits for it to report "done". If the worker crashes, or takes too long, before reporting, the queue assumes the job was lost and hands it to another worker. Sometimes the first worker had actually finished, so the job runs twice.

The fix is to make handlers idempotent: running them twice has the same effect as running them once. For the welcome email, record that it was sent and check first:

queue.process('send-welcome-email', async (job) => {
  const user = await db.findUser(job.data.userId);
  if (user.welcomeEmailSentAt) return; // already done
  await email.send(user.email, 'Welcome!');
  await db.markWelcomeEmailSent(user.id);
});

This does not close every gap (a crash between sending and marking can still send twice), but it turns "every retry sends again" into a rare edge case. For payments, where twice is never acceptable, use an idempotency key that the payment provider itself checks.

When the queue grows faster than it drains

A queue absorbs bursts: 500 signups in a minute become 500 jobs that workers finish over the next few minutes. But if jobs arrive faster than workers can finish them for a long time, an unbounded queue just keeps growing. Jobs wait longer and longer, memory in the queue store climbs, and eventually it runs out and falls over, taking every waiting job with it.

Watch the queue length and the age of the oldest job. If they keep rising, add workers, make jobs faster, or put a limit on the queue so producers get a clear error instead of silently piling up work nobody will reach in time.

Dead-letter queues for jobs that keep failing

Some jobs will never succeed: the user was deleted, or the data is malformed. If the queue retries them forever, they waste worker time and clutter the logs. Instead, give each job a retry limit. After the last attempt fails, move it to a dead-letter queue, a separate holding area for failed jobs.

Nothing in the dead-letter queue is processed automatically. A developer looks at it, fixes the bug or the data, and replays the jobs that still matter. Set up an alert when it starts filling, because a sudden rush of dead letters usually means something broke.

Resources

Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.