Minimal Senior-Level Job Queue
This is a deliberately small PostgreSQL-backed job queue for the interview assignment.
The important parts are:
- one internal priority queue
- two tables:
jobsandjob_events - atomic claiming with PostgreSQL row locks and
SKIP LOCKED - deterministic status transitions
- automatic retries with exponential backoff
- lease renewal for long-running jobs
- lease timeout cleanup when workers die
- idempotent job creation
- demo UI with job state and event polling
Run
copy .env.sample .env
npm.cmd --prefix frontend install
npm.cmd --prefix frontend run build
docker compose up --build
Open:
http://localhost:5173
API docs:
http://localhost:8000/api/docs/
Admin panel:
http://localhost:8000/admin/
Create an admin user:
docker compose exec backend python manage.py createsuperuser
Worker logs default to INFO. To see every poll, claim, lease renewal, progress update, retry, and completion, set:
JOB_WORKER_LOG_LEVEL=DEBUG
Then recreate the worker:
docker compose up -d --build worker
docker compose logs -f worker
Architecture
React UI
|
Django API ---- PostgreSQL
|
Django worker process
|
N worker threads from env
There is no queue table and no worker table. Workers are ephemeral process threads with generated ids. The queue is internal and ordered by:
priority DESC, available_at ASC, created_at ASC, id ASC
Statuses
queued -> running
running -> succeeded
running -> queued retry after failure or timeout
running -> failed attempts exhausted
failed -> queued manual retry
The database also validates row shape:
- queued jobs cannot have locks or finish timestamps
- running jobs must have a lock owner and lease deadline
- terminal jobs must have a finish timestamp and no lock
At-Least-Once Execution
This queue provides at-least-once execution, not exactly-once execution.
A worker can perform an external side effect and crash before marking a job succeeded. The lease will expire and the job can run again. Real handlers should therefore be idempotent.
Why PostgreSQL
The assignment requires PostgreSQL, and PostgreSQL gives a compact solution for safe concurrent claiming through SELECT ... FOR UPDATE SKIP LOCKED. This keeps the implementation transactional, inspectable, and easy to demo.
For a high-throughput distributed production queue, Redis-backed systems such as BullMQ or Sidekiq-style designs are common. That is documented as the next architecture, not implemented here.
Useful Commands
Run backend tests locally with SQLite fallback:
$env:TEST_DATABASE_ENGINE="sqlite"
python manage.py test
Run pytest:
$env:TEST_DATABASE_ENGINE="sqlite"
python -m pytest
Run worker locally:
python manage.py run_job_workers
References
- PostgreSQL
SKIP LOCKED: https://www.postgresql.org/docs/current/sql-select.html - pg-boss: https://github.com/timgit/pg-boss
- Solid Queue: https://github.com/rails/solid_queue
- BullMQ concurrency: https://docs.bullmq.io/guide/workers/concurrency
- Distributed task queue article: https://medium.com/@sindhukripa007/i-built-a-distributed-task-queue-from-scratch-to-actually-understand-how-they-work-37fa0452ff9b