# Minimal Senior-Level Job Queue This is a deliberately small PostgreSQL-backed job queue for the interview assignment. The important parts are: - one internal priority queue - two tables: `jobs` and `job_events` - atomic claiming with PostgreSQL row locks and `SKIP LOCKED` - deterministic status transitions - automatic retries with exponential backoff - lease renewal for long-running jobs - lease timeout cleanup when workers die - idempotent job creation - demo UI with job state and event polling ## Run ```powershell copy .env.sample .env npm.cmd --prefix frontend install npm.cmd --prefix frontend run build docker compose up --build ``` Open: ```text http://localhost:5173 ``` API docs: ```text http://localhost:8000/api/docs/ ``` Admin panel: ```text http://localhost:8000/admin/ ``` Create an admin user: ```powershell docker compose exec backend python manage.py createsuperuser ``` ## Architecture ```text React UI | Django API ---- PostgreSQL | Django worker process | N worker threads from env ``` There is no queue table and no worker table. Workers are ephemeral process threads with generated ids. The queue is internal and ordered by: ```text priority DESC, available_at ASC, created_at ASC, id ASC ``` ## Statuses ```text queued -> running running -> succeeded running -> queued retry after failure or timeout running -> failed attempts exhausted failed -> queued manual retry ``` The database also validates row shape: - queued jobs cannot have locks or finish timestamps - running jobs must have a lock owner and lease deadline - terminal jobs must have a finish timestamp and no lock ## At-Least-Once Execution This queue provides at-least-once execution, not exactly-once execution. A worker can perform an external side effect and crash before marking a job succeeded. The lease will expire and the job can run again. Real handlers should therefore be idempotent. ## Why PostgreSQL The assignment requires PostgreSQL, and PostgreSQL gives a compact solution for safe concurrent claiming through `SELECT ... FOR UPDATE SKIP LOCKED`. This keeps the implementation transactional, inspectable, and easy to demo. For a high-throughput distributed production queue, Redis-backed systems such as BullMQ or Sidekiq-style designs are common. That is documented as the next architecture, not implemented here. ## Useful Commands Run backend tests locally with SQLite fallback: ```powershell $env:TEST_DATABASE_ENGINE="sqlite" python manage.py test ``` Run pytest: ```powershell $env:TEST_DATABASE_ENGINE="sqlite" python -m pytest ``` Run worker locally: ```powershell python manage.py run_job_workers ``` ## References - PostgreSQL `SKIP LOCKED`: https://www.postgresql.org/docs/current/sql-select.html - pg-boss: https://github.com/timgit/pg-boss - Solid Queue: https://github.com/rails/solid_queue - BullMQ concurrency: https://docs.bullmq.io/guide/workers/concurrency - Distributed task queue article: https://medium.com/@sindhukripa007/i-built-a-distributed-task-queue-from-scratch-to-actually-understand-how-they-work-37fa0452ff9b