DevOps Engineer Needed: Diagnose & Fix Docker Container Crash-Loop
Бюджет: $35.0 - $60.0
HOURLY / PART_TIME
⭐ 0.00 (0)
United Arab Emirates
amazon-web-services, containerization, docker, python, amazon-ec2, troubleshooting, docker-compose, linux-system-administration
Предпочтительная квалификация
- Опыт: Эксперт
We have a multi-container Docker Compose application running on an AWS EC2 instance. One of the containers, the worker, starts normally but crashes after approximately 15–90 seconds and is automatically restarted by Docker.
We’re looking for an experienced DevOps/Docker engineer who can quickly reproduce the issue, identify the root cause, implement a permanent fix, and document the solution.
SSH access to the AWS EC2 instance will be provided immediately.
Environment:
- AWS EC2 instance
- Ubuntu 22.04 LTS
- Docker Engine 24.x
- Docker Compose v2
- 2 vCPU / 4 GB RAM
- Docker Compose stack includes:
api — Python / FastAPI — stable
worker — Python / Celery — affected by the issue
redis — stable
postgres — stable
Current Issue:
- All containers start successfully.
- The worker runs for approximately 15–90 seconds and then exits.
- Docker automatically restarts the worker.
- The worker sometimes exits with code 137.
- Worker logs show normal startup and task processing before stopping, without a clear application error.
- Increasing available memory reduces the frequency of the issue but does not completely resolve it.
- Host-level logs show possible memory-related events around some of the crashes.
- Redis, PostgreSQL, and the API container remain healthy.
What We've Already Checked:
- Verified Redis connectivity and Celery configuration.
- Confirmed the Docker image builds successfully.
- Confirmed the worker can run successfully in a simpler environment.
- Checked for obvious process/PID issues.
- Confirmed the issue is not caused by a Docker healthcheck.
What You'll Be Given:
- SSH access to a dedicated AWS EC2 instance
- Docker Compose configuration
- Dockerfile
- Relevant environment configuration
- Worker logs
- Host system logs
If additional files or configuration are required during the investigation, we can provide them.
What We Need You To Do:
- Reproduce the issue on the provided EC2 instance.
- Identify the actual root cause using evidence from Docker, Linux, and application logs.
- Implement a permanent fix.
- Verify that the worker runs reliably without entering the crash/restart loop.
- Provide a short explanation of the root cause, fix, and any relevant recommendations.
Deliverables:
- Working fix
- Updated configuration/code or patch as required
- Brief root-cause explanation
- Confirmation that the issue has been resolved
Required Skills:
- Strong Docker & Docker Compose experience
- Linux troubleshooting
- Experience diagnosing Docker container crashes and restart loops
- Understanding of Linux resource and memory issues
- Python/Celery experience
- Ability to troubleshoot systematically using logs and system-level tools
Ideal Candidate:
We're looking for someone who can quickly diagnose technical issues rather than simply trying random configuration changes.
Please include a brief example of a similar Docker, Linux, Celery, or container troubleshooting issue you've successfully diagnosed and fixed.
Estimated Effort:
Intermediate-Advanced
Expected effort: 3–4 hours.
Important:
This is primarily a troubleshooting and root-cause analysis task. We need someone who can identify why the container is crashing, provide evidence for the diagnosis, implement the appropriate fix, and verify the result.
Открыть заказ
AI-черновик отклика
Короткий текст отклика для копирования в оффер: интерес + готовность работать.
Войдите, чтобы сгенерировать AI-черновик.
Войти