وصف الوظيفة
الوصف الوظيفي
حول الدور
نبحث عن مهندس DevOps لبناء الأساس التخطيطي للبنية التحتية الذي يجعل منتجنا المعزز بالذكاء الاصطناعي سريعًا وموثوقًا وآمنًا وفعال من حيث التكلفة. ستقوم بإعداد السحابة وأنظمة النشر والمراقبة التي تسمح لفريق صغير بالإطلاق بثقة وتوسيع النطاق بسلاسة خلال فترة التشغيل الأولي وما بعدها.
ما ستقوم به
- تصميم وتوفير وإدارة بنية سحابية (AWS، GCP، أو Azure)
- بناء وصيانة خطوط أنابيب CI/CD للنشرات السريعة والآمنة والم automatable
- تنفيذ مفهوم البنية كرمز (Terraform، Pulumi، أو مماثل) لبيئات قابلة لإعادة الإنتاج
- تحويل الخدمات إلى حاويات وتنسيقها (Docker، Kubernetes) مع نمو النظام
- إعداد الرصد: المراقبة، التسجيل، التنبيهات، والتتبع عبر المكدس
- إدارة وتحسين الكلفة والكمون والموثوقية لأعباء العمل المتعلقة بـ LLM والذكاء الاصطناعي
- تملك الأمن، إدارة الأسرار، والتحكم في الوصول عبر البيئات
- تأسيس ممارسات النسخ الاحتياطي والتعافي من الكوارث والاستجابة للحوادث للمخطط التجريبي
المؤهلات
- 4+ سنوات خبرة في DevOps أو SRE أو هندسة البنية التحتية
- خبرة عملية قوية مع مزود سحابة رئيسي واحد على الأقل (AWS/GCP/Azure)
- إتقان أدوات البنية كرمز (Terraform، Pulumi، CloudFormation)
- خبرة في الحاويات والتنسيق (Docker، Kubernetes)
- خبرة قوية في CI/CD (GitHub Actions، GitLab CI، CircleCI، أو ما يُشابه)
- مهارات برمجة قوية (Bash، Python، أو Go)
- خبرة في تنفيذ أدوات المراقبة والرصد (Prometheus، Grafana، Datadog، إلخ)
- عقلية أولوية الأمن وخبرة في إدارة الأسرار والولوج
معلومات إضافية
- خبرة في إدارة بنية تحتية للذكاء الاصطناعي/التعلم الآلي أو أعباء عمل تركز على LLM
- الإلمام بتحسين التكاليف لاستخدام واجهات برمجة تطبيقات عالية الإنتاجية
- خبرة في بنية GPU التحتية أو خدمة الاستدلال
- خبرة مبكرة في بناء بنية تحتية من الصفر
Job description
Job Description
About the Role
We're looking for a DevOps Engineer to build the infrastructure foundation that keeps our AI-powered product fast, reliable, secure, and cost-efficient. You'll set up the cloud, deployment, and monitoring systems that let a small team ship confidently and scale smoothly through the pilot and beyond.
What You'll Do
- Design, provision, and manage cloud infrastructure (AWS, GCP, or Azure)
- Build and maintain CI/CD pipelines for fast, safe, automated deployments
- Implement infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible environments
- Containerize and orchestrate services (Docker, Kubernetes) as the system grows
- Set up observability: monitoring, logging, alerting, and tracing across the stack
- Manage and optimize the cost, latency, and reliability of LLM and AI workloads
- Own security, secrets management, and access controls across environments
- Establish backup, disaster-recovery, and incident-response practices for the pilot
Qualifications
- 4+ years of DevOps, SRE, or infrastructure engineering experience
- Strong hands-on experience with at least one major cloud provider (AWS/GCP/Azure)
- Proficiency with infrastructure-as-code tools (Terraform, Pulumi, CloudFormation)
- Experience with containerization and orchestration (Docker, Kubernetes)
- Solid CI/CD experience (GitHub Actions, GitLab CI, CircleCI, or similar)
- Strong scripting skills (Bash, Python, or Go)
- Experience implementing monitoring and observability tooling (Prometheus, Grafana, Datadog, etc.)
- Security-first mindset and experience with secrets and access management
Additional Information
- Experience managing infrastructure for AI/ML or LLM-heavy workloads
- Familiarity with cost optimization for high-throughput API usage
- Experience with GPU infrastructure or inference serving
- Early-stage experience standing up infrastructure from scratch