الوصف
الدور
نحن نبحث عن مهندس تعلم آلي كبير عملي لقيادة تطوير وتشغيل أنظمة الذكاء الاصطناعي عالية الإنتاجية عبر نماذج اللغة الكبيرة (LLMs) والتعرّف البصري على الأحرف (OCR) والصوت. يجمع هذا الدور بين الملكية التقنية العميقة وقيادة فريق عالي الأداء من المهندسين المبتدئين، بالإضافة إلى التواصل المباشر مع أصحاب المصلحة لتشكيل حلول الذكاء الاصطناعي.
المسؤوليات:
- قيادة وتوجيه فريق من مهندسي ML الموهوبين عبر:
• مراجعات الشفرة، ومراجعات التصميم، والتوجيه التقني
• تطبيق ممارسات هندسة البرمجيات وتعلم الآلة القوية
- تصميم ونشر وتشغيل أنظمة ذكاء اصطناعي قابلة للتوسع مع التركيز على الموثوقية والأداء
- قيادة نشر الإنتاج لنماذج LLMs وأنظمة متعددة الوسائط (RAG، OCR، الصوت)
- امتلاك أداء النموذج نهاية إلى نهاية، مع الجمع بين التقييم، الرصد، وتحسين العتاد:
• بناء خطوط تقييم (معايير المقارنة، اختبارات الرجوع، LLM كحكم)
• تنفيذ رصد عميق (التعقب، الكمون، تتبع الأخطاء)
• تحسين استخدام GPU (خدمة متعددة-GPUs، التجميع، التكميم، ضبط الذاكرة)
• تحسين الإنتاجية والكمون والكفاءة من حيث التكلفة باستمرار
- تصميم وإدارة بنية GPU التحتية:
• تقديم النماذج، توزيع الحمولة، واستراتيجيات التوسع
• النشر الواعي للمعدات وضبط الأداء
- بناء وصيانة خطوط MLOps قوية:
• إدارة النماذج/الإصدارات، CI/CD، الاختبار الآلي، واستراتيجيات التراجع
• المراقبة ودورات التغذية المرتدة للتحسين المستمر
- التفاعل المباشر مع العملاء وأصحاب المصلحة لـ:
• جمع وتClarifying متطلبات العمل
• ترجمة الاحتياجات غير التقنية إلى مشاكل تقنية محددة جيداً
• توصيل الحلول والتوازنات والتقدم من خلال وثائق وتقارير ومقترحات واضحة
- الإسهام بنشاط في تصميم النظام والتنفيذ والتصحيح واحتواء الحوادث الإنتاجية
المتطلبات
- خبرة مثبتة في نشر نماذج LLMs في الإنتاج
- خبرة قوية في الاستنتاج المعتمد على GPU والتحسين
- مهارات هندسة خلفية قوية (Python، APIs، أنظمة موزعة)
- خبرة في MLOps وأنظمة ML في الإنتاج
- خبرة في OCR/ذكاء اصطناعي للمستندات و/أو أنظمة الصوت (STT/TTS)
- خبرة مع Docker وKubernetes
- فهم قوي لهندسات الذكاء الاصطناعي الحديثة (RAG، قواعد بيانات المتجهات، سلاسل الوكلاء)
- خبرة في توجيه أو قيادة المهندسين
- مهارات تواصل قوية مع قدرة على ربط مجالات الأعمال والتقنية
يفضّل وجوده:
- خبرة مع النماذج مفتوحة الوزن (Qwen، Llama، DeepSeek، Gemma)
- خبرة في النشر المحلي / الذكاء الاصطناعي السيادي
- خبرة مع LoRA / الضبط الدقيق
- خبرة في اللغات المتعددة أو NLP العربية
Description
The role
We are looking for a hands-on senior ML engineer to lead the development and operation of production-grade AI systems across LLMs, OCR, and voice. This role combines deep technical ownership with leadership of a high-performing team of junior engineers, as well as direct engagement with stakeholders to shape AI solutions.
Responsibilities:
- Lead and mentor a team of highly talented junior ML engineers through:
• Code reviews, design reviews, and technical direction
• Enforcement of strong software engineering and ML best practices
- Design, deploy, and operate scalable AI systems with a focus on reliability and performance
- Lead production deployment of LLMs and multimodal systems (RAG, OCR, voice)
- Own model performance end-to-end, combining evaluation, observability, and hardware optimization:
• Build evaluation pipelines (benchmarks, regression testing, LLM-as-judge)
• Implement deep observability (tracing, latency, error tracking)
• Optimize GPU utilization (multi-GPU serving, batching, quantization, memory tuning)
• Continuously improve throughput, latency, and cost efficiency
- Architect and manage GPU infrastructure:
• Model serving, load balancing, and scaling strategies
• Hardware-aware deployment and performance tuning
- Build and maintain robust MLOps pipelines:
• Model/version management, CI/CD, automated testing, and rollback strategies
• Monitoring and feedback loops for continuous improvement
- Engage directly with clients and stakeholders to:
• Gather and clarify business requirements
• Translate non-technical needs into well-defined technical problems
• Communicate solutions, trade-offs, and progress through clear documentation, reports, and proposals
- Contribute hands-on to system design, implementation, debugging, and production incident resolution
Requirements
- Proven experience deploying LLMs in production
- Strong experience with GPU-based inference and optimization
- Solid backend engineering skills (Python, APIs, distributed systems)
- Experience with MLOps and production ML systems
- Experience with OCR/document AI and/or voice systems (STT/TTS)
- Experience with Docker and Kubernetes
- Strong understanding of modern AI architectures (RAG, vector DBs, agent workflows)
- Experience mentoring or leading engineers
- Strong communication skills with the ability to bridge business and technical domains
Nice to have:
- Experience with open-weight models (Qwen, Llama, DeepSeek, Gemma)
- Experience with on-prem / sovereign AI deployments
- Experience with LoRA / fine-tuning
- Multilingual or Arabic NLP experience