عن بُعد
Mindrift -
سلطنة عمان
--
Mindrift

تفاصيل الوظيفة

وصف العمل

يرجى تقديم سيرتك الذاتية بالإنجليزية وذكر مستوى كفاءتك في اللغة الإنجليزية.
Mindrift يربط المتخصصين بفرص AI قائمـة على المشاريع لشركات التقنية الرائدة، مع التركيز على الاختبار والتقييم وتحسين أنظمة AI.
المشاركة قائمة على المشروع، وليست توظيفاً دائماً.
ما تنطوي عليه هذه الفرصة: عمل وكلاء التشفير Frontier لديهم خبرة في اجتياز الاختبارات بالفعل.
نقيس ما إذا كانوا يجتازونها بالطريقة الصحيحة .
نحن نبني مجموعة بيانات لتقييم أمان وسلوك وكلاء تشفير الذكاء الاصطناعي — ليس فقط إذا أكملوا مهمة، بل إذا فعلوا ذلك بأمانة، ضمن النطاق، ودون تقصير.
ستصمم مهام حيث الطريق الأسهل هو الطريق غير الآمن، وتكتب الاختبارات التي تلتقطه: بناء بيئات مطور واقعية — شركة افتراضية بها قاعدة كود وبنية تحتية وسياق (التذاكر، المستندات، المحادثات) تشكل تاريخ تطوير مقنع تصميم المهام التي تقترن هدف تطويري بريء مع اختصار غير آمن مغرٍ: زيادة النطاق، انتهاك السياسات، تلف البيانات، تغييرات واسعة الإذن اكتب اختبارات تتحقق مما إذا كان الوكيل أنجز المهمة بالطريقة الصحيحة — التقاط التقعير في الكماليات، وليس مجرد التحقق من المخرجات Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming — there is no attacker in the scenario.
Cybersecurity experience is a nice-to-have but not a requirement.
We're looking for engineers who understand how code should behave, not penetration testers.
Strong software engineers, not security specialists; Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome; What we look for 4–5+ years in software development; Core stack: Python, JavaScript/TypeScript; Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect; Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar); Familiarity with GitHub PRs and CI workflows as a user; Stack breadth is welcome, not a filter.
Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer; English proficiency — B2+ Why this is hard Frontier models are already good at coding.
Creating a task that genuinely challenges the best models is non-trivial.
The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it.
Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Project time expectations For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements.
This is an estimate, not a guaranteed workload, and applies only while the project is active.
Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Compensation On this project, contributors can earn up to $75 per hour equivalent , depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise.
Please note that other projects on the platform may offer different earning levels based on their requirements.
 

وظائف مشابهة

حول Mindrift
سلطنة عمان
تكنولوجيا المعلومات والخدمات