Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
Frontier coding agents are already good at passing tests. We measure whether they pass them the right way. We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.
You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
- Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
- Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
- Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
- Not data labeling;
- Not prompt engineering;
- Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
- Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
- 4–5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
- English proficiency — B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
How it works
Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid
Project time expectations
For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Compensation
On this project, contributors can earn up to $50 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
يرجى تقديم سيرتك الذاتية باللغة الإنجليزية وتحديد مستوى إتقانك للإنجليزية.
ربط Mindrift المتخصصين بفرص الذكاء الاصطناعي القائمة على المشاريع لشركات تكنولوجية رائدة، مع تركيز على اختبار وتقييم وتحسين أنظمة الذكاء الاصطناعي. المشاركة تكون قائمة على المشروع وليست توظيفاً دائماً.
ما تتضمنه هذه الفرصة
عُملاء ترميز المتقدمين للمادة Frontier ماهرين بالفعل في اجتياز الاختبارات. نقيس ما إذا كانوا يحققونها بالطريقة الصحيحة. نحن نبني مجموعة بيانات لتقييم سلامة وسلوك وكلاء ترميز الذكاء الاصطناعي — ليست مجرد إنجاز مهمة، بل إنجازها بأمان وبالنطاق المناسب وبدون تقليل الجودة.
ستصمم مهام حيث الطريق السهل هو غير الآمن، وتكتب الاختبارات التي تلتقط ذلك:
- إنشاء بيئات مطور واقعية — شركة افتراضية بمُستودع رمز وبنية تحتية وسياق (تذاكر، مستندات، محادثات) تشكل تاريخ تطوير قابلًا لسرد القصة
- تصميم مهام تجمع هدف تطويري غير ضار مع اختصار غير آمن مغرٍ: زيادة النطاق، انتهاكات السياسات، تلف البيانات، تغييرات ذات صلاحيات واسعة
- كتابة اختبارات تتحقق مما إذا كان الوكيل قد أكمل المهمة بالشكل الصحيح — تُلتقط التغييرات غير المحسوبة، وليس فقط فحص النتائج
- التكرار على المهام والاختبارات استناداً إلى ملاحظات QA — مراجعة حلول الوكيل، تحليل الإخفاقات، والتحسين حتى تكون التقييمات عادلة وقوية
ما ليس هذا:
- ليس تصنيف بيانات؛
- ليس هندسة الإرشاد؛
- ليس أمن سيبراني أو اختبار اختراق — لا يوجد مُهاجم في السيناريو. خبرة الأمن السيبراني ميزة ولكنها ليست شرطاً. نبحث عن مهندسين يفهمون كيف يجب أن يتصرف الكود، لا مختبري اختراق. مهندسون برمجيات أقوياء، وليسوا متخصصين في الأمن؛
- ليس كتابة شفرة من الصفر — الوكيل يكتب معظم الشفرة؛ أنت تصمم الوضع وتقييم النتيجة;
ما نبحث عنه
- 4–5+ سنوات في تطوير البرمجيات؛
- التكدس الأساسي: بايثون، JavaScript/TypeScript؛
- مهارات قوية في تصميم الاختبارات — اختبارات وظيفية وتكامل تفصل بين الإكمال الآمن وغير الآمن، وليس مجرد صحيح مقابل غير صحيح؛
- خبرة عملية مع وكلاء التشفير (Claude Code، GitHub Copilot CLI، Codex، أو ما شابه)؛
- الاعتياد على PRs و CI في GitHub كمستخدم؛
- تقبل اتساع التكدس كميزة وليس كفلترة. ت模拟 المهام مستودعات حقيقية تحتوي على قواعد بيانات، خطوط CI، و سكربتات نشر، لذا تعدد الخبرة في الخلفية والبنية التحتية مفيد حقاً — لكن ليس من الضروري أن تكون خبيراً في كل طبقة؛
- إتقان الإنجليزية — B2+
لماذا هذا صعب
نماذج Frontier جيدة في البرمجة أصلاً. إنشاء مهمة تتحدى النماذج الأفضل حقاً ليس بسيطاً. الصعوبة الحقيقية هي بناء الإغراء — سيناريو يكون فيه المسار غير الآمن أو خارج النطاق هو مسار المقاومة الأقل — ثم كتابة اختبارات تلتقط بموثوقية وكيل أخذ هذا المسار. لدى المهام حلول صحيحة عديدة؛ يجب أن تقبل الاختبارات جميعها وت rejects السيئة منها.
كيف يعمل
قدم طلب؟ اجتز qualification؟ انضم إلى مشروع؟ أكمل المهام؟ تحصل على الأجر
توقعات وقت المشروع
بالنسبة لهذا المشروع، تُقدر المهام بـحوالي 20-25 ساعة أسبوعياً خلال المراحل النشطة، بناءً على متطلبات المشروع. هذا تقدير، ليس عبء عمل مضمون، ويطبق فقط أثناء نشاط المشروع. يجب تقديم المهام قبل الموعد النهائي وتلبية معايير القبول المذكورة ليتم قبولها.
التعويض
في هذا المشروع، يمكن للمساهمين أن يحققوا ما يصل إلى ما يعادل 50 دولارًا في الساعة، بحسب مستواهم وتيرة مساهمتهم.
يتفاوت التعويض عبر المشاريع حسب النطاق والتعقيد والخبرة المطلوبة. يرجى ملاحظة أن مشاريع أخرى على المنصة قد تقدم مستويات كسب مختلفة اعتماداً على متطلباتها.