Remote job
AI Benchmark Engineer | Native Language Specialist - Arabic (UAE) - Remote
Job details
About this role
Role overview Contribute to a rigorous, verifiable benchmark suite that measures how well large language models handle multilingual software challenges in terminal workflows. This remote, freelance role is open to a native Arabic-speaking software engineer based in the United Arab Emirates.
Responsibilities - Engineer benchmark tasks designed to evaluate coding agents across realistic CLI scenarios - Construct task environments using Arabic-language assets that remain in the target language to genuinely test multilingual handling - Surface AI failure points through Arabic prompting and translation probing - Develop reference solutions and write reliable, deterministic verifier scripts, with rubric-based judging reserved for narrow cases - Analyze execution logs and tune task difficulty from Easy to Very Hard across multiple model sizes - Take part in a four-layer human quality assurance process alongside automated LLM-based checks to safeguard fairness and integrity
Requirements - At least five years of professional software engineering experience - Background at leading technology firms and/or a degree from a top-tier engineering university - Native or near-native Arabic fluency with strong command of grammar and register, plus high English proficiency - Strong Python, shell scripting, and data processing skills - Deep CLI development experience and working familiarity with coding agents - Solid grasp of multilingual text processing issues, including Unicode normalization, locale-dependent formatting, safe string operations, and RTL handling
Nice to have - Experience with typography, font fallback behavior, or non-Latin rendering in user interfaces and artifacts