Remote job
AI Benchmark Engineer | Native Language Specialist - Arabic (Saudi Arabia) - Remote
Job details
About this role
Role overview Help build a verifiable evaluation suite of Terminal-Bench tasks that stress-tests large language models on multilingual software challenges. This is a remote, freelance engagement for a native Arabic-speaking software engineer based in Saudi Arabia.
Responsibilities - Design and engineer benchmark tasks that evaluate coding agents in terminal-based workflows - Build realistic task environments using Arabic-language datasets and files that remain in the target language - Probe failure points where AI does not work, using native Arabic prompting and translation analysis - Develop reference implementations and write deterministic verifier scripts, using rubric-based judging only when strictly necessary - Analyze execution logs and calibrate task difficulty from Easy to Very Hard across multiple model tiers - Participate in a four-layer human quality control process alongside automated LLM-based checks
Requirements - Five or more years of industry software engineering experience - Track record at leading technology companies and/or a degree from a top-tier engineering university - Native or near-native Arabic fluency with strong grasp of grammar, register, and phrasing, plus high English proficiency - Strong proficiency in Python, standard shell scripting, and data processing - Extensive terminal and CLI-based development experience, with working familiarity with coding agents - Deep understanding of multilingual text processing pitfalls, including Unicode normalization, locale-dependent conventions, text I/O, and bidirectional or RTL handling
Nice to have - Familiarity with font fallback behavior and rendering or typography in user interfaces and artifacts - Comfort designing benchmark tasks that intentionally avoid English translation crutches