Remote job
AI Benchmark Engineer | Native Language Specialist - Chinese (Taiwan) - Remote
Job details
About this role
Role overview Design and validate multilingual benchmark tasks that push large language models on real-world software challenges in terminal environments. This is a remote, freelance opportunity aimed at a native Mandarin Chinese speaker based in Taiwan.
Responsibilities - Build benchmark tasks that evaluate coding agents in realistic CLI workflows - Create task environments with Chinese-language datasets and files that stay in the target language to genuinely probe multilingual handling - Identify AI failure points by designing Chinese-language prompts and translation probes - Write reference implementations and deterministic verifier scripts, falling back to rubric-based judging only when unavoidable - Calibrate task difficulty across Easy to Very Hard levels by analyzing execution logs against several model tiers - Contribute to a four-layer human quality control process paired with automated LLM-based checks for fairness and integrity
Requirements - Five or more years of professional software engineering experience - Track record at leading technology companies and/or a degree from a top-tier engineering university - Native or near-native Mandarin Chinese fluency with strong understanding of grammar, register, and phrasing, plus high English proficiency - Strong proficiency in Python, standard shell scripting, and data processing - Extensive terminal and CLI workflow experience, with working knowledge of coding agents - Deep knowledge of multilingual text processing pitfalls, including Unicode normalization, locale conventions, text I/O, and CJK or font fallback considerations
Nice to have - Familiarity with typography and rendering concerns for non-Latin scripts in user interfaces or artifacts