Remote job
Prompt Engineer & AI Evaluator
Job details
About this role
Role overview This role focuses on shaping how enterprise large language models behave at scale, combining prompt design with rigorous evaluation work. The successful candidate will iterate on system prompts, design structured tests, and surface subtle failure modes such as hallucinations, unsafe outputs, and inconsistent formatting before they reach end users.
Responsibilities - Design, draft, and refine system prompts for production LLM features across multiple use cases. - Build and run systematic evaluation suites that measure accuracy, safety, and output consistency over time. - Investigate model failures, reproduce edge cases, and document findings for engineering and product teams. - Apply version control and structured formats (JSON, Markdown) when iterating on prompts and eval datasets. - Use basic Python scripting to automate test runs, parse model outputs, and aggregate results.
Requirements - 1+ years of hands-on experience with prompt engineering or AI evaluation in a production environment. - Fluent or native-level English (C1–C2), since the majority of prompting and evals are conducted in English. - Strong analytical and critical thinking skills, with the ability to spot subtle errors, bias, or hallucinations. - Working familiarity with JSON, Markdown, and basic Python scripting for automation and data handling. - Indonesian citizenship and current residence in Indonesia; the position is 100% full remote.
Nice to have Experience with LLM APIs, evaluation frameworks, or annotation tools used in NLP quality work.