Manus logo

Agent Evaluation Engineer

Manus

Singapore; BeijingFull-timePosted Oct 9, 2026

Job description

Manus is a general-purpose AI agent that bridges mind and action, handling complex tasks from start to finish. It enables people to delegate entire workflows across deep research, data analysis, and software development, turning advanced AI into reliable, practical results.

About the Role

Develop evaluations grounded in product needs, user tasks and how AI models and agents work. Use metrics, experiments and failure analysis to assess capability changes and investigate gaps in existing evaluations, informing system development, post-training and model selection.

Responsibilities

Design online and offline metrics that translate user tasks, output quality and practical value into measurable, testable evaluation criteria. Design evaluation tasks and experiments around agent planning, tool use, context and feedback, comparing performance before and after system changes and identifying what affects results.

Support post-training evaluations and comparisons of third-party model quality, defining use cases, metric definitions and the conditions to which results apply. Develop new tasks, metrics or experimental methods for capabilities and user experience issues that existing evaluations miss, and test their validity, bias and reproducibility.

Analyze evaluation results and failures, distinguish score changes from real capability changes, and validate improvements and refine methods with product, engineering and model teams.

Requirements

Practical experience evaluating agent products or model post-training, with concrete evidence of metric design and validation skills. Deep understanding of AI model and agent mechanisms, including task planning, tool use, context management and feedback, and the ability to use this knowledge to analyze system behavior.

Data analysis, engineering and research skills, including the ability to design experiments, process evaluation data and analyze uncertainty in results. Ability to investigate open-ended problems, develop well-reasoned new evaluation methods, examine experimental bias and how samples support conclusions, and test judgments through experiments.

Independent judgment on AI output quality and real user value, with the ability to explain findings, their applicability and limitations clearly. Familiarity with user research methods such as interviews, observation or usability testing, and the ability to translate findings into evaluation tasks and criteria, are preferred.

More jobs at Manus

See all openings at Manus
See all open roles at Manus or explore thousands more companies on TheJobsMap.