Agent Evaluation Engineer
Turn agent quality into reproducible evidence across repeated runs, adversarial inputs and changing models.
- Location
- Kyiv / remote
- Experience
- Mid / Senior
- Working arrangement
- Arrangement to be discussed
Requirements
- Experience testing complex software or evaluating ML systems.
- Ability to analyse noisy results and distinguish model, data and harness failures.
- Strong Python skills and clear written communication of experimental findings.
Responsibilities
- Create task suites, scoring rules and reproducible execution environments.
- Investigate incorrect tool calls, data leakage, brittle recovery and inconsistent outcomes.
- Build regression checks that connect model behaviour to release decisions.
Useful evidence
- An evaluation or test suite you designed and a failure it exposed.
- Experience with security testing, statistics or agent benchmarks is useful.
Working terms
- Startup culture, a goal-oriented team, and a research mindset
- The opportunity to apply your engineering skills to tools and systems for fellow engineers and help shape the future of AI
- Latest-generation MacBook Pro
- An in-house GPU cluster for training and experimentation
- 20 working days of annual leave
- English courses, educational events, and conferences
- Medical insurance
Tools & systems
PythonInspect AIBraintrustpytestDocker
Relevant experience matters more than knowing every tool listed.