Research
Research Engineer, Evaluation
- Department
- Research
- Location
- Remote (UK / EU time zones)
- Work mode
- Remote
- Type
- Full-time
About ORYN
ORYN builds AI systems that understand, reason and take action. We are a small team working on hard problems in reasoning, knowledge and infrastructure, and we care about building things that hold up under real use.
About the role
Evaluation is the part of this field most likely to be done badly, and the part that most determines whether anything improves. You will own how we know whether a change to a reasoning system made it better.
That means building harnesses, designing tasks that measure something real, and being the person who says the improvement is within noise when it is.
What you will do
- Design evaluations that measure capability on the work our systems actually do.
- Build and maintain the harness that runs them, including the unglamorous parts.
- Report results honestly, including the ones that say a change did nothing.
- Push back when a metric is being optimised at the expense of the thing it proxies for.
What we are looking for
- Experience evaluating machine learning or language systems beyond standard benchmarks.
- Statistical literacy — you know what a confidence interval is for and when a result is noise.
- Enough engineering to build your own tooling rather than waiting for someone else to.
Nice to have
Not requirements. Apply if the section above describes you.
- Published work on evaluation methodology, or a strong opinion about why most of it is inadequate.
- Experience with human evaluation protocols and their failure modes.
Apply
Apply for Research Engineer, Evaluation
We read every application. If the fit is not right we will tell you, rather than leaving you wondering.