Role Overview
We are building benchmarks with vertical AI companies. This role would mostly involve building evals. You will spend most of your time deciding whether an AI system actually did the job, then building the datasets, graders, and checks that proves it. We are a small YC team. You work with the founders. If you want a 9-to-5 or a remote-first job, skip this. If you want to own how we measure things, keep reading.
About Ressl ai
Train, evaluate and build autonomous agents - Ressl AI helps teams bridge the gap from prototype to production by providing the platform needed to train, benchmark, and deploy autonomous agents that actually work in real-world business systems.
Key Responsibilities
- Turn real failures into tasks we can score
- Build datasets, gold labels, graders, and human review loops
- Find ways to prevent reward hacking
- Find LLM incapabilities in production systems
Requirements & Eligibility
- You have shipped agents.
- You have built evals, graders, or datasets for LLMs or agents
- You can sit with a lot of traces and find the few that matter
Required Skills & Tech
EvalsAI AgentsLLMsGradersDatasets