Ergon is in stealth, so this posting describes the shape of the work rather than the thing being built. That is not coyness for its own sake: on the first call we tell you exactly what it is, including the claim you would be asked to test, before you spend any real time on us.
We make a strong claim about outcomes. That claim is worth nothing until it is measured properly, against baselines chosen to be hard rather than flattering. Our research lead sets the questions. This role builds and runs the experiments that answer them, and reports what they say on the days when the answer is inconvenient for everyone else in the company.
What you will do
- Design evaluations against honest baselines, not strawmen
- Build the harness, not just read the numbers
- Report effect sizes and uncertainty rather than headline wins
- Keep a durable record of what works, so the company learns instead of re-deciding
- Say plainly when a change made things worse, to the founders and to the room
What we are looking for
- Experimental design and statistics you can defend, including where the design is weak
- Enough engineering to build the harness yourself
- Skepticism as a working method rather than as a personality
- Clear writing. Results nobody reads do not count.
Useful, not required
- Published or internal work on LLM evaluation and its failure modes
- Human annotation pipelines and inter-rater reliability
How we hire
- 01A 45-minute call with the CTO.
- 02A discussion of an evaluation you designed, including what you would change.
- 03A working session designing an evaluation for one of our open questions.
- 04A conversation with our research lead about what you would measure first.
- 05References, then an offer.
Ergon is an equal opportunity employer. We consider all qualified applicants without regard to any characteristic protected by law.
What you send is used to evaluate this application and nothing else, and it is deleted on a schedule. You can ask us to delete it sooner. Ask here.