Skip to content

Research

Research Scientist

Work on the hardest question here: how do you measure whether a person actually understood the passage?

Full-time

Remote, Americas or Europe

Apply for this role

Standard language model benchmarks tell us almost nothing about whether Trinity taught someone well. This role builds the science we are missing: evaluations for teaching quality, faithfulness to a cited text, and the difference between an answer a reader accepted and one they understood. Expect to publish some of it.

What you will do

  • Design and run evaluations for faithfulness, citation accuracy, and pedagogical quality.

  • Study where the system fails on genuinely hard passages and translate that into shipped fixes.

  • Build the human-evaluation process, including recruiting readers across traditions and reading levels.

  • Work with the content lead to turn theological standards into measurable criteria.

  • Write up what we learn, internally always and externally where it is useful to others.

What we are looking for

  • Graduate-level research training in machine learning, NLP, cognitive science, education, or a near neighbour.

  • Experience evaluating language models beyond leaderboard scores.

  • Strong Python and the discipline to make an experiment reproducible by someone else.

  • Willingness to work on a problem where the metric does not exist yet.

Nice to have

  • Background in learning science or assessment design.

  • Experience with retrieval-augmented systems and their failure modes.

  • Published work on evaluation methodology.

How to apply

Email hello@gettrinity.app with a short note about why this role, and anything you have made that you are still proud of. A link is worth more than a cover letter. We read every message and reply either way.