Research
Research Scientist
Work on the hardest question here: how do you measure whether a person actually understood the passage?
Apply for this roleStandard language model benchmarks tell us almost nothing about whether Trinity taught someone well. This role builds the science we are missing: evaluations for teaching quality, faithfulness to a cited text, and the difference between an answer a reader accepted and one they understood. Expect to publish some of it.
What you will do
Design and run evaluations for faithfulness, citation accuracy, and pedagogical quality.
Study where the system fails on genuinely hard passages and translate that into shipped fixes.
Build the human-evaluation process, including recruiting readers across traditions and reading levels.
Work with the content lead to turn theological standards into measurable criteria.
Write up what we learn, internally always and externally where it is useful to others.
What we are looking for
Graduate-level research training in machine learning, NLP, cognitive science, education, or a near neighbour.
Experience evaluating language models beyond leaderboard scores.
Strong Python and the discipline to make an experiment reproducible by someone else.
Willingness to work on a problem where the metric does not exist yet.
Nice to have
Background in learning science or assessment design.
Experience with retrieval-augmented systems and their failure modes.
Published work on evaluation methodology.
How to apply
Email hello@gettrinity.app with a short note about why this role, and anything you have made that you are still proud of. A link is worth more than a cover letter. We read every message and reply either way.