Three measures shape every recommendation: what you're seeking, how familiar a passage is, and how hard it is to read. Here's how each is built — and where it's uncertain.
Your reflection becomes a small set of values. We compare their meaning — not their exact words — against the themes of each passage, so a search for "forgiveness" can surface a passage that never uses the word. Nothing is a fixed checklist; the match is open-ended.
We're measuring whether a passage is recognized — the verse on the stadium sign, the line in the film — not whether it's theologically central. We show a language model many small batches of passages and ask which are most and least recognizable, then count how often each is chosen across hundreds of comparisons.
Done broadly, this reliably separates the famous few from the long tail. A second, focused round puts the famous passages up against each other to rank the head. We measure on a widely-read reference text — for the Bible, the World English Bible — so familiarity reflects shared recognition, then carry that score across whichever translation you actually read.
You choose the lens. A verse famous to an evangelical reader isn't always the one a secular or Catholic reader knows — "the bread of life" rises for some, "the truth shall make you free" for others. We score through several lenses and let you weight them. The famous few are famous to everyone; your lens reorders what comes just behind them.
Where it's soft: the model leans slightly toward famous chapters (a well-known opening lifts its neighbors), and recognition is culture-bound — "Jesus wept" is famous in English, less so in translation. We blend in behavioral signals (how often a passage is actually read online) to anchor the head and catch these biases.
Difficulty rests on how common a passage's words actually are in the language you're learning — blended from graded vocabularies (CEFR, HSK) and real-usage frequency (subtitles, web). We weight toward the hardest words, not the average, since one rare word is what stops a learner. Archaic or unusual senses are flagged even when the word looks common.
These are estimates, not verdicts. We keep every source we draw on, run the model against behavioral ground truth rather than trusting it alone, and store each signal separately so you can see — and weight — what went into a recommendation.