How It's Built

Behind the reader is a data project: it links published translations, source-language morphology, manuscript signals, commentary, and cross-references to the words on the page.

The most distinctive technical layer is a translation-specific reverse interlinear: word-level alignment between the English text a reader chose and the original Hebrew, Aramaic, or Greek beneath it — keyed to Strong's numbers so every link is stable, comparable, and lockable.

OpenScripture data path

Evidence to alignment to reader

1

Source data

Hebrew, Aramaic, Greek, Latin, English, manuscript signals, notes, canons, and licensing terms all enter the system as separate evidence streams.

2

Alignment graph

OpenScripture builds a reverse interlinear for every supported translation: reviewed links between each Strong's-numbered original-language word and each translation's actual English wording.

3

Reader signals

The app uses those links for interlinear study, Word Locks, translation comparison, certainty signals, and AI translation context.

Why This Matters

Many Bible apps make digital text fast and searchable. OpenScripture is aiming at a different layer of accessibility: helping readers see why translations differ, how English words connect back to the source languages, and where manuscript evidence is more or less settled.

That means treating technical infrastructure as part of the mission. A tap can open the original-language word. A small marker can explain a real translation difference. A personal Word Lock can help someone learn recurring vocabulary while reading naturally.

The alignment spine: Strong's numbers and manuscripts

A traditional interlinear Bible keeps the original Hebrew or Greek word order and places English beneath it — useful to a linguist, disorienting to most readers. A reverse interlinear works the other way: the English text stays exactly as the translator wrote it, and the matching original-language data appears beneath each word in natural reading order. Building this per-translation, for many different English Bibles, is the central technical challenge of OpenScripture.

The shared backbone that makes per-translation alignment possible is Strong's numbers — a stable lexical ID system introduced by James Strong in 1890 and still the most widely used cross-translation linking scheme in Bible software. A Strong's number stays constant across translations: whether an English Bible renders a word as "LORD" or "Yahweh," both point to the same Hebrew lexeme (H3068). That shared key is what makes Word Locks work: lock a Strong's number once, and the preference propagates wherever that word appears, across every translation on the same alignment spine. Strong's numbers are paired with morphological tagging — part of speech, grammatical case, verb tense, and stem — that describes how a word functions in its sentence, not just what it means.

Three manuscript foundations

A translation should be aligned to the text it was actually translated from. Most Protestant English Bibles follow the Hebrew Masoretic Text and the Greek New Testament — but translations descended from the Greek Septuagint or the Latin Vulgate need their own layer. Forcing them onto the Hebrew spine produces false links.

Hebrew & Aramaic OT

Westminster Leningrad Codex (Masoretic Text)

Morphology from the OpenScriptures Hebrew Bible (OSHB / morphhb), CC BY 4.0. The traditional Hebrew source for most Protestant English Bibles.

Greek NT

SBL Greek New Testament (SBLGNT)

Edited by Michael W. Holmes (SBL & Logos). Paired with MorphGNT morphology: every Koine Greek word tagged by lemma, part of speech, case, tense, and stem.

Septuagint & Vulgate

Separate original-language layers

The Greek Septuagint (LXX) and the Latin Clementine Vulgate are maintained as their own layers. Translations that descend from them are aligned to those texts — not forced onto the Hebrew Masoretic spine.

Where the word links come from

Word links do not all come from the same place. Each translation uses the most reliable source available to it:

1

Scholar-tagged

Some translations carry scholar-made word-level tagging: the KJV via STEPBible's TAGNT/TAHOT data, the Berean Standard Bible via its published interlinear, and the unfoldingWord Literal Text via its hand-built alignment. These are the most dependable links available for their translation families, though they still need checking where a link is wrong, or too broad to lock a single word to.

2

Publisher-supplied

Some publishers ship word-alignment data alongside their translation text. The NET Bible provides its own alignment data and is treated as authoritative for that translation.

3

Computed, then checked

For translations with no ready-made word links, the links are worked out automatically, then scored against links checked by hand — and by a reviewer for the hardest cases — before a reader sees them.

Feature by Feature

Each feature has a visible reader experience and a less visible data problem underneath it. The work underneath has to be thorough so the page itself stays calm and quick to read.

The big data layer

Reverse Interlinear: Translation-Specific Alignment

How it is built

A traditional interlinear keeps the original Hebrew or Greek word order and contorts the English around it. A reverse interlinear keeps the English exactly as the translator wrote it, and places the matching original-language data beneath each word in natural reading order — per translation, for every supported Bible. Where an authoritative alignment source exists (KJV via STEPBible, the Berean Standard Bible's published interlinear, the unfoldingWord Literal Text), that data is used directly. For other translations, the build step runs a computer-assisted matching engine: it gives each original-language word a set of possible English matches, then scores them using lexical glosses, learned translation probabilities (how likely this translation renders a given Hebrew or Greek word with a specific English word), part-of-speech and lemma matching, and positional evidence. Links a person has checked are shown as verified; the rest stay marked approximate until they pass the same check.

Technical complications

Biblical Hebrew is often Verb-Subject-Object; English is Subject-Verb-Object, so the source word rarely lands directly below its English rendering. The engine models positional displacement — how far a word is likely to drift between source and target — using the same diagonal model that drives statistical machine translation systems. One Hebrew word can become an English phrase and vice versa; translators also supply words English grammar requires that have no separate original-language word (the possessive in a Hebrew construct chain, the copula, the article). The build step records those supplied words as supplied rather than forcing false one-to-one matches. Disambiguation is harder still: 'the LORD God' in Genesis 2 maps to two Hebrew words whose glosses overlap — the system separates them using learned co-occurrence, morphology, and local positional evidence. Transliterated names that vary across translations (Nebuchadnezzar vs Nabuchodonosor) are matched using edit distance on the orthographic form.

Desired result

The operating principle is "zero wrong anchors": a confident wrong link is worse than a visible gap. Computed links are scored against a set of links checked by hand; only links that score well enough are shown as verified, and the rest stay marked approximate. Hard cases — free paraphrase, idiomatic collapses, English idioms where no source word cleanly corresponds — are flagged for review rather than guessed. When a reader taps a word in the English text, they see original-language data the system can stand behind.

Learning by reading

Word Locks and Personalised Bibles

How it is built

Word Locks are keyed to source-language identity, normally a Strong's number plus the matching original-language word. In Composite mode, a reader can choose a preferred rendering for a word, and the app applies it wherever the alignment is suitable. Verse Locks and Word Locks share the same personalisation model. By default, Word Locks still apply inside a verse-locked verse; My OpenScripture → Lock Appearance lets readers choose verse-first wording instead.

Technical complications

This only works if the alignment is compact enough for substitution. If one original-language word accidentally owns a whole English clause, a Word Lock would damage the sentence. The build step therefore uses a substitution test: replacing one direct anchor should leave the surrounding English intelligible. Connector words such as articles and prepositions stay out of the lock path so locking the meaningful word does not quietly swallow its neighbours. Publisher policies also matter, so the write path has to respect translation-specific restrictions rather than relying on the button being hidden.

Desired result

A reader can build vocabulary in context. Instead of studying a word once in a separate lexicon, they can see that word reappear across Scripture with their chosen rendering, while the rest of the verse remains connected to published translation text.

Why translations disagree

Translation Difference Symbols

How it is built

OpenScripture stores precomputed divergence data by verse. Entries are classified by what kind of difference is present: source-text or canon split, theological or interpretive rendering, or translation philosophy. The reader sees circled markers in context, and the drawer explains the difference with the relevant renderings grouped by tradition.

Technical complications

The hard part is judgment. A visible wording difference is not automatically a meaningful disagreement. We have to avoid inflating ordinary style differences into manuscript issues, avoid hiding important textual variants, keep explanations short, and store only phrase-level renderings so copyright boundaries stay respected.

Desired result

Readers get a small signal exactly where it helps: this verse is translated differently, and here is why. The goal is not to push a preferred wording, but to help people notice the scholarly landscape behind familiar English phrases.

Manuscript context without overload

Textual Certainty Signals

How it is built

Textual certainty data is stored sparsely at word or passage level. The build step can draw on documented reference variants, translation editorial brackets, and SBLGNT/MorphGNT signals, then attach scores and reasons to individual words. The reader setting decides how strongly those signals appear.

Technical complications

Textual certainty is adjacent to translation disagreement, but it is not the same thing. A translation can differ because of style even when the source text is stable, or because the underlying manuscript reading is genuinely contested. The app keeps those signals separate. It also has to respect licensing limits around critical apparatus material, storing only what OpenScripture is allowed to store.

Desired result

Stable readings stay quiet. More debated readings can be marked when the reader wants that level of detail. The result is a Bible reader that can surface manuscript uncertainty without turning every chapter into a specialist apparatus.

Independent, evidence-bound, and not yet published

AISE, AIRE, and AIDE Bible editions

How it is built

OpenScripture is developing three fresh editions: the AISE Bible (Study Edition), which shows its sources, the accessible AIRE Bible (Reader Edition), and the literary AIDE Bible (Dynamic Edition). Each edition will translate independently from certified source and research packets rather than rewriting another edition or a copyrighted English Bible. The earlier experimental renderings have been withdrawn. Read the translation philosophy and current plan.

Technical complications

Fluent wording is not enough. Every passage needs an exact source profile, source-backed lexical and textual evidence, independent translation review, every source word accounted for, word links settled by review, notes written to one standard, and a fixed, named version recorded before release. Hebrew, Aramaic, Greek, Latin, Septuagintal, Deuterocanonical, and other source routes cannot silently borrow one another's authority.

Desired result

No AI Bible wording is currently published in the reader. The AI mode is a development placeholder until a complete edition has passed its research, translation, alignment, apparatus, and release gates.

The quiet infrastructure

Translation Ingestion, Notes, and Canons

How it is built

Each new translation has to pass through licensing, metadata, verse text ingestion, publisher notes, copyright notices, reader visibility, search, comparison, and word-data checks. Where a publisher provides notes, introductions, commentary, or cross-references, those sources are normalised so the drawer can show the right material for the verse the reader tapped.

Technical complications

Publishers deliver data in different formats. Word documents, USFM, JSON, public-domain files, study notes, cross-reference lists, and commentary all behave differently. Versification can differ. Canon scope can differ. Formatting such as italics, bold, paragraphing, quotation layout, and note anchors is part of the meaning, so the parser cannot simply flatten everything into plain text.

Desired result

The desired result is a broad, respectful reader across Protestant, Catholic, Orthodox, Ecumenical, Jewish, and Independent traditions, with each translation shown on its own terms and connected to the same study surfaces where the data allows.

How the data is built

The same discipline shows up across translation ingestion, word alignment, divergence explanations, and certainty data. OpenScripture keeps the original source, the automatic draft, and the reader-facing claim separate until the evidence is strong enough.

  1. 1

    Pick the most reliable source of word links for this translation: scholar tagging, publisher-supplied links, or links worked out automatically.

  2. 2

    Ingest the publisher text and notes from the most authoritative available source.

  3. 3

    Normalise books, chapters, verses, notes, formatting, copyright terms, and reader visibility.

  4. 4

    Attach grammar details, Strong's numbers, the original-language words, and the place each one takes in this translation's English.

  5. 5

    Draft the word links and difference explanations automatically, then score them against examples checked by hand.

  6. 6

    Publish the reviewed data so the reader, the verse drawer, the interlinear view, and locks can all use it.

  7. 7

    Keep uncertainty visible: approximate links stay approximate, checked links earn stronger language, and gaps remain labelled rather than hidden.

Research direction — not yet fully shipped

Where the method is heading

Automatically computed word links are being improved around a principle from modern natural language processing research: combine several independent lines of evidence and trust a link when they agree. A link corroborated by lexical glosses, learned translation probabilities, part-of-speech and lemma matching, position, and semantic similarity is more trustworthy than one supported by a single signal.

The evidence the build step is designed to combine includes: lexical glosses, translation-probability models trained on confirmed word links, part-of-speech and lemma matching, agreement across multiple translations sharing the same Strong's spine, local semantic embeddings, statistical aligners, syntactic structure, and orthographic similarity for transliterated name variants. Agreement across independent sources earns higher confidence; disagreement flags a case for human review.

Literal translations align almost word-for-word and move through this system cleanly. Dynamic or paraphrase translations reorder, expand, and supply more — so their alignment is partial by nature, and the system is designed to say so rather than fabricate coverage. Where a layer is not yet finished, the app shows what data exists rather than guessing. See current status for live word-link and AI-generation progress.

A contribution to digital Bible accessibility

The broader movement is about making the depth behind the text easier to reach: original languages, translation philosophy, textual history, commentary, cross-references, and personal study patterns.

OpenScripture contributes by building a reader where those layers can be available without overwhelming the page.

  • Make deep study tools understandable for ordinary readers, not only specialists.
  • Let many translation traditions sit beside each other without flattening their differences.
  • Show where the data stops, so digital convenience does not become false certainty.
  • Use modern software, licensed sources, and reviewable data to make Scripture study more accessible over time.

Sources and Methods

OpenScripture is built on open scholarship. The word-matching work re-implements ideas from the computational linguistics research below; it does not vendor these tools directly. License terms are listed as published by each project — verify at the linked source before relying on them.

Data Sources and Texts

Tagged Greek New Testament and Tagged Ancient Hebrew/Aramaic OT, including KJV word-level alignment and Strong's links.

Westminster Leningrad Codex (the Masoretic Text) with full morphological tagging for every Hebrew and Aramaic word.

Edited by Michael W. Holmes (Society of Biblical Literature & Logos Bible Software). Paired with MorphGNT morphology.

Word-level interlinear alignment published alongside the Berean Bible translation.

Hand-built word-level alignment to the original-language texts, published by unfoldingWord.

Translation text and word-alignment data for the NET Bible.

Syntactic trees and additional linguistic annotation for biblical Hebrew and Koine Greek.

Brenton's English Septuagint

Public domain

English translation of the Greek Old Testament (Septuagint / LXX), used as base for LXX-aligned translations.

Clementine Vulgate

Public domain

Latin Vulgate base text, used as the original-language layer for Vulgate-descended translations such as the Douay-Rheims.

Strong's Exhaustive Concordance (James Strong, 1890)

Public domain

The original lexical numbering system for biblical Hebrew and Greek — still the most widely used cross-translation linking scheme in Bible software.

Methods and Research

The alignment engine adapts ideas from the following computational linguistics literature. These are the intellectual lineage of the method; OpenScripture re-implements the core concepts for the biblical alignment domain rather than vendoring any of these tools.

  1. 1

    Brown, P. F., et al. (1993). The mathematics of statistical machine translation: parameter estimation. Computational Linguistics, 19(2).

    IBM Models 1–5: the foundational statistical word alignment framework underpinning most aligners.

  2. 2

    Vogel, S., Ney, H., & Tillmann, C. (1996). HMM-based word alignment in statistical translation. COLING.

    Introduced the positional (diagonal) displacement model for word order differences between source and target.

  3. 3

    Och, F. J., & Ney, H. (2003). A systematic comparison of various statistical alignment models. Computational Linguistics, 29(1).

    Defined Alignment Error Rate (AER): precision and recall across sure and possible links.

  4. 4

    Melamed, I. D. (2000). Models of translational equivalence among words. Computational Linguistics, 26(2).

    Competitive linking — resolving many-to-one and one-to-many alignment conflicts.

  5. 5

    Moore, R. C. (2004). Improving IBM word alignment Model 1. ACL.

    Practical improvements to translation probability estimation for small or sparse parallel corpora.

  6. 6

    Liang, P., Taskar, B., & Klein, D. (2006). Alignment by agreement. NAACL.

    Training source→target and target→source models to agree — key inspiration for corroborated-evidence alignment.

  7. 7

    Koehn, P., et al. (2007). Moses: Open source toolkit for statistical machine translation. ACL.

    Reference SMT implementation including symmetrisation heuristics (grow-diag-final) used across the field.

  8. 8

    Dyer, C., Chahuneau, V., & Smith, N. A. (2013). A simple, fast, and effective reparameterization of IBM Model 2. NAACL.

    fast_align: efficient alignment with a strong diagonal prior, widely used for new language pairs.

  9. 9
  10. 10
  11. 11

    Dou, Z., & Neubig, G. (2021). Word alignment by fine-tuning embeddings on parallel corpora. EACL.

    awesome-align: fine-tuned contextual embeddings achieving state-of-the-art alignment on standard benchmarks.

  12. 12

    Imani, A., et al. (2021). Graph-based multilingual word alignment. EACL.

    Multi-parallel consensus alignment — leveraging agreement across multiple translation pairs sharing a common source.

  13. 13
  14. 14

    Levenshtein, V. I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8).

    Edit distance — orthographic similarity for transliterated name variants (Nebuchadnezzar / Nabuchodonosor).

  15. 15

    Bird, S., Klein, E., & Loper, E. (2009). Natural Language Processing with Python. O'Reilly.

    NLTK — reference NLP toolkit for tokenization and lemmatization.

  16. 16

How this shows up in the app

The reader stays simple because the work happens before it: texts ingested from their licensed sources, word links checked before they are shown, and gaps left visible instead of filled in.