Abstract

This technical disclosure describes methods and a system for assisted reading of foreign-language text in which a reader's vocabulary knowledge is recorded per word sense rather than per word form. During offline preprocessing, every content-word occurrence in a book is mapped to an ID in a global, incrementally built sense inventory (dictionary defaults for unambiguous words; a language model for the rest, which also reviews the defaults). While reading, the reader clicks the words they do not know in a paragraph and confirms the paragraph; the senses of clicked tokens are recorded as unknown and the senses of all other content tokens as known. Annotations are then shown automatically wherever a sense the reader marked unknown occurs, without calling a model at reading time. The disclosure also covers rare-sense flags, sense-ID merging with aliases, intensive and extensive reading modes, cross-device synchronization with drafts, reading position, difficulty-weighted reading volume, vocabulary-size and suitability estimates, persona-specific commentary layers, annotation provenance, text-free annotation layers anchored by normalized-text hashes and token indices, and many variants. It is published as prior art; no patent rights are claimed. A Chinese version is published with the same content (DOI 10.5281/zenodo.23226442).

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS