Abstract

A system and method for detecting entity mentions in multilingual text are disclosed. A shared multilingual transformer encoder produces a pooled representation from a tokenized representation of an item of text. A named-entity-recognition (NER) feature extraction module produces an NER feature vector from the item of text and a configured entity list, the NER feature vector including exact-match, fuzzy-similarity, entity-overlap, organization-entity, position-score, and mention-frequency components. A late-fusion concatenation module concatenates the pooled representation with the NER feature vector after and external to encoder layers of the shared multilingual transformer encoder to form a joint representation. A classification head produces an entity-mention probability from the joint representation. A multi-guard post-filter chain compares the entity-mention probability to a configurable threshold and applies a plurality of sequential guards, with at least one guard short-circuited in response to a script family of the item of text being a non-Latin script family.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS