Abstract

We disclose a method for automatically constructing and validating Natural Language Inference (NLI) training pairs from unlabeled regulatory corpora in the accounting and tax domain (e.g., accounting standards, standard-setter Q&A, tax rulings, tax-tribunal decisions, audit opinions, regulatory-filing notes). The core is a provenance-conditional arbitration rule that assigns label adoption authority according to the source type of each pair, combined with asymmetric agreement thresholds, structure-relation label derivation, domain antonym contradiction synthesis, and numeric-equivalence pair generation. Publishing this defensively secures freedom to operate and blocks third-party patents on the technique.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS