Inventor(s)

Abstract

This paper presents a vendor-agnostic, search-augmented large language model (LLM) pipeline that validates, enriches, and normalizes company records using grounded information from public web sources. The pipeline addresses enterprise requirements for official names, addresses, tax identifiers, revenue, industry classifications, and related compliance metadata where manual validation, database lookups, and rule-based integration pipelines struggle with incomplete, inconsistent, multilingual, and region-specific inputs. The architecture uses a central control framework to orchestrate native-search and tool-calling web-enabled LLM providers under production constraints, including context limits, tokens-per-minute (TPM), requests-per-minute (RPM), maximum batch size, concurrency, incomplete outputs, safety filtering, and provider-specific response formats. It combines batch-admission control, adaptive retry, fallback routing, and structured output validation to improve throughput and robustness. Experiments on 20 company records show that both evaluated architectures achieve high enrichment completeness, while native-search models provide lower latency and stronger resilience under increased concurrency. The findings show that production-scale search-augmented validation depends equally on model capability and operational stability.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution-Share Alike 4.0 License.

Share

COinS