Abstract
Conversational artificial intelligence systems for enterprise databases may face a trade-off between the potentially high cost and latency of large models and the potentially lower accuracy of smaller models when processing complex, domain-specific schemas. A system is described for natural language query decomposition that can use a constrained-parameter small language model (SLM). The system can process large data schemas by decomposing them into smaller, parallel batches. Specialized prompt engineering, such as an inverted prompt structure for caching, and a structured reasoning template can guide the SLM. A sequential ambiguity resolver may then consolidate outputs from the parallel workers into a final query plan. This approach can allow a smaller model to generate structured subqueries from natural language requests with potentially reduced latency, which may facilitate interaction with large data catalogs.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Badjatiya, Pinkesh, "Natural Language Query Decomposition Using a Small Language Model with Parallel Schema Processing", Technical Disclosure Commons, (September 07, 2026)
https://www.tdcommons.org/dpubs_series/11617