Abstract

A system is proposed herein for reducing the time between upload of enterprise knowledge and a first useful answer in a retrieval-augmented generation (RAG) environment. The system makes a newly uploaded document rapidly searchable through lightweight text extraction, chunking, metadata extraction, and lexical indexing while semantic enrichment continues in the background. Pending ingestion work is decomposed into document-type-aware tasks and prioritized using document characteristics, live query demand, retrieval-quality signals, expected knowledge utility, processing cost, service-level objectives, and tenant fairness. A global multi-tenant scheduler can continuously reprioritize embedding, optical character recognition (OCR), table extraction, entity extraction, summarization, image understanding, and related tasks so that higher-value knowledge becomes available earlier. The framework supports cold-start operation, query-driven adaptation, and concurrent multi-tenant workloads, and can evolve from deterministic scoring to machine-learning and reinforcement-learning scheduling while preserving the same progressive-ingestion architecture.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS