Abstract
Virtual try-on (VTO) systems using generative models can be a challenge. To address this, an agentic framework can evaluate and correct errors in generated VTO images. The framework may operate as an iterative feedback loop, where a multimodal evaluation module can analyze an image for pose and semantic discrepancies. Based on the analysis, a routing system can direct the task to a targeted corrective pathway, such as a full structural regeneration or a localized semantic repair, guided by a dynamic prompt engine. This iterative process of evaluation and targeted correction may continue until a set of criteria is met, providing a scalable method for producing high-fidelity VTO outputs with a potential reduction in the need for manual review.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Garware, Bhushan; Barapatre, Darshan; and Athalye, Anibha, "The Agentic Framework for Iterative Pose and Semantic Correction in Virtual Try-on", Technical Disclosure Commons, (September 09, 2026)
https://www.tdcommons.org/dpubs_series/11681