Abstract

Classifying custom, domain-specific object subcategories can present challenges, as retraining classical models may be resource-intensive and using generative models alone might produce unreliable results. A technique is described for a hybrid computational pipeline to address such challenges. For example, a classical object detection model can identify general object categories within an image. These detections may then be filtered according to a user-defined mapping that links general categories to specific, custom subclasses of interest. For remaining candidate objects, a large vision model may generate a textual description, and a large language model can perform a detailed classification against a user's list of custom subclasses, optionally using exemplar images for context. This method can provide an adaptable system for classifying custom object sets with potentially reduced data requirements by combining classical detection with the reasoning capabilities of generative models.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS