Abstract

Methods, systems, and computer program products are provided for fairness-aware, test-time prompt tuning of a vision-language machine learning model. An example method includes receiving text prompts and images, inputting the text prompts and the images to a vision-language machine learning model, generating text embeddings based on the text prompts and image embeddings based on the images, calculating target similarity scores between the target text embeddings and the image embeddings and spurious similarity scores between the spurious text embeddings and the image embeddings, calculating a measure of target entropy loss based on the target similarity scores and a measure of spurious entropy loss based on the spurious similarity scores, and updating the vision-language machine learning model to minimize the measure of target entropy loss and to maximize the measure of spurious entropy loss.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS