Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

Foundation-Model-Project

Mitigating Hallucination in Vision-Language Models

This project explores techniques for reducing hallucinations in Vision-Language Models (VLMs). We investigate supervised fine-tuning, Direct Preference Optimization (DPO), and feedback-guided self-revision to improve the factual accuracy and visual grounding of model responses.

Project Overview

Vision-Language Models can generate incorrect information that is not grounded in the provided image. This project investigates approaches inspired by Fact-RLHF and Volcano to reduce these hallucinations.

We experimented with:

Models and Techniques

Supervised Fine-Tuning

BLIP-2 was fine-tuned using the LLaVA-Instruct dataset. We trained on 5,000 visual question-answer pairs using LoRA adapters and float16 precision.

Direct Preference Optimization

We generated multiple responses to image-question pairs using BLIP-2 and used Qwen2-VL as a judge to construct preference pairs.

The resulting preference dataset was used to train BLIP-2 with the DPOTrainer from the Hugging Face TRL library.

Fact-Augmented Self-Revision

We implemented a feedback-guided revision process inspired by Volcano.

The model:

  1. Generates an initial response.
  2. Generates a caption containing factual information about the image.
  3. Uses the factual information to critique and revise its response.
  4. Produces a final answer grounded in the visual content.

We evaluated the approach both with and without captions.

Datasets

LLaVA-Human-Preference LLaVa Human Preference 10k

Used to generate preference data for DPO training.

The dataset contains approximately 10K questions about COCO imagesCOCO train2017.

MMHAL-Bench

MMHAL-Bench evaluates hallucinations in multimodal models using open-ended image-question pairs.

POPE

POPE evaluates object hallucination using yes/no questions about whether objects are present in an image.

Results

BLIP-2 DPO / SFT

Model POPE Accuracy MMHAL Avg. Score Hallucination Rate
BLIP-2 Base 0.73 1.31 0.65
BLIP-2 SFT 0.61 1.35 0.68
BLIP-2 DPO 0.73 1.25 0.66

Self-Revision with Captions

Adding image captions improved performance on both POPE and MMHAL-Bench.

Model Setting POPE Accuracy
BLIP-2 Without Captions 0.31
BLIP-2 With Captions 0.47
LLaVA-OneVision Without Captions 0.31
LLaVA-OneVision With Captions 0.50

On MMHAL-Bench:

Model Setting Avg. Score Hallucination Rate
BLIP-2 Without Captions 1.54 0.58
BLIP-2 With Captions 1.85 0.51
LLaVA-OneVision Without Captions 1.80 0.52
LLaVA-OneVision With Captions 2.30 0.50

Technologies

  • Python
  • PyTorch
  • Hugging Face Transformers
  • Hugging Face TRL
  • BLIP-2
  • LLaVA-OneVision
  • Qwen2-VL
  • LoRA
  • Direct Preference Optimization (DPO)
  • Supervised Fine-Tuning (SFT)

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages