This project explores techniques for reducing hallucinations in Vision-Language Models (VLMs). We investigate supervised fine-tuning, Direct Preference Optimization (DPO), and feedback-guided self-revision to improve the factual accuracy and visual grounding of model responses.
Vision-Language Models can generate incorrect information that is not grounded in the provided image. This project investigates approaches inspired by Fact-RLHF and Volcano to reduce these hallucinations.
We experimented with:
- Supervised Fine-Tuning (SFT)
- Direct Preference Optimization (DPO)
- Fact-augmented self-guided revision
- BLIP-2 BLIP VQA (400M)
- LLaVA-OneVision( https://huggingface.co/llava-hf/llava-onevision-qwen2-0.5b-ov-hf)(0.5B)
BLIP-2 was fine-tuned using the LLaVA-Instruct dataset. We trained on 5,000 visual question-answer pairs using LoRA adapters and float16 precision.
We generated multiple responses to image-question pairs using BLIP-2 and used Qwen2-VL as a judge to construct preference pairs.
The resulting preference dataset was used to train BLIP-2 with the DPOTrainer from the Hugging Face TRL library.
We implemented a feedback-guided revision process inspired by Volcano.
The model:
- Generates an initial response.
- Generates a caption containing factual information about the image.
- Uses the factual information to critique and revise its response.
- Produces a final answer grounded in the visual content.
We evaluated the approach both with and without captions.
LLaVA-Human-Preference LLaVa Human Preference 10k
Used to generate preference data for DPO training.
The dataset contains approximately 10K questions about COCO imagesCOCO train2017.
MMHAL-Bench evaluates hallucinations in multimodal models using open-ended image-question pairs.
POPE evaluates object hallucination using yes/no questions about whether objects are present in an image.
| Model | POPE Accuracy | MMHAL Avg. Score | Hallucination Rate |
|---|---|---|---|
| BLIP-2 Base | 0.73 | 1.31 | 0.65 |
| BLIP-2 SFT | 0.61 | 1.35 | 0.68 |
| BLIP-2 DPO | 0.73 | 1.25 | 0.66 |
Adding image captions improved performance on both POPE and MMHAL-Bench.
| Model | Setting | POPE Accuracy |
|---|---|---|
| BLIP-2 | Without Captions | 0.31 |
| BLIP-2 | With Captions | 0.47 |
| LLaVA-OneVision | Without Captions | 0.31 |
| LLaVA-OneVision | With Captions | 0.50 |
On MMHAL-Bench:
| Model | Setting | Avg. Score | Hallucination Rate |
|---|---|---|---|
| BLIP-2 | Without Captions | 1.54 | 0.58 |
| BLIP-2 | With Captions | 1.85 | 0.51 |
| LLaVA-OneVision | Without Captions | 1.80 | 0.52 |
| LLaVA-OneVision | With Captions | 2.30 | 0.50 |
- Python
- PyTorch
- Hugging Face Transformers
- Hugging Face TRL
- BLIP-2
- LLaVA-OneVision
- Qwen2-VL
- LoRA
- Direct Preference Optimization (DPO)
- Supervised Fine-Tuning (SFT)