DAVID NKWO
All work
05Multimodal AI · Data Systems

Vision-Language Model Evaluation Pipeline

Built a reproducible evaluation pipeline for benchmarking vision-language models across standardized multimodal question-answering datasets.

  • PySpark
  • Parquet
  • BLIP-2
  • InstructBLIP
  • LLaVA
Evaluation heatmap
Evaluation heatmap from actual results.

01

What it does

Built a reproducible evaluation pipeline for benchmarking vision-language models across standardized multimodal question-answering datasets.

The system uses PySpark and Parquet to normalize benchmark records, provides model interfaces for BLIP-2, InstructBLIP and LLaVA-style workflows, performs answer-ranking evaluation and produces comparative results and heatmaps.

02

Demonstrations

Evaluation pipeline diagram
Evaluation pipeline
Sample evaluation record
Actual sample evaluation record
Model comparison table
Model comparison from supplied results

03

Results

10,301
ScienceQA records
1,000
VMCBench records
11,301
Evaluation items per model
49.11%
InstructBLIP weighted accuracy
45.16%
BLIP-2 weighted accuracy

04

Technology

  • PySpark
  • Parquet
  • BLIP-2
  • InstructBLIP
  • LLaVA
  • Python

Next project

01 — Physical AI Autonomous Delivery Robot