DAVID NKWO
All work
02Deployed · Robot Perception

Adaptive Visual Grounding for Robot Vision

A deployed robot-vision system that converts natural-language instructions into localized visual targets across still images, recorded video and browser-camera inputs.

  • YOLO
  • OWL-ViT
  • GPT
  • Visual Grounding
Person / lobby grounding result
Grounding output — person in a lobby scene.

01

What it does

A deployed robot-vision system that converts natural-language instructions into localized visual targets across still images, recorded video and browser-camera inputs.

The system parses objects, attributes, quantities and spatial relationships, then routes requests across YOLO, OWL-ViT and GPT-guided OWL-ViT while producing structured bounding boxes, confidence, timing information and annotated outputs.

To determine which approach worked best, I compared Grounding DINO, OWL-ViT, GPT Vision and a custom configuration on the same 105-image evaluation set. Based on those results, I designed and implemented a custom GPT Vision-guided OWL-ViT pipeline to improve grounding accuracy. The pipeline combines OWL-ViT open-vocabulary proposals with GPT Vision-based instruction interpretation and candidate selection, followed by local refinement, bounded box checks and tracking/re-acquisition for video and live-camera use.

02

Demonstrations

Table / flower relationship grounding
Relationship grounding — table / flower
Door / sign object and attribute grounding
Object & attribute grounding — door / sign
Second-closest package spatial / ordinal grounding
Spatial / ordinal grounding — second-closest package
Grey-door grounding and sustained tracking demonstration
Robot target grounding and tracking
Correct-package-on-chair grounding
Hidden-package spatial grounding

03

System overview

Request routing

  1. Natural-language instruction
  2. Parse objects · attributes · relations
  3. Route: YOLO / OWL-ViT / GPT-guided OWL-ViT
  4. Boxes · confidence · timing

04

Results

105
Evaluation images
0.473
Mean IoU
78.1%
Success @ IoU ≥ 0.25
57.1%
Success @ IoU ≥ 0.50

Results for the strongest configuration: the custom GPT Vision-guided OWL-ViT pipeline. The custom pipeline achieved the strongest mean IoU at 0.473, compared with OWL-ViT at 0.344, GPT Vision at 0.300 and Grounding DINO at 0.178 on the shared 105-image evaluation set.

05

Model comparison & custom pipeline

  1. 01Structure the natural-language task into targets, attributes, relations and anchors.
  2. 02Generate and rank open-vocabulary proposals with OWL-ViT.
  3. 03Use GPT Vision for instruction-aware candidate selection and ambiguity resolution.
  4. 04Refine locally and apply bounded box checks.
  5. 05Track and stabilize between model calls, recover from short losses and re-acquire when necessary.

06

Technology

  • Python
  • YOLO
  • OWL-ViT
  • GPT
  • Browser camera

Next project

03 — ROS 2 Hardware Integration & Robotic Arm Bring-Up