Evaluate Β· Replay Β· Inspect Data β All-in-one toolkit for real-world VLA deployment debugging
π Quick Start Β· πΈ Screenshots Β· π― Features Β· π§ Installation
Deploying VLA models to real robots is hard. You face:
- π΅οΈ Black-box inference β Can't see what the model "sees" or why it fails
- β±οΈ Hidden latencies β Transport delays, inference bottlenecks, control loop timing issues
- π Fragmented logging β Every framework logs differently, making cross-model comparison painful
- π Tedious debugging β Replaying failures requires manual log parsing and visualization
VLA-Lab solves this. A unified logging format + interactive visualization dashboard covering the debugging loop from open-loop evaluation to real-robot replay and dataset inspection.
|
Start offline: compare predicted actions against ground truth before touching the robot. Inspect MSE / MAE summaries, temporal alignment, error heatmaps, and 3D trajectory overlays. Replay deployment logs step-by-step: multi-camera observations, state/action curves, action chunks, latency breakdowns, attention overlays, and Rerun-compatible recordings. |
Inspect LeRobot / GR00T / Zarr datasets frame-by-frame with video playback, state/action/gripper curves, sampled frames, and workspace distributions. When real-world success stays low, compare deployment observations against the dataset to catch out-of-distribution lighting, camera framing, object placement, or task-state mismatches. |
| Framework | Status |
|---|---|
| Diffusion Policy | β Supported |
| Isaac-GR00T | β Supported |
| Pi 0.5 | β Supported |
| DreamZero | β Supported |
| VITA | β Supported |
VLA-Lab uses a unified logging protocol β adapting a new framework takes only a few lines of glue code.
- Run open-loop evaluation first. If predicted actions do not match ground truth offline, fix the checkpoint, action normalization, horizon, or model inputs before running the robot.
- Then inspect real-robot Runs. Check whether camera observations, model paths, prompts, states, action chunks, timing, and attention maps look sane during deployment.
- Finally inspect the dataset for OOD cases. If open-loop eval and real-robot wiring both look healthy but success remains low, compare the run observations against training episodes. The failure may come from lighting, camera placement, object configuration, or states that were absent from the dataset.
pip install vlalabInstall with full dependencies (including Zarr dataset support):
pip install vlalab[full]Or install from source:
git clone https://github.com/ky-ji/VLA-Lab.git
cd VLA-Lab
pip install -e .import vlalab
# Initialize a run
run = vlalab.init(project="pick_and_place", config={"model": "diffusion_policy"})
# Log during inference
vlalab.log({"state": obs["state"], "action": action, "images": {"front": obs["image"]}})import vlalab
# Initialize with detailed config
run = vlalab.init(
project="pick_and_place",
config={
"model": "diffusion_policy",
"action_horizon": 8,
"inference_freq": 10,
},
)
# Access config anywhere
print(f"Action horizon: {run.config.action_horizon}")
# Inference loop
for step in range(100):
obs = get_observation()
t_start = time.time()
action = model.predict(obs)
latency = (time.time() - t_start) * 1000
# Log everything in one call
vlalab.log({
"state": obs["state"],
"action": action,
"images": {"front": obs["front_cam"], "wrist": obs["wrist_cam"]},
"inference_latency_ms": latency,
})
robot.execute(action)
# Auto-finishes on exit, or call manually
vlalab.finish()# Default web dashboard (FastAPI + Next.js)
vlalab view# 1. Install Python package
pip install -e .
# 2. Install web dependencies once
cd web
npm install
cd ..
# 3. Start FastAPI + Next.js together
vlalab viewUseful variants:
# Start only the API backend
vlalab serve --no-frontend
# Change the frontend port
vlalab view --port 3100
# Override the run directory used by the API
vlalab view --run-dir /path/to/vlalab_runsThe web UI covers the main workflows:
- deploy dashboard
- overview / run list / run detail
- dataset viewer
- open-loop eval viewer
Run β A single deployment session (one experiment, one episode, one evaluation)
Step β A single inference timestep with observations, actions, and timing
Artifacts β Images, point clouds, and other media saved alongside logs
vlalab.init() β Initialize a run
run = vlalab.init(
project: str = "default", # Project name (creates subdirectory)
name: str = None, # Run name (auto-generated if None)
config: dict = None, # Config accessible via run.config.key
dir: str = "./vlalab_runs", # Base directory (or $VLALAB_DIR)
tags: list = None, # Optional tags
notes: str = None, # Optional notes
)vlalab.log() β Log a step
vlalab.log({
# Robot state
"state": [...], # Full state vector
"pose": [x, y, z, qx, qy, qz, qw], # Position + quaternion
"gripper": 0.5, # Gripper opening (0-1)
# Actions
"action": [...], # Single action or action chunk
# Images (multi-camera support)
"images": {
"front": np.ndarray, # HWC numpy array
"wrist": np.ndarray,
},
# Timing (any *_ms field auto-captured)
"inference_latency_ms": 32.1,
"transport_latency_ms": 5.2,
"custom_metric_ms": 10.0,
})RunLogger β Advanced API
For fine-grained control over logging:
from vlalab import RunLogger
logger = RunLogger(
run_dir="runs/experiment_001",
model_name="diffusion_policy",
model_path="/path/to/checkpoint.pt",
task_name="pick_and_place",
robot_name="franka",
cameras=[
{"name": "front", "resolution": [640, 480]},
{"name": "wrist", "resolution": [320, 240]},
],
inference_freq=10.0,
)
logger.log_step(
step_idx=0,
state=[0.5, 0.2, 0.3, 0, 0, 0, 1, 1.0],
action=[[0.51, 0.21, 0.31, 0, 0, 0, 1, 1.0]],
images={"front": image_rgb},
timing={
"client_send": t1,
"server_recv": t2,
"infer_start": t3,
"infer_end": t4,
},
)
logger.close()# Launch the default FastAPI + Next.js dashboard
vlalab view [--port 3000] [--api-port 8000]
# Launch FastAPI backend and optional Next.js frontend
vlalab serve [--api-port 8000] [--web-port 3000] [--frontend/--no-frontend]
# Convert legacy logs (auto-detects format)
vlalab convert /path/to/old_log.json -o /path/to/output
# Inspect a run
vlalab info /path/to/run_dirVLA-Lab now treats attention extraction as a model-provided backend with a stable interface.
The built-in default still auto-discovers the existing Isaac-GR00T backend, but new models should register their own backend explicitly:
export VLALAB_ATTENTION_BACKEND=/abs/path/to/backend.py
export VLALAB_ATTENTION_PYTHON=/abs/path/to/python # optionalInterface contract and migration checklist:
vlalab_runs/
βββ pick_and_place/ # Project
βββ run_20240115_103000/ # Run
βββ meta.json # Metadata (model, task, robot, cameras)
βββ steps.jsonl # Step records (one JSON per line)
βββ artifacts/
βββ images/ # Saved images
βββ step_000000_front.jpg
βββ step_000000_wrist.jpg
βββ ...
- Core logging API & unified run format
- Streamlit visualization suite (5 pages)
- Diffusion Policy adapter
- Isaac-GR00T adapter
- Pi 0.5 adapter
- DreamZero adapter
- VITA adapter
- Open-loop evaluation pipeline
- Cloud sync & team collaboration
- Real-time streaming dashboard
- Automatic failure detection
- Integration with robot simulators
We welcome contributions!
git clone https://github.com/ky-ji/VLA-Lab.git
cd VLA-Lab
pip install -e ".[dev]"MIT License β see LICENSE for details.
β Star us on GitHub if VLA-Lab helps your research!
Built with β€οΈ for the robotics community


