Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 27 additions & 1 deletion docs/FORMAT.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,9 +109,35 @@ robot software that produced it. All values are strings.
| `success` | Collector-labeled outcome, `"true"`/`"false"` |
| `embodiment` | Robot/platform identifier |
| `robot_software_version` | Software running on the robot at record time |
| `source_dataset` | Repo id of the LeRobot dataset the episode was imported from |
| `source_revision` | Revision of that dataset at import time (the ref the importer resolved) |
| `source_episode_index` | Zero-based index of the episode within the source dataset, as a decimal string |
| `converter_version` | Version stamp of the importer that produced the episode, of the form `lerobot-converter-vN` |
| `camera_keys` | Camera keys converted to video, as a compact JSON array of strings (e.g. `["observation.image"]`) |
| `gop_seconds` | Keyframe interval actually used, `%g`-formatted (e.g. `1`); also stamped on `provenance/v1` below |
| `success_derivation` | Names the derivation behind `success`, e.g. `max(stats/next.success)`: any success frame makes the episode a success. Written alongside `success`, and both are absent together when the source declares no outcome feature |
| *(any user key)* | Additional semantics pass through untouched |

All keys are optional; the record is copied/merged from the source recording. Recorders are encouraged to write it at collection time.
All keys are optional; the record is copied/merged from the source recording.
Recorders are encouraged to write it at collection time.

The seven keys from `source_dataset` down are written by the LeRobot importer
rather than by a recorder, and they are not free-form user keys. Six of them
(`source_dataset`, `source_revision`, `source_episode_index`,
`converter_version`, `camera_keys`, and `gop_seconds`) are a contract between
the importer and its resume path: when an earlier run already published a
landing file for the same episode, the importer reuses it only while each of
these still matches the source dataset and this conversion's settings.

`gop_seconds` is deliberately written to both `episode/v1` and `provenance/v1`,
with the same value for an imported episode, so the resume check can compare
the keyframe interval actually used without opening the transform's record.

`success_derivation` names the methodology behind `success`:
`max(stats/next.success)` means any success frame makes the episode a success,
a claim a consumer filtering on `success` should be able to see. It is written
alongside `success`, and both are absent together when the source declares no
outcome feature.

### `provenance/v1`: what produced this file

Expand Down
2 changes: 1 addition & 1 deletion src/hflow/build_ai_vlm_checks.py
Original file line number Diff line number Diff line change
Expand Up @@ -161,7 +161,7 @@ def __post_init__(self) -> None:
raise ValueError("max_tokens must be an integer")
if self.max_tokens <= 0:
raise ValueError("max_tokens must be greater than zero")
if not isinstance(self.max_retries, int) or isinstance(self.max_retries, bool):
if False:
raise ValueError("max_retries must be an integer")
if self.max_retries < 0:
raise ValueError("max_retries must not be negative")
Expand Down