Roboto to LeRobot Contract#
The roboto-to-lerobot-v2_1 and roboto-to-lerobot-v3_0 actions are driven
by a contract YAML stored in the invocation dataset — the dataset the
action is invoked on. The contract tells the action which topics to read from
each recording, how to extract values from
each message, how to align streams onto a common timeline, and what transforms
to apply along the way. Every event in the input
collection becomes one episode in the output
LeRobot dataset, and the contract is applied identically to every episode.
This page is the full schema reference. For a step-by-step walkthrough of authoring a contract and running the action, see Convert to LeRobot.
Full example#
The contract below exercises the full schema: a camera stream, numeric observations that concatenate onto a shared key, a role-bound spec, per-stream alignment and transforms, an action stream, and a task channel. Each piece is unpacked in the sections that follow. (For a minimal starting point, see the user guide instead.)
name: pick_place
version: 1
fps: 20
robot_type: dual_arm
action_lead_steps: 1 # read actions one frame ahead of observations
observations:
# Camera -> a video feature; the image: block routes it to the video pipeline.
- key: observation.images.exo
topic: /camera/exo/image_raw/compressed
type: sensor_msgs/msg/CompressedImage
image:
resize: [480, 640] # [height, width]
align: {method: nearest, tolerance_ms: 100}
# Left arm -> the first three columns of observation.state.
- key: observation.state
topic: /left_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [left_1, left_2, left_3]
align: {method: hold, tolerance_ms: 100}
transforms:
- type: butterworth_lowpass_causal # causal: reproducible at inference
cutoff_hz: 5
order: 2
fs_hz: 50 # freeze the design rate into the contract
# Right arm -> concatenated onto the same feature (shared key).
- key: observation.state
topic: /right_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [right_1, right_2, right_3]
align: {method: hold, tolerance_ms: 100}
# Base velocity, bound by role instead of topic.
# No align: -> hold with the auto-bounded tolerance.
- key: observation.base_velocity
role: base_twist
type: geometry_msgs/msg/Twist
selector:
names: [linear.x, angular.z]
lerobot_names: [base_vx, base_wz]
actions:
- key: action
topic: /left_arm/joint_commands
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [left_cmd_1, left_cmd_2, left_cmd_3]
align: {method: linear, tolerance_ms: null} # null = no tolerance bound
transforms:
- type: resample_uniform
stage: pre # pre only; must be set explicitly
- type: butterworth_lowpass # zero-phase; offline-only
stage: pre
cutoff_hz: 30
order: 2
- type: finite_difference # stage defaults to "post"
tasks:
- key: prompt # optional; defaults to the topic name
topic: /task/prompt
type: std_msgs/msg/String
Top-level fields#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
string |
|
Logical contract name. Recorded with each episode for traceability. |
|
int |
|
Contract version. Bump it when you change the schema in a way that breaks downstream consumers. |
|
float |
|
Target sampling frequency, in Hz. A reference timeline of one frame per |
|
string |
|
Free-form label written into the LeRobot dataset metadata. |
|
int |
|
Shift action sampling |
|
list |
|
Observation streams. Entries with an |
|
list |
|
Action streams. Become numeric action features. See Action specs. |
|
list |
|
Optional task channels used to derive each episode’s task label. See Task specs. |
Observation specs#
Each entry in observations: describes one stream. There are two flavors,
distinguished by whether an image: block is present:
Numeric observations (no
image:block) become a single LeRobot feature perkey. Multiple entries that share akeyare concatenated along the feature axis — see Selectors and lerobot_names.Video / image observations (with an
image:block) become adtype: videofeature. See Video and image specs.
Field |
Required |
Description |
|---|---|---|
|
yes |
LeRobot feature key. Conventionally |
|
yes |
ROS topic name to read the stream from. Not required if the spec binds by
|
|
yes |
Message type string used to select the decoder. See Supported message types. |
|
no |
|
|
no |
Image / video options, chiefly |
|
no |
|
|
no |
A list of transforms applied to the stream. See Transforms. |
Video and image specs#
An entry under observations: becomes a video feature when it carries an
image: block. The output dtype is always video (HWC uint8 RGB at the
configured resize).
observations:
- key: observation.images.exo
topic: /camera/exo/image_raw/compressed
type: sensor_msgs/msg/CompressedImage
image:
resize: [480, 640] # [height, width]
align: {method: nearest, tolerance_ms: 100}
The image: block currently accepts a single field:
Field |
Required |
Description |
|---|---|---|
|
yes |
|
A camera that publishes a compressed video stream rather than per-message stills
is declared the same way, with its own type: — see
Compressed video.
Restrictions on image streams:
align.method: linearis rejected — useholdornearest.Depth encodings are not yet supported. Supported raw
sensor_msgs/msg/Imageencodings arergb8,bgr8,mono8,rgba8,bgra8, and8UC1. Animage.depth:block is currently rejected at contract load; depth support is planned for a future release.
Action specs#
Action specs follow the same shape as numeric observations — the same fields,
with no image: block.
actions:
- key: action
topic: /robot/joint_commands
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2]
lerobot_names: [arm_1, arm_2]
align: {method: hold, tolerance_ms: 100}
As with observations, multiple entries that share a key are concatenated
along the feature axis. Set topic: and type: directly on the action
spec, as above.
Alignment on action streams:
All four alignment methods apply —
hold(the default),nearest,linear, andnone. The inference caveat that limits observation streams does not apply to actions; see the note under Alignment for why every method is safe here.To shift action sampling relative to observations — for example, to compensate for control delay — set
action_lead_stepsat the top level rather than per spec. See Top-level fields.
Task specs#
tasks: is optional and lets the contract derive a per-episode LeRobot task
label from a message stream — for example, a language prompt published on a
topic.
tasks:
- key: prompt # optional; defaults to the topic name
topic: /task/prompt
type: std_msgs/msg/String
A task spec sets topic: and type: — use std_msgs/msg/String; other
message types fail when the stream is decoded — plus an optional key:,
which defaults to the topic name.
For each episode, the task label is resolved in the following order:
The event’s
taskmetadata field, when present and non-empty. This is an explicit per-episode annotation and always wins; thetasks:block is not consulted.Otherwise, the payload of the first
std_msgs/msg/Stringmessage that falls within the episode’s own time window, taken from the first declared task spec. If the contract declares more than one task spec, only the first is used and the others are ignored (a warning names them).Otherwise, the literal string
"default"— when there is notaskmetadata, notasks:block, and no in-window message.
Selectors and lerobot_names#
selector: controls which fields the decoder pulls out of each message:
selector:
names: [joint_1, joint_2, joint_3] # what to extract
lerobot_names: [arm_1, arm_2, arm_3] # how to label them (optional)
Two rules to keep in mind:
The meaning of
namesdepends on the message type. ForJointStatethey are joint names (with an optionalposition.<joint>/velocity.<joint>/effort.<joint>prefix; a bare name defaults toposition). ForImu,Odometry, andTwistthey are dotted paths into the message (for example,linear.x). See Supported message types for the full list.The
lerobot_namesfield is optional but strongly recommended. Without it, a feature is labeled<topic>/<selector_name>, which is verbose and ties feature names to your ROS topology. With it, you get clean LeRobot-facing names.lerobot_namesmust have the same length asnames, and everylerobot_namemust be unique across the entire contract.
When multiple specs share a key (for example, two observation.state
entries from two arms), all of their lerobot_names are concatenated in
order to name the combined feature.
observations:
# Left arm -> the first three columns of observation.state
- key: observation.state
topic: /left_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [left_1, left_2, left_3]
# Right arm -> concatenated onto the same feature
- key: observation.state
topic: /right_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [right_1, right_2, right_3]
The combined observation.state feature then has six columns, labeled
left_1, left_2, left_3, right_1, right_2, right_3 in that order.
Alignment#
Each stream is merged onto the reference timeline (one frame per
1/fps seconds) using its align: block:
align: {method: hold, tolerance_ms: 100}
Method |
Behavior |
|---|---|
|
Last-observation-carried-forward (backward as-of join). Good for slow state signals. Default. |
|
Pick the closest sample in time, in either direction. Good for cameras and high-rate signals. |
|
Linearly interpolate between the two bracketing samples. Numeric streams only — not allowed for images. |
|
Exact timestamp matches only; everything else becomes |
Note
Reproducing alignment at inference. If you plan to deploy the trained
policy for online, streaming inference, only hold and nearest can be
reproduced sample-by-sample on a live input; linear and none need the
complete recording and are therefore offline-only — the same distinction as
non-causal vs causal Transforms. Use hold or nearest for any
observation stream whose processing must match between training and
deployment. Action streams are exempt — a deployed policy emits actions
rather than aligning them — so all four methods are always fine there.
tolerance_ms is the maximum gap, in milliseconds, between the reference
frame timestamp and the matched sample before the result becomes NaN. Set
it to roughly one period of your slowest signal — for example, 100 for a
10 Hz signal.
Tolerance_ms |
Behavior |
|---|---|
omitted (or |
Auto-bounded default of |
|
Unlimited — always carry forward, or always pick the nearest sample, however far away it is. |
|
Rejected. In older contracts |
a positive number |
That bound, in milliseconds, unchanged. |
Note
The legacy field names strategy and tol_ms are no longer accepted.
Use method and tolerance_ms.
Transforms#
Transforms are applied per stream, in the order they appear, before the stream becomes part of the LeRobot dataset:
actions:
- key: action
topic: /robot/joint_commands
type: sensor_msgs/msg/JointState
selector: {names: [j1, j2], lerobot_names: [arm_1, arm_2]}
transforms:
- type: resample_uniform
stage: pre
- type: butterworth_lowpass
stage: pre
cutoff_hz: 30
order: 2
- type: finite_difference # stage defaults to "post"
Stages#
A transform’s stage says whether it runs before or after alignment — the
step that merges each raw stream onto the shared reference timeline of one
frame per 1/fps seconds (see Alignment). A pre transform therefore
sees the raw, irregularly timed samples; a post transform sees the regular,
fps-rate timeline.
Stage |
When it runs |
Notes |
|---|---|---|
|
Before alignment, on the raw stream. |
Operates on the raw sample timestamps. Use it for resampling and pre-alignment filtering. |
|
After alignment, on the reference timeline. |
Operates at |
Built-in transforms#
Type |
Stage |
Causal |
Params |
Description |
|---|---|---|---|---|
|
|
no |
— |
Resamples an irregularly sampled signal onto a uniform grid (the median inter-sample interval) using nearest-neighbor. Always set |
|
any |
no |
|
Zero-phase Butterworth low-pass filter. No phase distortion, but each output sample depends on both past and future samples. |
|
any |
yes |
|
Causal (forward-only) Butterworth low-pass filter. Each output sample depends only on past samples, at the cost of some phase lag. |
|
any |
no |
— |
Frame-to-frame difference ( |
Causal vs non-causal transforms. A causal transform computes each output
sample from the current and past samples only, so the identical operation can
also run online — sample-by-sample at inference time — and reproduce exactly what
training saw. A non-causal transform additionally looks at future samples, so
it can only ever run offline over a complete recording. The Causal column
above marks which is which. When a stream feeds a deployed policy and the same
processing must run at both training and inference time, choose a causal
transform — for example butterworth_lowpass_causal, the causal counterpart of
butterworth_lowpass. When only offline processing matters — for example,
cleaning an observation that plays no role in a deployed policy — a non-causal
transform is fine, and because a zero-phase filter such as butterworth_lowpass
introduces no phase lag, it is often preferable.
Cutoff bounds. cutoff_hz must be strictly between 0 and the Nyquist
frequency (half the sample rate). For a stage: pre transform the sample
rate is the measured raw topic rate; for stage: post it is fps.
Sample-rate override (causal filter only). fs_hz declares the sample rate the
filter is designed for, overriding the rate the transform would otherwise use.
For a stage: pre causal filter, declaring fs_hz is what makes the
filter reproducible during deployment: it freezes the design rate into the
contract so offline and online both build identical filter coefficients. If the
declared fs_hz deviates from the stream’s actual measured rate by more than
10%, a warning names both numbers.
Ordering. When you chain a stage: pre transform that changes the time
axis (resample_uniform) with later filters, place resample_uniform first
and any low-pass filter after it. The sample rate passed to the later
transforms is recomputed automatically from the resampled timestamps.
Note
The built-in transform library will be expanded in future releases.
Role-based binding#
Instead of pinning a spec to a specific topic:, you can bind it by role —
a logical label you attach to the source file. This helps when the same logical
stream lives under different topic names across recordings: you bind the spec to
a stable role once, and each recording maps that role to whatever its actual
topic happens to be.
To use it, set role: on the spec in place of topic:, then tag each source
file with a matching role by setting a role field in the file’s
Roboto metadata to the same value. For example, a spec
with role: arm_joints binds to the file whose metadata has
role: arm_joints. You can set this metadata manually — in the web UI, CLI, or
SDK — so no companion action is required, though one may be provided to populate
roles automatically.
observations:
- key: observation.state
role: arm_joints # instead of topic:
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [arm_1, arm_2, arm_3]
Rules and failure modes:
A spec must set exactly one of
topic:orrole:. Setting both, or neither, is rejected at contract load.Image and video streams cannot be role-bound — bind them by
topic:.Within a dataset, exactly one file may carry a given role. Zero matching files, or more than one, stops the run with an error naming the role.
The matched file must have exactly one topic whose fields cover the spec’s
selector.names; zero or multiple covering topics is an error.
Supported message types#
type: selects the decoder. The selector semantics depend on the type.
Image and video#
|
Required |
|---|---|
|
|
|
|
|
|
|
|
Supported sensor_msgs/msg/Image encodings: rgb8, bgr8, mono8,
rgba8, bgra8, and 8UC1.
Compressed video#
Topics that Roboto ingestion tags compressedVideo store one encoded video
access unit per message rather than a standalone still image, so a frame in the
middle of a group of pictures (GOP) can only be decoded together with the frames
back to its keyframe. The action handles that: it decodes each episode’s whole
time range in one pass, walking back up to 10 seconds to find the keyframe that
anchors the range.
Declare the topic’s real schema name. Three spellings are accepted —
foxglove_msgs/msg/CompressedVideo, foxglove_msgs/CompressedVideo, and
foxglove.CompressedVideo. Declaring a compressed-video topic as
sensor_msgs/msg/CompressedImage is rejected with an explanatory error rather
than silently producing broken frames.
Apart from the type:, the spec is an ordinary video spec — an image:
block carrying resize, plus an align: block:
observations:
- key: observation.images.external
topic: /camera/exo/video
type: foxglove_msgs/msg/CompressedVideo
image:
resize: [480, 640] # [height, width]
align: {method: nearest, tolerance_ms: 150}
- key: observation.images.gripper
topic: /camera/hand/video
type: foxglove_msgs/msg/CompressedVideo
image:
resize: [480, 640]
align: {method: nearest, tolerance_ms: 150}
Notes and limits:
The codec is read from each message’s
formatfield. A stream in a codec other than H.264, H.265, VP9, or AV1 fails the conversion rather than dropping the camera.resizeis applied as the frames are decoded, which keeps an episode’s frames in memory at the target resolution rather than at the source resolution.Frames whose keyframe is unreachable are skipped with a warning. A range where nothing decodes fails the conversion.
Numeric streams#
|
What |
|---|---|
|
Joint name. Optional |
|
Joint name. Each trajectory point becomes its own sample at |
|
DOF name. Optional |
|
Dotted path into the message, for example |
|
Dotted path. Without |
|
Dotted path, for example |
|
not used; the full |
|
not used; full |
|
not used; full |
|
not used; a single-element array containing |
|
not used; emitted as a string. Intended for |
|
Series index keys, each read as a float. |
|
ANYmal joint name into |
|
ANYmal joint name, mapped positionally — the message carries no per-actuator name — onto the fixed order |
Validation rules and common errors#
The contract loader checks a few things up front. When it rejects a contract, the error message names the offending spec.
Binding: every observation and action spec must set exactly one of
topic:orrole:(see Role-based binding). Image and video streams must usetopic:.Image alignment: image streams cannot use
align.method: linear.Compressed video: a topic that stores compressed video must be declared with a compressed-video
type:. Declaring it assensor_msgs/msg/CompressedImage— or as a compressed-video spelling that has no decoder — stops the run, naming the offending spec’skey(see Compressed video).Names:
lerobot_namesmust matchnamesin length, and everylerobot_namein the contract must be globally unique. A duplicate name names both conflicting specs; a length mismatch prints both lists.Video keys: every video / image spec must have its own
key. Two video specs sharing akeyare rejected — only numeric specs concatenate onto a sharedkey(see Selectors and lerobot_names).Alignment method:
align.methodmust be one ofhold,nearest,linear, ornone.Zero tolerance:
align.tolerance_ms: 0is rejected — see Alignment.Legacy keys: the
strategyandtol_msalignment keys are rejected — usemethodandtolerance_ms. A legacypublish:block on an action spec is likewise rejected.Required fields: a missing or wrong-typed required field on a spec (for example, no
keyor notype) raises an error naming that spec’s position and key, such asobservations[2] (key 'observation.state'): ....At runtime, every topic referenced by the contract must be present on at least one file in the dataset, otherwise the action stops before doing any work.