Roboto to LeRobot Contract#

The roboto-to-lerobot-v2_1 and roboto-to-lerobot-v3_0 actions are driven by a contract YAML stored in the invocation dataset — the dataset the action is invoked on. The contract tells the action which topics to read from each recording, how to extract values from each message, how to align streams onto a common timeline, and what transforms to apply along the way. Every event in the input collection becomes one episode in the output LeRobot dataset, and the contract is applied identically to every episode.

This page is the full schema reference. For a step-by-step walkthrough of authoring a contract and running the action, see Convert to LeRobot.

Full example#

The contract below exercises the full schema: a camera stream, numeric observations that concatenate onto a shared key, a role-bound spec, per-stream alignment and transforms, an action stream, and a task channel. Each piece is unpacked in the sections that follow. (For a minimal starting point, see the user guide instead.)

name: pick_place
version: 1
fps: 20
robot_type: dual_arm
action_lead_steps: 1     # read actions one frame ahead of observations

observations:
  # Camera -> a video feature; the image: block routes it to the video pipeline.
  - key: observation.images.exo
    topic: /camera/exo/image_raw/compressed
    type: sensor_msgs/msg/CompressedImage
    image:
      resize: [480, 640]   # [height, width]
    align: {method: nearest, tolerance_ms: 100}

  # Left arm -> the first three columns of observation.state.
  - key: observation.state
    topic: /left_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [left_1, left_2, left_3]
    align: {method: hold, tolerance_ms: 100}
    transforms:
      - type: butterworth_lowpass_causal   # causal: reproducible at inference
        cutoff_hz: 5
        order: 2
        fs_hz: 50        # freeze the design rate into the contract

  # Right arm -> concatenated onto the same feature (shared key).
  - key: observation.state
    topic: /right_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [right_1, right_2, right_3]
    align: {method: hold, tolerance_ms: 100}

  # Base velocity, bound by role instead of topic.
  # No align: -> hold with the auto-bounded tolerance.
  - key: observation.base_velocity
    role: base_twist
    type: geometry_msgs/msg/Twist
    selector:
      names:         [linear.x, angular.z]
      lerobot_names: [base_vx, base_wz]

actions:
  - key: action
    topic: /left_arm/joint_commands
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [left_cmd_1, left_cmd_2, left_cmd_3]
    align: {method: linear, tolerance_ms: null}   # null = no tolerance bound
    transforms:
      - type: resample_uniform
        stage: pre                  # pre only; must be set explicitly
      - type: butterworth_lowpass   # zero-phase; offline-only
        stage: pre
        cutoff_hz: 30
        order: 2
      - type: finite_difference     # stage defaults to "post"

tasks:
  - key: prompt          # optional; defaults to the topic name
    topic: /task/prompt
    type: std_msgs/msg/String

Top-level fields#

Field

Type

Default

Description

name

string

"contract"

Logical contract name. Recorded with each episode for traceability.

version

int

1

Contract version. Bump it when you change the schema in a way that breaks downstream consumers.

fps

float

20.0

Target sampling frequency, in Hz. A reference timeline of one frame per 1/fps seconds is built at this rate, and every observation and action is aligned onto it.

robot_type

string

null

Free-form label written into the LeRobot dataset metadata.

action_lead_steps

int

0

Shift action sampling N frames into the future to compensate for control delay. 0 means action[t] is read at the same instant as observation[t].

observations

list

[]

Observation streams. Entries with an image: block become video features; entries without become numeric observation features. See Observation specs.

actions

list

[]

Action streams. Become numeric action features. See Action specs.

tasks

list

[]

Optional task channels used to derive each episode’s task label. See Task specs.

Observation specs#

Each entry in observations: describes one stream. There are two flavors, distinguished by whether an image: block is present:

  • Numeric observations (no image: block) become a single LeRobot feature per key. Multiple entries that share a key are concatenated along the feature axis — see Selectors and lerobot_names.

  • Video / image observations (with an image: block) become a dtype: video feature. See Video and image specs.

Field

Required

Description

key

yes

LeRobot feature key. Conventionally observation.state, observation.images.<name>, and so on.

topic

yes

ROS topic name to read the stream from. Not required if the spec binds by role: instead; see Role-based binding.

type

yes

Message type string used to select the decoder. See Supported message types.

selector

no

{names: [...], lerobot_names: [...]}. Selects which fields the decoder extracts and how they are labeled. See Selectors and lerobot_names.

image

no

Image / video options, chiefly resize: [H, W]. Its presence routes the spec to the video pipeline. See Video and image specs.

align

no

{method, tolerance_ms}. Defaults to hold with the auto-bounded tolerance. See Alignment.

transforms

no

A list of transforms applied to the stream. See Transforms.

Video and image specs#

An entry under observations: becomes a video feature when it carries an image: block. The output dtype is always video (HWC uint8 RGB at the configured resize).

observations:
  - key: observation.images.exo
    topic: /camera/exo/image_raw/compressed
    type: sensor_msgs/msg/CompressedImage
    image:
      resize: [480, 640]   # [height, width]
    align: {method: nearest, tolerance_ms: 100}

The image: block currently accepts a single field:

Field

Required

Description

resize

yes

[height, width] in pixels. Frames are resized to this exact shape.

A camera that publishes a compressed video stream rather than per-message stills is declared the same way, with its own type: — see Compressed video.

Restrictions on image streams:

  • align.method: linear is rejected — use hold or nearest.

  • Depth encodings are not yet supported. Supported raw sensor_msgs/msg/Image encodings are rgb8, bgr8, mono8, rgba8, bgra8, and 8UC1. An image.depth: block is currently rejected at contract load; depth support is planned for a future release.

Action specs#

Action specs follow the same shape as numeric observations — the same fields, with no image: block.

actions:
  - key: action
    topic: /robot/joint_commands
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2]
      lerobot_names: [arm_1, arm_2]
    align: {method: hold, tolerance_ms: 100}

As with observations, multiple entries that share a key are concatenated along the feature axis. Set topic: and type: directly on the action spec, as above.

Alignment on action streams:

  • All four alignment methods apply — hold (the default), nearest, linear, and none. The inference caveat that limits observation streams does not apply to actions; see the note under Alignment for why every method is safe here.

  • To shift action sampling relative to observations — for example, to compensate for control delay — set action_lead_steps at the top level rather than per spec. See Top-level fields.

Task specs#

tasks: is optional and lets the contract derive a per-episode LeRobot task label from a message stream — for example, a language prompt published on a topic.

tasks:
  - key: prompt          # optional; defaults to the topic name
    topic: /task/prompt
    type: std_msgs/msg/String

A task spec sets topic: and type: — use std_msgs/msg/String; other message types fail when the stream is decoded — plus an optional key:, which defaults to the topic name.

For each episode, the task label is resolved in the following order:

  1. The event’s task metadata field, when present and non-empty. This is an explicit per-episode annotation and always wins; the tasks: block is not consulted.

  2. Otherwise, the payload of the first std_msgs/msg/String message that falls within the episode’s own time window, taken from the first declared task spec. If the contract declares more than one task spec, only the first is used and the others are ignored (a warning names them).

  3. Otherwise, the literal string "default" — when there is no task metadata, no tasks: block, and no in-window message.

Selectors and lerobot_names#

selector: controls which fields the decoder pulls out of each message:

selector:
  names:         [joint_1, joint_2, joint_3]   # what to extract
  lerobot_names: [arm_1,   arm_2,   arm_3]     # how to label them (optional)

Two rules to keep in mind:

  1. The meaning of names depends on the message type. For JointState they are joint names (with an optional position.<joint> / velocity.<joint> / effort.<joint> prefix; a bare name defaults to position). For Imu, Odometry, and Twist they are dotted paths into the message (for example, linear.x). See Supported message types for the full list.

  2. The lerobot_names field is optional but strongly recommended. Without it, a feature is labeled <topic>/<selector_name>, which is verbose and ties feature names to your ROS topology. With it, you get clean LeRobot-facing names. lerobot_names must have the same length as names, and every lerobot_name must be unique across the entire contract.

When multiple specs share a key (for example, two observation.state entries from two arms), all of their lerobot_names are concatenated in order to name the combined feature.

observations:
  # Left arm -> the first three columns of observation.state
  - key: observation.state
    topic: /left_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [left_1, left_2, left_3]

  # Right arm -> concatenated onto the same feature
  - key: observation.state
    topic: /right_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [right_1, right_2, right_3]

The combined observation.state feature then has six columns, labeled left_1, left_2, left_3, right_1, right_2, right_3 in that order.

Alignment#

Each stream is merged onto the reference timeline (one frame per 1/fps seconds) using its align: block:

align: {method: hold, tolerance_ms: 100}

Method

Behavior

hold

Last-observation-carried-forward (backward as-of join). Good for slow state signals. Default.

nearest

Pick the closest sample in time, in either direction. Good for cameras and high-rate signals.

linear

Linearly interpolate between the two bracketing samples. Numeric streams only — not allowed for images.

none

Exact timestamp matches only; everything else becomes NaN.

Note

Reproducing alignment at inference. If you plan to deploy the trained policy for online, streaming inference, only hold and nearest can be reproduced sample-by-sample on a live input; linear and none need the complete recording and are therefore offline-only — the same distinction as non-causal vs causal Transforms. Use hold or nearest for any observation stream whose processing must match between training and deployment. Action streams are exempt — a deployed policy emits actions rather than aligning them — so all four methods are always fine there.

tolerance_ms is the maximum gap, in milliseconds, between the reference frame timestamp and the matched sample before the result becomes NaN. Set it to roughly one period of your slowest signal — for example, 100 for a 10 Hz signal.

Tolerance_ms

Behavior

omitted (or align: omitted entirely)

Auto-bounded default of max(2/fps, 50 ms). This is usually a reasonable starting point.

null

Unlimited — always carry forward, or always pick the nearest sample, however far away it is.

0

Rejected. In older contracts 0 meant “unlimited”; that spelling collided with “zero tolerance” and is no longer accepted. Use null for unbounded, or omit tolerance_ms for the auto-bounded default.

a positive number

That bound, in milliseconds, unchanged.

Note

The legacy field names strategy and tol_ms are no longer accepted. Use method and tolerance_ms.

Transforms#

Transforms are applied per stream, in the order they appear, before the stream becomes part of the LeRobot dataset:

actions:
  - key: action
    topic: /robot/joint_commands
    type: sensor_msgs/msg/JointState
    selector: {names: [j1, j2], lerobot_names: [arm_1, arm_2]}
    transforms:
      - type: resample_uniform
        stage: pre
      - type: butterworth_lowpass
        stage: pre
        cutoff_hz: 30
        order: 2
      - type: finite_difference     # stage defaults to "post"

Stages#

A transform’s stage says whether it runs before or after alignment — the step that merges each raw stream onto the shared reference timeline of one frame per 1/fps seconds (see Alignment). A pre transform therefore sees the raw, irregularly timed samples; a post transform sees the regular, fps-rate timeline.

Stage

When it runs

Notes

pre

Before alignment, on the raw stream.

Operates on the raw sample timestamps. Use it for resampling and pre-alignment filtering.

post

After alignment, on the reference timeline.

Operates at fps. This is the default when stage: is omitted.

Built-in transforms#

Type

Stage

Causal

Params

Description

resample_uniform

pre only

no

—

Resamples an irregularly sampled signal onto a uniform grid (the median inter-sample interval) using nearest-neighbor. Always set stage: pre explicitly — this transform cannot run after alignment, and the stage: default is post.

butterworth_lowpass

any

no

cutoff_hz (required), order (default 2)

Zero-phase Butterworth low-pass filter. No phase distortion, but each output sample depends on both past and future samples.

butterworth_lowpass_causal

any

yes

cutoff_hz (required), order (default 2), fs_hz (optional)

Causal (forward-only) Butterworth low-pass filter. Each output sample depends only on past samples, at the cost of some phase lag.

finite_difference

any

no

—

Frame-to-frame difference (out[i] = in[i+1] - in[i]). The last row is repeated to keep the shape constant.

Causal vs non-causal transforms. A causal transform computes each output sample from the current and past samples only, so the identical operation can also run online — sample-by-sample at inference time — and reproduce exactly what training saw. A non-causal transform additionally looks at future samples, so it can only ever run offline over a complete recording. The Causal column above marks which is which. When a stream feeds a deployed policy and the same processing must run at both training and inference time, choose a causal transform — for example butterworth_lowpass_causal, the causal counterpart of butterworth_lowpass. When only offline processing matters — for example, cleaning an observation that plays no role in a deployed policy — a non-causal transform is fine, and because a zero-phase filter such as butterworth_lowpass introduces no phase lag, it is often preferable.

Cutoff bounds. cutoff_hz must be strictly between 0 and the Nyquist frequency (half the sample rate). For a stage: pre transform the sample rate is the measured raw topic rate; for stage: post it is fps.

Sample-rate override (causal filter only). fs_hz declares the sample rate the filter is designed for, overriding the rate the transform would otherwise use. For a stage: pre causal filter, declaring fs_hz is what makes the filter reproducible during deployment: it freezes the design rate into the contract so offline and online both build identical filter coefficients. If the declared fs_hz deviates from the stream’s actual measured rate by more than 10%, a warning names both numbers.

Ordering. When you chain a stage: pre transform that changes the time axis (resample_uniform) with later filters, place resample_uniform first and any low-pass filter after it. The sample rate passed to the later transforms is recomputed automatically from the resampled timestamps.

Note

The built-in transform library will be expanded in future releases.

Role-based binding#

Instead of pinning a spec to a specific topic:, you can bind it by role — a logical label you attach to the source file. This helps when the same logical stream lives under different topic names across recordings: you bind the spec to a stable role once, and each recording maps that role to whatever its actual topic happens to be.

To use it, set role: on the spec in place of topic:, then tag each source file with a matching role by setting a role field in the file’s Roboto metadata to the same value. For example, a spec with role: arm_joints binds to the file whose metadata has role: arm_joints. You can set this metadata manually — in the web UI, CLI, or SDK — so no companion action is required, though one may be provided to populate roles automatically.

observations:
  - key: observation.state
    role: arm_joints          # instead of topic:
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [arm_1, arm_2, arm_3]

Rules and failure modes:

  • A spec must set exactly one of topic: or role:. Setting both, or neither, is rejected at contract load.

  • Image and video streams cannot be role-bound — bind them by topic:.

  • Within a dataset, exactly one file may carry a given role. Zero matching files, or more than one, stops the run with an error naming the role.

  • The matched file must have exactly one topic whose fields cover the spec’s selector.names; zero or multiple covering topics is an error.

Supported message types#

type: selects the decoder. The selector semantics depend on the type.

Image and video#

type:

Required image: fields

sensor_msgs/msg/CompressedImage

resize

sensor_msgs/msg/Image

resize

video, avi_video, mp4_video

resize. File-backed video, one still frame per row.

foxglove_msgs/msg/CompressedVideo

resize. A compressed video stream (H.264, H.265, VP9, or AV1). See Compressed video.

Supported sensor_msgs/msg/Image encodings: rgb8, bgr8, mono8, rgba8, bgra8, and 8UC1.

Compressed video#

Topics that Roboto ingestion tags compressedVideo store one encoded video access unit per message rather than a standalone still image, so a frame in the middle of a group of pictures (GOP) can only be decoded together with the frames back to its keyframe. The action handles that: it decodes each episode’s whole time range in one pass, walking back up to 10 seconds to find the keyframe that anchors the range.

Declare the topic’s real schema name. Three spellings are accepted — foxglove_msgs/msg/CompressedVideo, foxglove_msgs/CompressedVideo, and foxglove.CompressedVideo. Declaring a compressed-video topic as sensor_msgs/msg/CompressedImage is rejected with an explanatory error rather than silently producing broken frames.

Apart from the type:, the spec is an ordinary video spec — an image: block carrying resize, plus an align: block:

observations:
  - key: observation.images.external
    topic: /camera/exo/video
    type: foxglove_msgs/msg/CompressedVideo
    image:
      resize: [480, 640]   # [height, width]
    align: {method: nearest, tolerance_ms: 150}

  - key: observation.images.gripper
    topic: /camera/hand/video
    type: foxglove_msgs/msg/CompressedVideo
    image:
      resize: [480, 640]
    align: {method: nearest, tolerance_ms: 150}

Notes and limits:

  • The codec is read from each message’s format field. A stream in a codec other than H.264, H.265, VP9, or AV1 fails the conversion rather than dropping the camera.

  • resize is applied as the frames are decoded, which keeps an episode’s frames in memory at the target resolution rather than at the source resolution.

  • Frames whose keyframe is unreachable are skipped with a warning. A range where nothing decodes fails the conversion.

Numeric streams#

type:

What selector.names means

sensor_msgs/msg/JointState

Joint name. Optional position.<joint> / velocity.<joint> / effort.<joint> prefix; a bare name defaults to position. Without names: all joint positions in message order.

trajectory_msgs/msg/JointTrajectory

Joint name. Each trajectory point becomes its own sample at header.stamp + time_from_start. Intended for action streams.

control_msgs/msg/MultiDOFCommand

DOF name. Optional values.<dof> / values_dot.<dof> prefix; a bare name defaults to values. Without names: all values followed by all values_dot.

sensor_msgs/msg/Imu

Dotted path into the message, for example orientation.x or angular_velocity.z. Without names: returns [quat, ang_vel, lin_acc] (10 values).

nav_msgs/msg/Odometry

Dotted path. Without names: returns [pos.xyz, quat.xyzw] (7 values).

geometry_msgs/msg/Twist

Dotted path, for example linear.x. Without names: returns [linear.xyz, angular.xyz] (6 values).

std_msgs/msg/Float32MultiArray

not used; the full data array is emitted as float32.

std_msgs/msg/Float64MultiArray

not used; full data as float64.

std_msgs/msg/Int32MultiArray

not used; full data as int32.

std_msgs/msg/Float32 / Float64 / Int32 / Int64

not used; a single-element array containing data.

std_msgs/msg/String

not used; emitted as a string. Intended for tasks:.

string_typed_msg

Series index keys, each read as a float. names is required for this type.

anymal_msgs/AnymalState

ANYmal joint name into joints. Optional position.<joint> / velocity.<joint> / acceleration.<joint> / effort.<joint> prefix; a bare name defaults to position. Without names: all joint positions in message order.

series_elastic_actuator_msgs/SeActuatorReadings

ANYmal joint name, mapped positionally — the message carries no per-actuator name — onto the fixed order LF/RF/LH/RH × HAA/HFE/KFE. Optional field-path prefix, for example commanded.velocity.<joint> or state.joint_position.<joint>; a bare name defaults to commanded.position. Without names: all 12 commanded.position.

Validation rules and common errors#

The contract loader checks a few things up front. When it rejects a contract, the error message names the offending spec.

  • Binding: every observation and action spec must set exactly one of topic: or role: (see Role-based binding). Image and video streams must use topic:.

  • Image alignment: image streams cannot use align.method: linear.

  • Compressed video: a topic that stores compressed video must be declared with a compressed-video type:. Declaring it as sensor_msgs/msg/CompressedImage — or as a compressed-video spelling that has no decoder — stops the run, naming the offending spec’s key (see Compressed video).

  • Names: lerobot_names must match names in length, and every lerobot_name in the contract must be globally unique. A duplicate name names both conflicting specs; a length mismatch prints both lists.

  • Video keys: every video / image spec must have its own key. Two video specs sharing a key are rejected — only numeric specs concatenate onto a shared key (see Selectors and lerobot_names).

  • Alignment method: align.method must be one of hold, nearest, linear, or none.

  • Zero tolerance: align.tolerance_ms: 0 is rejected — see Alignment.

  • Legacy keys: the strategy and tol_ms alignment keys are rejected — use method and tolerance_ms. A legacy publish: block on an action spec is likewise rejected.

  • Required fields: a missing or wrong-typed required field on a spec (for example, no key or no type) raises an error naming that spec’s position and key, such as observations[2] (key 'observation.state'): ....

  • At runtime, every topic referenced by the contract must be present on at least one file in the dataset, otherwise the action stops before doing any work.