EMBODIED AI / REPRODUCIBILITY

Reproducible embodied AI experiments

Reproducibility is a record-keeping problem as much as a simulation problem: the experiment needs to preserve enough context to explain what actually ran.

Keep five layers connected

Task: what behavior is being evaluated and what counts as success.

Version: which task specification was used.

Policy: which policy or controller was evaluated.

Environment: which simulation engine and configuration executed the task.

Evidence: which seeds, status and measured metrics were produced.

Why explicit provenance matters

Two runs can have the same headline metric while representing different task definitions or configurations. Recording provenance alongside measurements makes later comparisons more defensible.

Make failed runs useful

A failed or incomplete experiment can still reveal configuration and execution information. Preserve its status and context instead of treating it as if no experiment happened.

Build toward replay

Replay is most useful when the stored artifact contains the information needed to reconstruct the run context. It should remain distinct from a claim that a new physical-world trial was performed.

See the workflow → Read the docs →

Related resources

Embodied AI evaluation · Robot behavior evaluation · MuJoCo simulation guide