Reproducible embodied AI experiments
Reproducibility is a record-keeping problem as much as a simulation problem: the experiment needs to preserve enough context to explain what actually ran.
Keep five layers connected
Task: what behavior is being evaluated and what counts as success.
Version: which task specification was used.
Policy: which policy or controller was evaluated.
Environment: which simulation engine and configuration executed the task.
Evidence: which seeds, status and measured metrics were produced.
Why explicit provenance matters
Two runs can have the same headline metric while representing different task definitions or configurations. Recording provenance alongside measurements makes later comparisons more defensible.
Make failed runs useful
A failed or incomplete experiment can still reveal configuration and execution information. Preserve its status and context instead of treating it as if no experiment happened.
Build toward replay
Replay is most useful when the stored artifact contains the information needed to reconstruct the run context. It should remain distinct from a claim that a new physical-world trial was performed.
Related resources
Embodied AI evaluation · Robot behavior evaluation · MuJoCo simulation guide