Humanoid robot benchmarking
A benchmark is easier to interpret when the task contract remains stable and the execution variables are visible.
Start with a benchmarkable behavior
Define one task with clear success criteria, constraints and observable outcomes. Avoid changing the behavior definition and the policy at the same time when the goal is to attribute a performance change.
Fix the comparison variables
Record the behavior version, policy version, engine and selected seeds. Stable comparison variables make differences easier to attribute.
Choose metrics that answer the task
Metrics should describe the behavior being evaluated. Depending on the task, that may include success, contacts, cost, completion state or other task-specific measurements.
Report the context with the number
A metric without provenance is difficult to interpret. Store the configuration and execution status alongside the measured output.
Benchmarking is a workflow, not a single score
Good evaluation includes definition, setup, execution, inspection and replay. The number is one part of the evidence, not the entire record.
Related resources
Humanoid robot behavior · Robot policy evaluation · Behavior-driven evaluation