IROS Artifact Evaluation
IROS has no formal artifact-evaluation track or badge, so "artifact" here means the evidence bundle a
reviewer could inspect and that a careful author prepares anyway. The audience is an embodied-systems
reviewer who trusts a logged trajectory over a claimed one, and who cannot rerun your robot but can read
your logs, configs, and code.
The robotics artifact stack
- Hardware ledger: platform, sensors (make/model/firmware), actuators, onboard computer, and power
budget — the embodiment the result depends on.
- Logs: time-stamped rosbag-equivalent recordings for representative trials, so a claim about a
contact force or trajectory is auditable.
- Configs: the exact parameters, calibration, and launch files that produced the reported runs.
- Code: the system as run, with an entry point and a one-minute orientation in the README.
- Data: any dataset or object set, with provenance and licensing; for restricted data, enough detail
for credible reproduction without violating terms.
What an IROS reviewer inspects first
| Claim type |
First artifact inspected |
Common failure caught |
| "Runs onboard at rate X" |
Timing logs and the compute/power spec |
Rate measured on a desktop, not the robot |
| "Reliable across trials" |
The trial log and failure record |
Success rate with resets and failures omitted |
| "Transfers from sim" |
Paired sim and real logs |
Only sim logs exist; the gap is asserted |
| "Beats prior system" |
Baseline configs on the same platform |
Baseline run with different hardware or tuning |
Because a reviewer will read a log far sooner than rebuild a robot, make the logs and configs
legible first, and polish the code second.
Review-time versus acceptance-time states
- Review time (double-anonymous): strip organization names, cluster paths, calibration files tagged
with a lab, commit authorship, and any URL carrying a lab identity. If you link code, use an anonymized
mirror or omit the link and describe the artifact.
- Acceptance time: replace anonymized stand-ins with a public, licensed, citable release; test every
link from a logged-out browser; and add the DOI or archival link the IEEE Xplore record can point to.
Worked vignette: packaging a field-robot result
A submission claims a lidar-inertial system runs onboard across three outdoor sites.
- Ship one representative rosbag per site plus the launch/config that processed it, so a reviewer can see
the sensor streams and the estimated trajectory line up.
- Record the compute-and-power measurement method, not just a headline rate.
- Emit the trajectory-error tables directly from the logged runs so the paper numbers and the artifact
cannot drift apart.
- State which site is easy and which stresses the system, since that mapping is what a reviewer grades.
Artifact layout:
/hardware.md platform, sensors, compute, power
/logs/<site>/ time-stamped bags for representative trials
/config/ calibration, params, launch files
/src/ system as run + README (1-minute orientation)
/data/ provenance + license (or access instructions)
Output format
[Artifact role] review-time anonymous / acceptance-time public
[Contents] <hardware/logs/configs/code/data>
[Anonymity risks] <paths/metadata/URLs/calibration tags>
[Auditability] logged / scripted / described / asserted
[Fixes before upload] <ordered list>
1---2name: iros-artifact-evaluation3description: Use when packaging IROS evidence for a skeptical reviewer even though IROS runs no formal artifact track — the robotics artifact stack of hardware ledger, logs, configs, code, and data, what an embodied-systems reviewer opens first, review-time anonymity versus acceptance-time public release, and making a claim auditable not asserted.4---56# IROS Artifact Evaluation78IROS has no formal artifact-evaluation track or badge, so "artifact" here means the evidence bundle a9reviewer *could* inspect and that a careful author prepares anyway. The audience is an embodied-systems10reviewer who trusts a logged trajectory over a claimed one, and who cannot rerun your robot but can read11your logs, configs, and code.1213## The robotics artifact stack1415- **Hardware ledger:** platform, sensors (make/model/firmware), actuators, onboard computer, and power16 budget — the embodiment the result depends on.17- **Logs:** time-stamped rosbag-equivalent recordings for representative trials, so a claim about a18 contact force or trajectory is auditable.19- **Configs:** the exact parameters, calibration, and launch files that produced the reported runs.20- **Code:** the system as run, with an entry point and a one-minute orientation in the README.21- **Data:** any dataset or object set, with provenance and licensing; for restricted data, enough detail22 for credible reproduction without violating terms.2324## What an IROS reviewer inspects first2526| Claim type | First artifact inspected | Common failure caught |27|---|---|---|28| "Runs onboard at rate X" | Timing logs and the compute/power spec | Rate measured on a desktop, not the robot |29| "Reliable across trials" | The trial log and failure record | Success rate with resets and failures omitted |30| "Transfers from sim" | Paired sim and real logs | Only sim logs exist; the gap is asserted |31| "Beats prior system" | Baseline configs on the same platform | Baseline run with different hardware or tuning |3233Because a reviewer will read a log far sooner than rebuild a robot, make the **logs and configs**34legible first, and polish the code second.3536## Review-time versus acceptance-time states3738- **Review time (double-anonymous):** strip organization names, cluster paths, calibration files tagged39 with a lab, commit authorship, and any URL carrying a lab identity. If you link code, use an anonymized40 mirror or omit the link and describe the artifact.41- **Acceptance time:** replace anonymized stand-ins with a public, licensed, citable release; test every42 link from a logged-out browser; and add the DOI or archival link the IEEE Xplore record can point to.4344## Worked vignette: packaging a field-robot result4546A submission claims a lidar-inertial system runs onboard across three outdoor sites.4748- Ship one representative rosbag per site plus the launch/config that processed it, so a reviewer can see49 the sensor streams and the estimated trajectory line up.50- Record the compute-and-power measurement method, not just a headline rate.51- Emit the trajectory-error tables directly from the logged runs so the paper numbers and the artifact52 cannot drift apart.53- State which site is easy and which stresses the system, since that mapping is what a reviewer grades.5455```text56Artifact layout:57 /hardware.md platform, sensors, compute, power58 /logs/<site>/ time-stamped bags for representative trials59 /config/ calibration, params, launch files60 /src/ system as run + README (1-minute orientation)61 /data/ provenance + license (or access instructions)62```6364## Output format6566```text67[Artifact role] review-time anonymous / acceptance-time public68[Contents] <hardware/logs/configs/code/data>69[Anonymity risks] <paths/metadata/URLs/calibration tags>70[Auditability] logged / scripted / described / asserted71[Fixes before upload] <ordered list>72```