Possibility is not readiness

A strong robot demonstration is valuable. It shows that a combination of hardware, perception, control and software can produce a desired behaviour at least once. But the conditions around that success are often invisible: object placement may be carefully chosen, the environment may be simplified, the operator may reset the scene, and failed attempts may never appear.

Real work is less forgiving. Objects vary, tools wear, lighting changes, people interrupt and systems need recovery paths. A single selected run cannot tell an operator how often the task succeeds, what preparation it requires or what happens after failure.

Evidence needs context

RobotBay starts with a public task brief and visible test conditions. The system configuration, environment, number of attempts and human interventions should be recorded. Representative failures matter because they reveal where the system boundary actually sits.

Repeatability does not mean pretending every task is identical. It means describing variation clearly enough that another person can understand what changed and why the result may differ. Evidence should make comparison possible without reducing a complex system to a promotional score.

Boundaries are useful outputs

A benchmark is not complete when it identifies a winner. It is useful when it explains what worked, under which conditions, with what intervention and where the result should not be generalized. A well-described limitation is not an embarrassment; it is practical knowledge for builders and operators.

RobotBay therefore separates observation from interpretation and interpretation from conclusion. We do not manufacture a claim about who is best. We maintain a record that allows others to inspect the method and make their own judgment.