Task Evaluation Runs · for site operators
Find out what a robot could do here, before anyone shows up.
You do not need a policy, a vendor, or an evaluation stack to start. Describe the job you would hand to a robot and the terms you would hand it under. We turn that into a testbed we maintain and a decision you can inspect.

Same service, different entry
You should not have to become an evaluation expert to get an answer.
Robot teams arrive with candidates and a threshold. You can arrive with a job and a set of conditions. Both end up in the same run.
Start with the job, not a robot
The workflow, the shifts, the conditions, the failures you will not accept. That is enough to scope a run.
See what is missing
The run separates what your site can already answer from what needs stronger evidence — or real hardware.
Judge candidates on your terms
When vendors or policies arrive, they enter the same run, against the same task and the same threshold.
Keep the parts that matter to you
Capture windows, restricted areas, privacy, who may use the evidence, and whether anyone tests on your floor.

02 · What you keep
Your floor, your rules — written down before anything is captured.
Restricted areas are excluded before capture rather than redacted afterwards, and what a run's evidence may be used for is recorded per artifact. Nobody tests on your floor without a separate, explicit yes.
- Capture windows
- You set when, where, and who is on site.
- Restricted areas
- Excluded before capture, not redacted after.
- Evidence use
- What the run may be used for is written down, per artifact.
- Physical access
- Any test on your floor is a separate, explicit yes.
How a run moves
One real task in. One inspectable answer out.
Candidates are optional at the start. The task and the terms are what make a run scopeable.
- 01
Bring one real task
A specific job at a specific site: the conditions, who can be there, what counts as success, and what must never happen.
Stage 1 of 5 - 02
We build the testbed
The site-task becomes a captured, versioned testbed we keep — so this answer and the next one are measured against the same thing.
Stage 2 of 5 - 03
You state the decision
Not “run a benchmark.” The actual call you are about to make, the candidates in front of you, and what a wrong yes would cost.
Stage 3 of 5 - 04
We screen on measurement
Reach, clearance, footprint, and sightlines come off the capture. Candidates the building will not take are out before anyone spends a rollout on them.
Stage 4 of 5 - 05
We order what survives
The remaining candidates are ranked on the same testbed version, with the margin, the interval on it, and the smallest gap the run can resolve.
Stage 5 of 5
What comes back
Including the answer nobody likes to sell.
If the evidence will not carry the decision, the run says so and names what would. That is the version of this service worth buying twice.
Ordered, with margin
The candidates are ranked and the gaps clear the run's resolution. You get the order, each margin, and the conditions it holds under.
Ruled out on measurement
A candidate does not physically fit the site. The cheapest finding in a run, and the one that names its own cause.
Ordered in part
Some pairs separate and some sit inside the resolution. You see which is which, instead of a full ranking implying precision the design does not have.
Inside the resolution
The gaps are smaller than this design can separate. The run reports the floor and what it would take to get under it.
Test this next
The least expensive experiment that would move the decision — more rollouts, a recapture, or real hardware.

Before a vendor arrives
Know what the honest answer is before someone pitches you one.
A maintained testbed of your own task means every candidate that shows up later is measured against the same job, the same conditions, and the same threshold — yours.
Where we stop
The limits, stated plainly.
A run is evidence, not permission
Nothing we return is a safety approval, a certification, or a licence to operate. Those stay with you and your regulator.
An ordering is bounded by its testbed
A ranking holds on the testbed version it was measured on, under the conditions stated. We have not measured how our orderings track real-world orderings, and we do not inherit anyone else's correlation figures as if they were ours.
Resolution is a property of the design
Every run can only separate gaps above a certain size. We publish that floor with the ordering, because a rank you cannot separate is not a result.
Some claims need hardware
Contact-heavy and safety-critical questions often cannot be settled short of real robots. When that is the case, the run says so.
Where we are strongest today
Navigation, mobile-base movement, and rigid pick-and-place in warehouse and logistics spaces. We will tell you when your task is outside that.
Start with the job
Tell us what you need to decide.
Describe the real workflow, the conditions, the failures you will not accept, and the terms of access. Candidates can be linked whenever they exist.
