Know which robot policy works at a real site — before field time.
For robot and foundation-model teams, Blueprint does one thing today: Turn a pre-sales site walkthrough into a maintained testbed: rule out candidate policies or checkpoints the site will not physically accept, then rank the remainder against that site-task with the margin, its interval, and the resolution floor of the design. It's an estimate and a decision-support screen — never a guarantee, a safety certification, or a deployment-readiness claim.
Longer-horizon, that neutral measurement is the first rung of a climb — toward the standard the market routes deployment decisions through, and the proprietary data that opens the door to prediction, site-specific policies, and deployment itself. Capture-first, provenance-true, the whole way up. Only rung one is a product you can buy today; everything above it is where we're heading.
- Robot data vs. text
- ~1B×
- SC3-Eval research
- 0.929
- One real eval
- 2,500+
- Reliability bar
- 99.99%
Smaller than internet text (Bessemer)
Published correlation; not a Blueprint result
Rollouts + 100+ human hours (AutoEval)
What industrial buyers expect (Bain)
Review support · not real-world proof“When bodies and brains are both plentiful, the scarce, valuable thing is a trustworthy way to compare them on a real site — and the data that comparison produces.”
A single rigorous real-world evaluation of one policy can take thousands of rollouts and a hundred hours of human labor. New research from NVIDIA, Physical Intelligence, and leading universities — SC3-Eval and OSCAR (2026) — suggests generated worlds may support some policy-comparison claims. Those are external research signals, not a Blueprint result. Blueprint starts with claim-specific qualification, a bounded decision or abstention, and the proprietary physical outcome data that lets the evidence improve.
The path up the stack
One product today. Four more rungs it's built to reach.
We don't skip rungs. Only rung one is a product you can buy today; each rung above it is a moat we'd deepen and the launchpad for the next. Everything above rung one is where the same capture-first foundation is taking us — a direction, not a shipped offer.
Know what the evidence supports — before field time.
Blueprint's Task Evaluation Runs evaluate the decision-relevant claims on a real captured site against task, success, cycle-time, intervention, and risk thresholds. A run may rank candidates only when the evidence supports ordering; partial decisions and abstention are first-class outcomes. Current virtual evidence is strongest for navigation, mobile-base movement, and rigid pick-and-place in warehouse and logistics spaces. Contact-rich or safety-critical claims require a stronger validation envelope and may require physical evidence. External research can motivate methods, but it is not a Blueprint run result.
Proof boundaryPer-claim evidence, validation envelope, and uncertainty — never a guaranteed field outcome or safety certification.
Where this goes over time — direction, not shipped product
Become the neutral standard both sides route decisions through.
The aspiration is that robot teams use our runs to prove readiness and win pilots, and that over time site operators come to ask for them before a robot reaches the floor. The goal is that a large share of deployment and pilot decisions eventually pass through one trusted, neutral measurement — the way credit ratings, UL safety marks, and MLPerf became the scoreboard their industries transact against. This is a decision layer we are building toward, not a marketplace and not a gate we operate today.
Proof boundaryNeutrality is the asset. Visible methodology, re-validation, and a conflict-of-interest firewall.
Predict real-world performance, and generate the data to improve it.
Every deployment decision routed through Blueprint is a labeled, ground-truth outcome. That proprietary, multi-site capture is the scarcest input in robotics — and research shows policy generalization scales with the diversity of real environments, not raw demo count. Today's evaluators are strong in-distribution but weaker on unfamiliar sites (SC3-Eval drops from 0.98 to ~0.87 out-of-distribution); every site we capture pulls more of the real world in-distribution. That is what powers site-specific post-training data and, over time, calibrated prediction: getting a 95% eval to mean ~95% in the real world.
Proof boundaryCalibrated prediction depends on multi-year world-model progress. We publish the dependency, not a promise.
Site-specific policies, measured on a neutral scoreboard.
Because we hold provenance-clean data for each site we capture, we can fine-tune policies specialized to that exact environment — and use our own neutral evaluation to test the honest question of whether they beat the alternatives on that site. An edge only counts when our own scoreboard says so.
Proof boundaryOnly claimed when the neutral eval measures it — and only behind a structural neutrality firewall.
Help run the deployment where we can prove we're the best operator.
As robot hardware commoditizes, the durable value moves to the intelligence and the operating relationship. Where a site is best served by our per-site policy — and only where our neutral evaluation proves it — Blueprint can help operate the deployment on commodity hardware. We keep the standard credibly independent from any operating arm.
Proof boundaryAn option we earn, never an assumption. Neutrality is protected structurally before this step.
More sites captured → better, more diverse evaluations → more deployment decisions routed through us → more proprietary real-world outcome data → better prediction and data → better site-specific policies → more deployments we can credibly serve → which funds more capture.
The first rungs deliver the product and compound the capture network at the same time. Request-scoped outcomes can improve future evaluation design only when rights and provenance permit that use; they are never silently repurposed as ground truth.
What never changes
Four commitments that hold on every rung.
The higher we climb, the more these matter. They are what keep the measurement trustworthy and the company honest.
Capture first, always.
Every stage is built on real, rights-clean, provenance-true site capture. That is the moat that grows stronger as models commoditize — not weaker.
The model backend stays swappable.
No stage couples the company to one checkpoint, provider, or world model. A better model later is a drop-in behind the adapter boundary, not a rebuild.
Estimates, never guarantees.
Rank fidelity and predicted success — with proof boundaries and missing-proof labels — all the way up. We never turn a correlation into a promise.
Neutrality is an asset we protect.
From the standard onward, independence is structural. Any move to build our own policies or operate deployments is gated on a credible firewall that keeps the measurement trusted.
Start at rung one
The vision is long. The product is real today.
Evaluate a policy on a real captured site before you spend field time — that's where the whole climb begins.
