Skip to content
Moduloa
← Journal
MOD-01 · Field notes

The measurement layer arrived, and it points at the robot

Research

Three companies in a single Y Combinator batch and one ISO working draft are all building standardized, reproducible measurement for robots. It is the framework layer this site keeps saying is missing, arriving fast, aimed at a different object. Nothing in it touches whether a qualified production process can move between two sites.

2026-09-12 · Field notes · 7 min read · By

Two weeks ago this site argued that Alphabet had published the cell and nobody had published the certificate. The hardware and the software for a standard work cell exist. The thing that would let a qualified process move between two of them, owned by two different companies, does not.

The obvious counter-argument was that the certification layer would arrive once there was enough robot to certify. The evidence this fortnight is that it is arriving. It is measuring the robot.

The batch

Three companies funded to grade a robot

Y Combinator Summer 2026 · Demo Day 10 September

The batch is roughly 235 companies, the largest Y Combinator has run, and it pitched on Thursday 10 September. Industrials took about 23 percent of it, up from 12.8 percent in the previous batch. The composition of that surge was read here on 5 September as humanoids, warehouse robotics, defense and training infrastructure, with no micro-factory or distributed-manufacturing entrant found. That reading was taken before Demo Day, and it missed something.

Three companies in this batch are building measurement itself. Robocurve, three people in San Francisco under founder Jay Chooi, describes itself as evals for robots: open-source benchmarks and independent real-world evaluations, on the argument that frontier labs currently evaluate in-house and publish demonstration videos, so nobody knows where the frontier actually sits. Its framework is called Inspect Robots. It is incorporated as a Public Benefit Corporation.

Instance, founded by Claire Mao and Lucy Cai, automates the evaluation loop rather than the benchmark. Its first product is a success detector that judges whether a robot actually completed its task, replacing the human who waits for the rollout, scores it, and resets the scene.

Hebbian Robotics launched publicly on 31 August with HFlow, an SDK that turns multimodal recordings from robots and human operators into standardized, quality-checked episodes and queryable dataset manifests. It is Apache-2.0 and pre-v1. The problem its founders describe is provenance: once a corpus grows past scripts, it becomes hard to know which code ran, why an episode was excluded, or whether a dataset can be reproduced.

That last sentence is a factory problem, written about a training corpus.

The standard

ISO is drafting the same object

ISO/WD 26264-1 · ISO/TC 299/WG 16 · working draft, June 2026

In working group 16 of ISO/TC 299, the robotics technical committee, a working draft on humanoid robot datasets splits robot data into a horizontal infrastructure covering lifecycle, provenance, quality, versioning and traceability, plus capability-specific modules for manipulation, locomotion and interaction. Its authors published their own account of the work in June, arguing that data standards are becoming foundational infrastructure for physical AI, and that the obstacle is non-cumulative data caused by high collection costs, data silos and inconsistent evaluation.

Read the four items together. Venture capital and a standards committee converged on the same object inside one quarter, apparently without coordinating. Provenance, versioning, traceability, reproducible evaluation. That is what a layer looks like while it is being born, and it deserves recording as such, because this site has spent three months saying no such layer was being built.

It was. Just not around production.

The gap

A production qualification does not travel

PPAP: 18 elements, five submission levels, re-run on a location change

Manufacturing has had its measurement layer for decades, and the shape of it is the problem. The Production Part Approval Process, from the Automotive Industry Action Group, is mandatory for tier-one automotive suppliers and has been carried into aerospace through AS9145, alongside AS9102 first article inspection. Eighteen elements, five submission levels, covering design records, process flow, control plans, measurement systems analysis, capability studies and a signed warrant.

It is triggered by new parts, by changes to materials or processes, and by changes in supplier or manufacturing location. That third trigger is the entire difficulty. A PPAP approves this part, made by this process, on this equipment, at this site. Move any of it and you run it again. The approval is not a property of the process. It is a property of the pairing of a process and a place.

P-10 predicts the opposite: validated production blueprints moving between certified factories, so that portable production becomes commercially real. That requires an approval which survives the move because both ends are certified, rather than one that is voided by it. Nothing in the new measurement layer touches this. A benchmark tells you what a robot can do anywhere. A qualification tells you that a part made here is accepted. Those are different objects, and a capability score does not become a production approval by getting more accurate.

The money

The referee keeps being built as a public good

A benefit corporation, an Apache-2.0 SDK, and a standards committee

Look at the legal forms rather than the products. Robocurve is a Public Benefit Corporation publishing an open-source framework. Hebbian's HFlow is Apache-2.0, with monetization described as managed workspaces and enterprise support for teams that would rather not run the runtime themselves. ISO is a standards body. These are the structures people choose when they do not expect to charge for the thing itself.

The precedent is close at hand. MLPerf began in 2018 and became MLCommons in 2020, a nonprofit, member-funded consortium of more than 125 organizations. The benchmark that settled how an entire industry talks about machine learning performance captured membership fees, not rent. VDA 5050 did the same for mixed-vendor mobile robot fleets: free, adopted, and worth nothing directly to the bodies that wrote it.

This is a counter-argument to P-18 arriving in an inconvenient form. P-18 predicts that the durable moat proved to be the framework, standards, data, certification and training, rather than the robots. The measurement half of that framework is now visibly being built, and it is being built in structures that do not monetize.

There is a reading that keeps P-18 intact, and it should be stated rather than assumed. The money in a certified network was never in publishing the standard. It is in operating the capacity the standard makes interchangeable, and in being the layer that routes work to it. Everything above is consistent with that reading. Nothing above is evidence for it.

The register

What this bears on, and what would change it

P-18, P-27, P-29, P-10

P-18 gets evidence on both of its halves, pointing opposite ways. The framework half is being built faster than this site expected. The moat half is weakened by the forms it is being built in. P-27, which asks for a certified process transferred between standardized cells owned by two different organizations by 30 June 2027 and carries the nearest open deadline here, gets nothing: a success detector that judges task completion is not a qualification crossing an ownership boundary, and the substrate question identified in August is unchanged. P-10 gets the PPAP location trigger as a clean statement of what it is actually up against.

P-29 is the interesting one. It predicts a routed order where the producing site was chosen partly on a living, quality-fed certification score rather than a static audit certificate. Continuous independent benchmarking as a service is the closest thing to that mechanism anyone has built, and it is currently pointed at robot policies instead of at sites. If it is ever pointed at sites, that is P-29's substrate. Nothing found this week says it will be. Nothing here is scored, and the register is untouched.

What would change the reading. An evaluation company scoring a site or a process rather than a policy. ISO 26264 or a successor extending from robot datasets to process records. A PPAP or first-article regime that accepts a transferred qualification without a full re-run. Or a physical-AI benchmark body that is straightforwardly for-profit and charges for scores, which would undercut the public-good reading directly.

Two limits, plainly. The batch figures disagree between compilations, at 235 companies and 23.0 percent industrials in one and 234 and 23.7 percent in another, and Demo Day is given as 10 September by Y Combinator itself and as 7 September elsewhere. The company-level finding does not depend on which is right. And this batch was read through directory listings and secondary compilations rather than by opening all 235 company pages, so the absence of a routing or certification entrant is a negative from a partial search, not an exhaustive one.

Sources

What this is built on

Y Combinator: 2026 Demo Day dates ·New Economies: Y Combinator Summer 2026 batch ·Foundevo: Y Combinator Summer 2026 Demo Day, complete startup list, founders, sectors and trends ·Y Combinator: Robocurve, evals for robots ·Y Combinator Launch: Robocurve, real-world evaluations of physical AI ·Robocurve ·Y Combinator: Instance ·Hacker News: Launch HN, Hebbian Robotics, build scalable robotics data pipelines ·GitHub: Hebbian-Robotics/hflow ·ISO: ISO/WD 26264-1, humanoid robot datasets, part 1, general requirements ·Liu et al.: Data Standards for Humanoid Robotics, The Missing Infrastructure for Physical AI ·Fictiv: PPAP, production part approval process guide for manufacturing ·Xometry: PPAP certification standards definition and audit requirements ·Lexco: understanding FAI, PPAP and AS9102, a guide for modern manufacturers ·MLCommons: about us ·VDA: VDA 5050

Company facts are taken from Y Combinator's own directory and launch pages and from each company's site, with Hebbian's launch date and score read from the Hacker News API rather than from the page. The ISO draft is cited through ISO's own catalogue entry and through a paper by members of the drafting working group, who are a primary source on their own draft and not an independent one on its merits. The PPAP location trigger is reported by two independent sources. The 5 September company-radar sweep referred to in the first section is internal Moduloa research and is not a source for any claim here.

Discussion

Add to this

Corrections, evidence, and disagreement are welcome. This is knowledge in the open. Anyone can read; sign in with GitHub only to post or react, and it appears here instantly.

Field notes are research, not decisions. They are dated, sourced, and open to correction. Republish this in full or in part under CC BY 4.0, with credit and a link back. · All journal entries →
Ask Datum