ISF Voices 2026: The Robotics Data Gap: How Training Data Pipelines Will Shape the Global AI Race
The next edition of ISF Voices.
In this special edition of SCSP’s newsletter, we continue our ISF Voices series, showcasing writing by alumni and fellows in SCSP’s International Strategy Forum (ISF) program. Each piece reflects the unique vantage points of emerging leaders from around the world working to shape the future of geopolitics, technology, and democracy. Today’s piece is by Phillip An, a 2026 ISF Fellow and the Co-Founder and CEO of Allston Labs, a laboratory specializing in AI and robotics data infrastructure. In this article, Phillip discusses the critical importance of robotics data in shaping the future of physical AI, and how America can catch up in this competition.
🎙️New Episode of Memos to the President
On the margins of the India-US Forum in New Delhi, Ylli sat down for a new episode of Memos to the President with S. Krishnan, Secretary of India's Ministry of Electronics and Information Technology (MeitY), alongside SCSP Senior Advisor Dr. Sameer Lalwani.
Secretary Krishnan goes in depth on India's AI vision, the $10 billion semiconductor strategy, and how India is positioning itself as a trusted tech partner for the Global South.
Available wherever you get your podcasts.
The views expressed in this article are solely those of the author and do not necessarily represent the views of the SCSP, the ISF, or any affiliated entities.
The Robotics Data Gap: How Training Data Pipelines Will Shape the Global AI Race
The Invisible Bottleneck
Today, at a facility on the outskirts of Beijing, a young woman wearing a VR headset and an upper-body exoskeleton is folding a pair of shorts. She will repeat this motion hundreds more times today. Every fold is recorded by joint torque sensors, multimodal cameras, and LiDAR arrays — every wrist rotation, every finger’s pressure, every adjustment of the fabric captured as structured data. Tomorrow, a humanoid robot will attempt to replicate her movements. She is one of hundreds of what Chinese media have dubbed “cyber laborers” — a new class of infrastructure that the rest of the world is only beginning to understand.
The global race to build autonomous robots has captured the attention of policymakers and investors across major economies. Governments debate chip export controls, think tanks analyze hardware supply chains, and venture capitalists pour billions into humanoid startups; yet the bottleneck that will most determine which nations lead in this technology receives almost no policy attention: the physical AI1 training data pipeline.
The disconnect is fundamental. Language models are trained on the internet — trillions of words at near-zero marginal cost. However, a robot cannot learn to fold clothes from Wikipedia. It requires teleoperated demonstrations, sensor recordings, and manipulation sequences — all physically collected through human labor on hardware rigs costing $50,000 to $150,000 each, and yielding fewer than 200 demonstrations per worker per day.
The global AI data annotation market, valued at $4.89 billion in 2025, is projected to reach $17.1 billion by 2030 — with robotics emerging as its fastest-growing segment. Micro1 has contracted thousands of workers across more than 50 countries, paying $15 per hour to strap iPhones to their heads and film household chores; DoorDash has launched a service called “Tasks” leveraging its couriers to record robot training data; Scale AI has logged more than 100,000 hours of in-house collection in its San Francisco labs.
The deeper insight missing from current policy discussions is that teleoperation facilities are not merely cost centers — they are infrastructure. Just as TSMC’s plants became indispensable bottlenecks in the semiconductor supply chain, dedicated robotic data collection facilities will become critical bottlenecks in the physical AI supply chain. Whoever establishes the trusted, well-managed pipelines will hold leverage comparable to a semiconductor fab. The question is not whether this infrastructure will be built — it is where, by whom, and under what standards.
Two natural questions arise. First, why not synthetic data, given projections it could comprise the majority of AI training data by 2030? Synthetic generation does scale demonstrations aggressively — Nvidia’s data factories augment a single captured motion into millions of variants — but it amplifies a real corpus; it cannot originate one. You cannot simulate contact dynamics, friction coefficients, or sensor noise patterns you have never observed. Every synthetic variant traces back to a real seed, which is why Nvidia identified demonstration data as the critical bottleneck for physical AI at GTC 2026 even while projecting heavy synthetic augmentation downstream. Second, why not use the iPhones strapped to gig workers’ heads? Phones record what a fold looks like, but cannot capture the joint torque, force feedback, depth, or exoskeleton kinematics that contact-rich manipulation depends on. A model trained on smartphone video learns the appearance of a successful fold; it cannot learn the proprioceptive signal that distinguishes a secure grip from a slip. For visually loose tasks, lower-cost capture is plausible; for the dexterous manipulation that defines real-world robotic value — household chores, eldercare, assembly — the sensor stack is non-negotiable.
This is the “robot data gap”: whoever can build the infrastructure to collect, standardize, and manage this data at scale will hold the key to the next era of industrial automation. Currently, one country is clearly building it fastest.
China’s National Robotics Data Strategy
China recognized the robot data gap as a strategic priority earlier and more comprehensively than any other country, offering both a model worth studying and a demonstration of the national ambition required for data supremacy in a critical emerging field.
While at Tsinghua University, I witnessed this transformation firsthand. In a fall 2025 seminar, a professor advising the Ministry of Industry and Information Technology’s (MIIT) committee on humanoid robot standardization outlined the government’s data collection roadmap: “embodied intelligence” — a term that appeared almost nowhere in Chinese policy documents prior to 2023 — has become a top research priority and a leading hiring category. Recruitment posters for “cyber laborer” positions now appear alongside traditional engineering roles on Tsinghua job boards.
In December 2025, MIIT appointed a 65-member standardization committee for humanoid robots and embodied intelligence — including the founders of Unitree and Zhiyuan Robotics (智元机器人), Huawei executives, and Tsinghua scholars. By March 2026, MIIT released its first national standard system, covering the entire lifecycle from data acquisition formats to safety protocols.
The 15th Five-Year Plan (2026–2030) elevated the field further, designating robotics one of eight “strategic emerging industries” and dedicating a section to embodied intelligence among the top ten “new industry tracks” — triggering mandatory coordination across central ministries and state financial institutions. Beyond $20 billion in direct subsidies, an NDRC guidance fund plans to inject $137 billion into AI and robotics over the next two decades.
The infrastructure surrounding Beijing offers a concrete lesson for any nation serious about physical AI. In 2025, more than 40 government-backed robot training centers were established across China, with approximately 20 already fully operational. The Beijing Humanoid Robot Data Training Center spans over 10,000 square meters, employs hundreds of workers, and covers 16 task categories from automotive assembly to elder care. In Zigong, Sichuan, a 6,000-square-meter facility opened in January 2026 is capable of generating three million high-quality data entries annually — comparable to the entire scale of the Open X-Embodiment dataset, which aggregated the work of 22 institutional contributors.2
Beyond standardization committees and industrial parks, Chinese firms are converging on a vertically integrated robotics stack, from sensors and rigs to data and foundation models.3 At the May 2026 JP Morgan China Summit, the depth of talent and capital behind Chinese physical AI was on open display.4
What makes this distinct is not the scale of labor but the degree of state coordination behind it. The PRC has done what no other market would: pre-financed and standardized the physical infrastructure of an entire training pipeline before downstream demand justified the investment.
America's Missing Strategy
Despite world-class robotics research institutions and a vibrant startup ecosystem, the United States lacks a national robotics data strategy. As the Special Competitive Studies Project’s (SCSP) National Robotics Strategy report5 lays out, China has spent the better part of a decade executing a coordinated national robotics strategy — accelerating adoption, scaling manufacturing, and securing its supply chain — while America has yet to mount a comparable effort.
The investment gap reflects this. The ARM Institute — America’s flagship public-private robotics accelerator — operates on a $30 million budget. In 2025, even Tesla — among the most prolific U.S. humanoid producers — is estimated to have shipped only 150 humanoid robots, while Chinese firms took the global lead. The gap reflects coordination and public investment, not lack of American talent or technology.
Washington is beginning to take notice. In June 2026, Senators John Hickenlooper, Dave McCormick, Martin Heinrich, and Todd Young introduced bipartisan legislation to establish a National Commission on Robotics tasked with developing recommendations to strengthen U.S. competitiveness, supply chains, workforce development, and national strategy in robotics. The American Security Robotics Act, introduced in March 2026, would restrict federal procurement of foreign-manufactured humanoid robots; the White House is reportedly drafting an executive order on robotics; Commerce Secretary Howard Lutnick has convened meetings with the CEOs of Apptronik, Boston Dynamics, and Tesla. These initiatives are largely reactive — focused on strategy development and procurement restrictions rather than on the proactive data infrastructure U.S. robotics companies need. The critical question is not what to block, but what to build.
Meanwhile, China is actively shaping international standards, leading the International Electrotechnical Commission’s (IEC) standard-setting efforts, such as for elder-care robots, and establishing domestic norms that could become the global default. This is what any leading producer would do. But it means other nations — including the United States — must proactively engage, or accept standards they played no part in designing.
A Three-Pillar Policy Framework
The robotics data gap requires a coordinated response. I propose a three-pillar framework addressing strategy, labor, and diplomacy.
Pillar One: A National Robotics Data Strategy
The United States should establish a coordinated national strategy for robotics training data — not by mimicking China’s centralized approach, but by drawing on the federal public-private partnerships that have historically driven American dynamism.
At the federal level: establish a National Robotics Data Initiative within the Department of Commerce to coordinate data collection standards across the Department of War, Department of Health and Human Services, and The National Institute of Standards and Technology; fund shared robotics data infrastructure (teleoperation facilities, sensor suites, annotation tools) through competitive grants; establish common data formats and interoperability standards to prevent fragmentation from proprietary pipelines; and launch a robotics data accelerator (modeled on the DARPA Challenges) matching funds for companies building scalable teleoperation operations.
At the industry level: form a robotics data consortium (modeled on the Semiconductor Research Corporation) where rival firms jointly invest in pre-competitive data infrastructure while retaining proprietary advantages in model training; invest in dedicated teleoperation facilities staffed by a specialized workforce; and adopt and expand open standards such as Nvidia’s Isaac Teleop and the Open X-Embodiment framework.
Pillar Two: International Labor Standards for Teleoperators
The emerging teleoperation workforce requires protections existing labor laws fail to provide — not as an ethical add-on, but because labor standards and data quality reinforce one another: better-trained workers in structured environments produce better data. This can also distinguish an American approach to teleoperated robotics data from the PRC’s approach.
At the federal level: establish a new occupational classification of “Data Demonstration Workers,” recognizing their role in generating training data for physical intelligence; institute minimum standards for wage transparency, including disclosure of how collected data will be used and its expected commercial value; and mandate that platforms recruiting international teleoperators meet labor standards of both the worker’s country of residence and the country where the commercial application is deployed.
At the international level: the International Labour Organization (ILO) should convene a working group on AI data labor and expand its framework for digital labor platforms to cover demonstration work; multilateral development banks should condition loans to the robotics industry on adherence to minimum standards for data workers.
Pillar Three: International Data-Sharing Agreements
No single country — including China — can independently gather the full diversity of data required to train robots; the diverse environments those robots will operate in demand demonstration data drawn from all of them. Workable precedents exist: ISO and IEC robotics safety standards (ISO 10218 for industrial robots, ISO 13482 for personal care robots) have for years brought American, European, Japanese, and Chinese engineers to the same technical tables, and the Bletchley (2023) and Seoul (2024) AI Safety Summits demonstrated that even strategic competitors can share evaluation methodologies on narrowly bounded technical questions.
At the government level: negotiate mutual recognition agreements for robot data standards, ensuring demonstration data collected in one country meets quality requirements for use in another; and pursue narrowly scoped data-sharing arrangements with China on robot safety, eldercare, and disaster-response applications. The IEC’s existing elder-care robot working group — which China leads but in which U.S. and European representatives also participate — is a working example of how technical cooperation can coexist with strategic competition.
At the standards-body level: actively participate in IEC and ISO data governance processes, collaborating rather than competing with China and other leading nations to develop interoperable global frameworks; and develop data provenance standards to establish a trusted supply chain for robot training data — analogous to the “Trusted Foundry” concept in semiconductors.
The Stakes
The robot data gap is not a distant policy issue — it is a present reality. China’s 15th Five-Year Plan projects that the market for embodied intelligence will reach 400 billion RMB ($56.5 billion) by 2030 and 1 trillion RMB by 2035. The training data infrastructure being built today — in China, and eventually elsewhere — will accumulate over time, creating a first-mover advantage increasingly difficult to replicate.
The first phase of the AI revolution was built on internet-scale text data and world-class research institutions. The second phase — physical AI — will be determined by who can master the collection, standardization, and governance of embodied data to deploy physical robotics platforms at scale.
The robotics data divide can be bridged. But only if policymakers recognize that the most impactful supply chain of the AI era is not silicon but the human demonstrations that teach machines how to move.
Nvidia and most U.S. industry literature use “physical AI”; Chinese policy documents use the parallel term “embodied intelligence” (具身智能). I use “physical AI” throughout for consistency, preserving “embodied intelligence” only when directly translating Chinese policy.
Open X-Embodiment is the largest collaborative open-source robotics dataset, aggregating demonstration data from 22 institutional contributors worldwide.
In contrast, American firms are running separate proprietary pipelines walled off from each other.
Among notable Chinese physical-AI players: AgileX Robotics' Cobot Magic teleoperation platform has become standard kit in many Chinese training centers; Tsinghua-affiliated Lumos Robotics, founded in late 2024 in Shenzhen, has closed pre-A rounds for a tactile-first humanoid platform built around its proprietary FastUMI data collection system; embodied-data specialists such as GenRobot AI and Datatang (数据堂) are productizing teleoperation pipelines as managed services rather than treating them as research artifacts; AgiBot's GO-1 foundation model is trained on data flowing through these facilities.
SCSP has also launched a National Security Commission on Robotics for Advanced Manufacturing—co-chaired by Senators Ted Budd (R-NC) and Elissa Slotkin (D-MI)—sending a clear signal that robotics is being elevated to the same strategic tier as semiconductors and critical minerals.



