Problem: Teleop Data Contains Cultural Knowledge

Teleoperated robot data will contain its operators' specific cultural knowledge about the physical world. If Mind Robotics isn't careful, that knowledge will be replicated in its models and weaken their utility as physical AI.

Everything we know about the physical world is colored by culture. When we need to pick up a pencil, we don't think about the pencil, its weight, or its shape in isolation. We think about how we'd hold the pencil, what we'd write with it, and how we'd need to put it down. The sociocultural function of things affects how we interact with them.

This will show up in teleoperator data, especially when operators are immersed in . A model trained on this data will replicate human cultural knowledge, which doesn't have a physical equivalent—so Mind Robotics' Physical AI will have knowledge it shouldn't have, biasing its knowledge of the physical world.

To illustrate this, let's imagine we need to install battery cells in a 2x4 grid. We're using a 7-axis arm like the Kuka iiwa, which can install one cell at a time.

setup

Most English-educated teleoperators are going to control the arm to install these cells left to right, row by row, just like they would write a paragraph.

setup

Models trained on teleoperator data—even if it's 10% of the total data—will carry over an association between cells in the same row, ordered left to right. This isn't a problem when doing the same task, but extrapolating to other tasks will create problems.

Now, we have an 8x8 array of bolt holes. We only need to install bolts in some of them.

Bolts

When it's trained on some teleoperator data, our model will be more likely to torque down bolts left-to-right, row-by-row, using way more travel than necessary:

Annotated

This is a pretty trivial problem. It'll add a fraction of a second to a 10-second task, and it can be easily fixed by using mostly RL training data and multiple embodiments (following OXE). But the cultural knowledge remains, and in more complex tasks, it will create unexpected problems. Things like:

  • Gaits (i.e. heel striking) will show up in locomotion tasks, even though a bot might move better with a different gait.
  • People are used to holding soldering irons delicately to avoid getting burned, even though the robots they're controlling can't get burned as easily.

This means that Mind's models will have culturally specific knowledge, reducing their utility as general physical intelligence. Mind needs to mitigate cultural knowledge in its teleoperated data to make better models.


Solution: Parallel Operator Interfaces

Multiple embodiments make VLAs and world models better. Why not try the same thing on the other side—collect data with multiple teleoperator interfaces in parallel?

Each of these will contain different aspects of an operator's spatial knowledge. For example:

  • A VR headset with a top-down view and a single controller immerses an operator in the robot, making them more likely to do the left-to-right row-by-row thing.
  • Kinesthetic teaching will make travel more difficult, forcing an operator to look for the most efficient way for the robot arm to install the cells, not their own arm. It'll rely more on problem solving than on cultural knowledge.

Teleoperators understand the installation task in different ways depending on the interface they use, so using more interfaces will reduce the impact of cultural knowledge. Some other interfaces you can use:

ALOHA
Leader-follower bimanual ALOHA setup

By training a model on teleoperator data from many of these interfaces, Mind will capture more of an operator's physical knowledge, building a more accurate representation of physical reality.