Skip to main content
Robotics Physical AI SO-101 Robot Learning

Why We’re Building a Robot Arm: A First Step Toward Adaptable Robots

Ashar Mirza
Toward Adaptable Robots — VoicePing’s robotics research. A human demonstrates a placement, the system interprets the goal, and a robot arm executes an existing skill.
Demonstrate the task, interpret its goal, and execute an existing skill—the starting point for our adaptation research.
In this article

VoicePing’s goal is to build intelligent robots that do useful work and adapt as that work changes. Knowledge gained from one task should help with the next, reducing the data, retraining and engineering needed for each new use case. We are starting with the SO-101 arm to investigate that goal across perception, planning and control.

Our evidence through September 18, 2026 includes a human demonstration interpreted and executed in simulation, plus a separate simulated sequence with two placements. Physical commissioning, successful transfer to the arm and broader skill learning remain unverified.

What should a robot be able to adapt to?

Imagine a workbench assistant sorting parts into trays. Moving a tray requires updated coordinates. Requesting familiar placements in a different order requires composing existing skills. Introducing a new shape or material may require a different grasp and more experience. The challenge is to recognize which kind of change has occurred and reuse what still works.

Three adaptation questions: locate a moved cube and tray while reusing the same skill; compose familiar placements in a new order; investigate a new grasp when shape or material changes.
A changed scene, sequence or object calls for a different kind of adaptation.

Open X-Embodiment shows that pooling experience across robot platforms can improve manipulation policies. It motivates our investigation of reuse; our own gains still need measurement. For deployment, the practical question is how much setup, new data and human help each change requires.

Why start with the SO-101?

The SO-101’s open mechanical design and LeRobot support make it an accessible platform for studying the complete manipulation loop. Even tabletop placement requires identifying an object, locating it, making contact and checking the outcome. We can vary one condition at a time and inspect what fails.

Operating the arm also exposes constraints that model evaluation alone misses: camera placement, mechanical behavior and the quality of recorded experiments. This develops reusable software and practical judgment before we choose larger platforms. Future hardware should follow the use case’s reach, payload, environment and reliability requirements.

What leading robotics platforms teach us

These examples address different layers of robotics, from commercial integration to research and factory trials. Primary sources were checked on September 24, 2026; the comparison is not a market-share or performance ranking.

Integration

Universal Robots

AI Accelerator

Available to order
3D vision
NVIDIA Jetson
PolyScope X

Evaluate the complete workflow

Setup · task success · fault diagnosis

Sensing & control

Franka Research 3

Research arm

Research hardware
Joint torque Sensing
1 kHz Low-level control

Match the hardware to the question

Precise force research may need another arm.

Robot models

Google DeepMind

Gemini Robotics 2 / ER 2

Preview models
ER 2 · Reasoning Public preview
Robotics 2 · Actions Private preview

ER 2 access: Google AI Studio / Gemini API

Test reasoning and action separately

Then validate their integration on hardware.

Deployment

Figure at BMW

Figure 03 + Helix 02

Production-site project
Figure 02 Body-shop pilot
Figure 03 Logistics sequencing

Helix 02: whole-body control, per Figure

Prove a complete operational task

A factory project has a defined scope.

Our starting pointVoicePing · SO-101
  • Ground goals
  • Reuse skills
  • Measure adaptation

How the whole system fits together

A demonstration or instruction supplies a goal. Perception locates the relevant objects; planning selects and sequences available skills; control produces movement. Fresh observations guide execution, while outcome checks determine whether to continue, retry or ask for help. Separating these responsibilities lets us diagnose and improve each component.

Simulation baseline: a human demonstration becomes a structured placement goal; unsupported or ambiguous tasks stop. RGB and depth locate the current objects, a programmed pick-lift-place skill uses fresh observations, and verification either advances to the next action or stops and records a failure.
The simulation baseline connects a demonstrated goal to an existing skill, with observations and outcome checks.

Our current system interprets a short human video as a supported pick-and-place goal and invokes a programmed geometric motor skill. The clip does not train a new manipulation policy. That baseline gives us a reference for evaluating learned perception, policies and task composition. Reliable execution from live perception remains open work.

Simulation and hardware should improve each other

Simulation lets us repeat tasks, vary conditions and compare methods before physical trials. Hardware then reveals what the simulator missed: lighting and occlusion, camera calibration, contact behavior or response delays. Measurements should feed back into the simulator, controller or training process.

Simulation varies conditions and tests candidates. Planned physical trials calibrate the camera and arm and measure behavior. Candidates move toward hardware; measurements return to simulation. Appearance, calibration, contact and timing are distinct transfer gaps.
Test candidates in simulation; use bounded hardware trials to measure the remaining gaps.

Domain randomization is one relevant approach: varying simulated appearances can reduce dependence on a single visual setting. The original study demonstrated transfer for object localization, without establishing that arbitrary robot behavior transfers automatically. We need a calibrated physical trial and a comparison of expected versus observed behavior to assess our own system.

How we will measure progress

Task success and the robot’s judgment of success are separate measurements. In our simulation, an object can reach the correct place while visual confirmation fails. An independent physics check helps distinguish that verification error from a failed grasp or placement; the controller does not receive the evaluator’s ground-truth state.

Four illustrative outcomes compare the actual placement with camera verification: correctly verified placement, false confirmation of a failed placement, successful placement obscured from view, and an action failure without confirmation.
Check what happened separately from what the robot believes happened.

Our proposed comparison keeps the task family and held-out layouts fixed, changes one component—perception, planning or the motor skill—and retains the programmed baseline. We would measure:

  • Complete-task success, including whether earlier placements stay correct.
  • Retries and human assistance, including failures the system cannot resolve.
  • Adaptation effort: extra demonstrations and tuning needed for each variation.

This makes an improvement useful only when its benefit, cost and failure conditions are visible.

What comes next

The next milestone is a bounded task performed and measured on the physical arm, followed by tests under changed conditions. Those results will guide which skills, interfaces and evaluation methods we carry into a broader task family and future robots.

For VoicePing, this is how a small research platform supports a larger ambition: building the capability to turn robot intelligence into dependable work. If these questions interest you, explore opportunities at VoicePing and our research articles .

Share this article

Help turn robotics research into useful systems

Interested in learning, perception, control or the systems that connect them? Explore opportunities at VoicePing and our technical work.

Video
0:00 0:00