When a Humanoid Robot Fights, Who Is Actually in Control?
Next time a humanoid throws a punch, watch the pilot first.
Watch a humanoid fight and the first thing you'll notice is the robot. The second thing you should notice is the person controlling it.
A gamepad. A VR headset. A motion-capture suit. They can all produce a humanoid that looks like it's fighting on its own, but they draw the line between human input and machine execution in very different places.
Across public robot fighting events, the useful question is not simply whether a robot is tele-operated. It is how much freedom the human has, how much work the robot has to do, and what each side is responsible for.
Every interface makes a different trade-off between human freedom and machine reliability. Once you understand that trade-off, the differences between today's robot fighters become much easier to see.
1. Game controller: the human chooses the action
This is the simplest interface, and it is also the one most visibly established in public robot fighting. At Unitree's Hangzhou tournament on 25 May 2025, operators used remote controllers to maneuver and strike. UFB has also built a remote-piloted competition format around Unitree G1s, using game controllers, keyboards and Joy-Cons.
The principle is straightforward: the human does not directly command every joint. The controller sends a higher-level command, while the robot's onboard software handles the fast execution, balance and coordination underneath.
The advantage is reliability. There is little information to transmit, the operator can react quickly, and the robot can execute motions that have already been tuned or learned. It is also the easiest system to learn.
The trade-off is expressiveness. The pilot can choose among the actions and behaviors the system exposes, but cannot simply invent an arbitrary movement and expect the robot to reproduce it.In other words: the human chooses the action; the robot produces the movement.
2. VR / XR: the human controls space
VR and mixed-reality interfaces give the pilot a more spatial way to control the robot. Instead of remembering which button produces which movement, the operator can move a controller, hand or body through a virtual or mixed-reality space.
Reflex Arc is an good example, they describe its system as a hybrid that combines motion mirroring with pretrained, AI-assisted fight primitives. The pilot supplies movement and tactical intent, while software translates that input into something the robot can execute.
The advantage is intuition and expressiveness. A pilot can reason spatially and react through movement rather than through a fixed button vocabulary.
The cost is complexity. Human motion has to be tracked, interpreted and mapped to a robot with different proportions and dynamics. Latency also matters much more when the operator is trying to react to a moving opponent.
The headset therefore tells us how the command entered the system. It does not tell us how much of the final movement the human actually authored.
3. Motion capture: the human becomes the interface
Source: https://www.arcleague.com/
Motion capture pushes the idea further. Instead of asking the pilot to translate movement into buttons or controller gestures, the system captures the pilot's body and uses that movement as the input.
ARC League, presented by Katena, is one of the clearest public robot-fighting examples. Its materials describe pilots using full-body motion capture to drive humanoid avatars in real time, with strikes and tactical decisions produced live while dynamic balance remains on the robot.
The advantage is freedom. The pilot can use their whole body instead of a fixed action menu, which makes the interface much closer to actually fighting through the machine.
But the machine now has a harder problem. A human and a humanoid do not have identical proportions, joint limits or dynamics. Captured movement therefore has to be retargeted and made physically feasible for the robot. Tracking noise, latency and balance all become part of the control problem.
The result is a useful reversal: the more freedom you give the human, the more responsibility you push onto the robot.
The three approaches at a glance

The pattern is the important part: moving from a controller to VR to full-body capture gives the human more freedom, but it also pushes more responsibility onto the robot.
What about exoskeletons?
Interesting, but not a fighting interface yet.
Exoskeletons are worth mentioning because they push human-robot synchronization even further. Instead of observing the operator from the outside, an exoskeleton can measure joint movement directly and, in some research systems, provide force feedback.
HOMIE uses 7-DoF exoskeleton arms, sensor gloves and a foot pedal to control Unitree G1 and Fourier GR-1, while a learned policy handles walking, squatting and balance.
But this is where the evidence matters: exoskeleton control has not emerged as a practical public robot-fighting interface. The hardware is bulky, the setup is specialized, and the benefits are not yet strong enough to offset the deployment burden of a controller, headset or motion-capture suit.
For this article, exoskeletons are therefore a research direction, not one of the main control methods already shaping robot competition.
So which one wins?
This is the part we still don't know.
GAMEPAD → VR / MR → MOTION CAPTURE → MORE AUTONOMY
The game controller is simple, reliable and already works in competition. VR is more intuitive and spatial. Motion capture gives the pilot much more freedom, but makes the robot's control problem substantially harder.
Exoskeletons may eventually offer tighter human-robot correspondence, but today they look more like a research path than something that can scale across a competitive league.
The more interesting question is what users will actually prefer. Will people want to feel like they are directly driving a humanoid? Will they prefer an interface that makes the robot feel like an extension of their own body? Or will the best experience eventually be the one that asks the human to do the least, with the robot handling most of the physical execution?
There may not be one winner. Different applications may settle on different points along the spectrum.
Next time a humanoid throws a punch, watch the pilot first.
Then ask how much of that punch the robot finished on its own.