On my way back from Pittsburgh last week, the connecting flight landed late in Chicago. With 10 minutes before my SFO-bound flight’s boarding closes, I had to sprint past about 20 gates. As I ran, I quietly thanked my actuators. Wait, my muscles and joints.
My sensors were working just as hard. My eyes scanned for the gate sign and gaps in the crowd. My body adjusted to the carry-on and backpack. My fingers kept hold of the lukewarm coffee I’d bought in Pittsburgh hours earlier. My ears listened for the final call.
If I were a humanoid, making that tight connection while keeping hold of my coffee would require all these sensor feedback working together. That’s the topic of today’s post - the senses of a robot. In autonomous vehicles, Waymo combines cameras, lidar and radar, while Tesla has built its driving system around camera-based vision. Sensor choices change what the hardware measures and what software has to infer.[1]
Recently, I asked a Formic deployment engineer what broke most often. His answer was sensors. Jake Panikulam of Main Street Autonomy, a startup that focuses on sensors calibration, shared with me that dirt on sensors as a source of failures.[2]
Let’s go through the six senses, their suppliers, and where deployment breaks. Then we’ll ask what technical progress would make robots more useful, and where specialists can compete with established suppliers.
In this piece, we cover:
What sensors do humanoid robots use
Sight: what’s around me?
Depth: how far away is it?
Balance and body position: where are my limbs, and how fast am I going?
Force: how much load am I carrying?
Touch: is the cup slipping?
Hearing: is the boarding gate closing?
Where deployment breaks
What technical progress would make robots more useful?
Will automotive suppliers lead, or can specialists win?
What sensors do humanoid robots use?

Sight: what’s around me?
How it works: A video camera captures images many times a second. 30 fps (frames per second) means 30 images each second. During a vision-guided task, the camera supplies a live stream of frames. Software then uses those images to recognize the gate sign, the cup or the person walking into the robot’s path.
What it measures: Light at each pixel, which software interprets as objects and scenes.
Where it’s located: Head cameras give a view of the surroundings. Wrist cameras can give a closer view of a grasp, although the hand or object can still block them.
Suppliers: The supply chain spans image-sensor chips, camera modules and integrated perception systems.[1]

Depth: how far away is it?
What it measures: Distance to surfaces across the scene.
How it works: Three common depth-camera approaches are:[2]
Stereo compares two views taken a known distance apart. Nearby objects shift more between the images than distant ones.
Structured light projects a known pattern and uses where it appears in the camera image to calculate distance.
Time of flight sends out light and measures its return timing, directly or through a change in phase, to estimate distance.
Many lidars use time of flight, the third approach above: they send out laser pulses and time the reflections. A spinning lidar sweeps its view around the robot, whereas a MEMS-scanned lidar steers light with a tiny moving mirror inside the unit. Scanning decides where it looks. Ranging determines how far away something is.[3]
Where it’s located: Depth cameras can sit on the head, torso or wrist, depending on the view needed. Unitree’s G1 combines a depth camera with 3D lidar.[4]
Suppliers: Hesai and RoboSense supply lidar. Orbbec and RealSense supply depth cameras, with some suppliers combining several sensing functions in one module.

Balance and body position: where are my limbs? and how fast am I going?
Proprioception is body awareness. If you know whether your elbow is bent with your eyes closed, you have proprioception.
How it works: Encoders read an optical or magnetic pattern as a joint rotates. Then, software takes joint angles, combined with the robot’s limb lengths, to calculate where its hands and feet sit relative to its body.[5]
An inertial measurement unit, or IMU, adds acceleration and rotation-rate measurements, with accelerometer, gyroscope.[6]
What it measures: Encoders report joint position, while IMUs report acceleration and rotation rate. Software combines IMU, joint and contact measurements to estimate motion and orientation, then the controller adjusts the motors to keep the robot upright.
Where it’s located: A body IMU can be mounted in the torso or pelvis, while encoders sit at the joints. The exact mounting depends on the robot.[5-1]
Suppliers: Bosch and ST supply IMU chips. VectorNav supplies inertial modules. Renishaw/RLS and Kübler supply encoders.
Force: how much load am I carrying?
How it works: Imagine a tiny metal spring inside the wrist. When the robot lifts or pushes something, the metal bends a little. Small electrical sensors attached to it, called strain gauges, stretch with it. Stretching changes how easily electricity flows through them. After testing with known loads, the robot can translate that change into how hard it is being pushed, pulled or twisted.[7]
What it measures: Push and pull are forces. Torque is their turning effect around an axis. A six-axis sensor measures pushes and pulls along three directions, plus twisting around those three directions. Force is measured in newtons. Torque is measured in newton-meters.
Where it’s located: At the wrist, it senses the combined load passing between the hand and arm. Ankle sensors can measure loads from the ground. Joint torque sensors are built into the joint or actuator, for example at the elbow or knee, to measure the twisting load transmitted through it.
For the coffee cup, wrist force sensing helps track the overall load, while fingertip touch tells me whether my cup is slipping.
Suppliers: ATI, Bota Systems and SRI.

Touch: is the cup slipping?
How it works: Tactile sensors turn contact at a surface into a signal. An optical design uses an internal camera to watch a soft material deform. A magnetic design detects how embedded magnets move as that material is pressed or sheared. Other designs use changes in resistance or capacitance.[8]
What it measures: A taxel is a sensing site, analogous to a pixel. An array can map contact across a fingertip or palm. What it reports depends on the design. Some measure local force, others capture deformation that software interprets. Temperature requires a temperature-capable sensor.
Where it’s located: Besides fingertips and the palm, touch can extend to feet, arms and the torso. That creates coverage beyond the camera’s view, but also more surfaces, wires and connections to maintain.
Suppliers: GelSight, XELA, PaXini, Daimon and Tekscan use different sensing approaches.

Why do tactile sensors need replacement?
Soft material deforming is part of how these sensors work. Wear happens when abrasion, tears or changes in material response alter the signal. Replacing the skin or coating may require calibration or model checks. Magnetic interference can also disturb readings, but is a separate issue from physical wear. Service life depends on the material, load and task.[9]
Hearing: Is the boarding gate closing?
How it works: Microphones capture sound. Software interprets voices or alarms. A robot can use a wake word like an Amazon Echo, a button, or continuous command recognition.[10]
What it measures: An audio signal.
Where it’s located: The microphone is usually on the head or torso, with motor noise and vibration influencing the design.[10-1]
Suppliers: Infineon and ST supply MEMS microphones.
Where deployment breaks
In deployment, measurements must line up in space and time, reach the controller quickly enough, and remain usable after repeated use and servicing. A missed task can come from the sensor, its calibration, the processing pipeline or the robot’s response.
The airport trip makes three failure modes concrete:
Occlusion and timing: Someone steps into its path. The robot needs a usable view to see the travelers around it, an accurate distance estimate and enough time to change course. A blocked camera or delayed measurement could leave it reacting too late. It needs to detect and avoid the person at its walking speed. A high camera frame rate helps refresh the scene, but processing and motor response also take time.
Grip feedback: The cup starts slipping. As the robot turns around a suitcase, the cup shifts in its hand. Its wrist sensor measures the overall load, while fingertip sensors help detect the slipping contact. The controller needs to adjust the grip before the cup falls, without squeezing hard enough to crush it. The robot needs to complete the trip without dropping, spilling or falling over.
Contamination and wear: Performance degrades during use. A smudged lens could make the gate sign harder to read, much like dirty glasses. Repeated use can also wear tactile surfaces or change calibration. Waymo’s built-in sensor cleaning offers a relevant lesson: keeping sensors usable is part of deployment.[11]
What technical progress would make robots more useful?
When it comes to sensors, it’s about getting “good enough” so it’s no longer a blocker for capabilities. Cameras, IMUs and force sensors are established product families from other industries, but their specifications still need to meet the robot’s operating conditions. [12] Below is a map of the gaps in the sensors offering today for robotics capabilities.




