Can an LLM act on contact?
We test whether structured tactile and force feedback changes hand roles, support decisions, and force limits during manipulation.
Two hands must push a tray through a gravity-loaded door. The LLM plans how to use a block to hold the door open; a 0.52M-parameter model uses touch to adjust each hand every 10 milliseconds. Measured contact and motion return to the LLM for its next decision.
Hold the door
Slide a block underneath
Unload and check support
Push with both hands
This project asks whether a general-purpose LLM can use non-visual physical feedback—force, contact, and motion—to revise manipulation decisions. The system adopts a human-inspired rate separation: a multimodal high-level controller updates task-level intent at a lower rate, while a restricted tactile controller updates local hand motion at a higher rate.
We test whether structured tactile and force feedback changes hand roles, support decisions, and force limits during manipulation.
The LLM supplies subgoals; a small tactile MLP maintains local contact between LLM calls. This rate separation targets a faster local control loop.
Study scope. This study evaluates information flow and rate separation in one two-hand MuJoCo task.
The LLM handles what to do next. The small model handles how to keep contact stable. They exchange a compact goal and measured outcome at different rates.
Observation summary Images, geometry, contact and motion measurements
Per LLM callChoose the next hand role and movement.
Two views + simulator state
Target XYZ · force cap · duration
Convert the goal and the latest contact history into a tiny hand motion.
4 tactile frames + 13 state values
ΔXYZ servo-target offset
Servo the two contact pads, then measure touch, pose, and motion.
Every 2 ms under an 8 N cap
Raw touch to MLP; summary to LLM
Immediate tactile loop Four recent frames and current hand state
Every 10 msThe policy receives target error, touch, velocity, and hand position, then produces a bounded 3D servo offset while the LLM is thinking.
4 × 441 tactile values
+ target error and hand state
ReLU hidden layers · tanh output
521,731 shared parameters
Clamped to 0.2 mm per step
same weights for left and right hands
Concrete example. For a target of 0.20 m and a hand position of 0.18 m, the network combines target error, touch, and motion to produce a bounded correction such as +0.2 mm; the servo target accumulates these corrections.
In B1, the local policy runs at 99.84 Hz with a 1.203–1.426 ms inference benchmark, whereas the first LLM response takes 25.06 s. These measurements characterize the local control loop. An end-to-end speed comparison requires a matched direct-LLM ΔXYZ baseline.
The rate split follows a human movement analogy: task-level intention is updated relatively slowly, while tactile and proprioceptive feedback support faster corrections during contact. We implement this analogy as an engineering hierarchy with multimodal LLM planning above a restricted tactile policy.
LLM sets a target
“Move behind the tray; use up to 7.5 N.”
MLP feels the contact
It sees recent touch, position, velocity, and target error.
Servo makes a small move
The next offset is applied; physics measures the result.
LLM plans again
Runtime sends a structured summary after the action interval.
Training and scope. The MLP is behavior-cloned from a local compliant teacher (180,900 training samples; 36,180 validation samples). Training covers local contact behavior; task-level sequencing remains with the LLM.
Three layout variants in each of A, B, and C. Every full-run replay lasts 15 seconds, including a 1.5-second hold on the final state. Eight runs pass the original criterion; A3 stops after a rejected target.
Rendered from saved states in their original time order. Sources, timing, and playback rates
Small changes to block position, block height, and tray friction.
Playback failed. Please open the MP4 link below.
Playback failed. Please open the MP4 link below.
Playback failed. Please open the MP4 link below.
One block is too short; the other can support the door.
Playback failed. Please open the MP4 link below.
Playback failed. Please open the MP4 link below.
Playback failed. Please open the MP4 link below.
The exit moves sideways, with an added lip on the support side.
Playback failed. Please open the MP4 link below.
Playback failed. Please open the MP4 link below.
Playback failed. Please open the MP4 link below.
8/9 pass the original criterion: at least 300 mm of tray travel plus all safety checks. 2/9 fully extract the tray: B1 and C2 exceed 365 mm, clearing the rear edge through the doorway, and pass safety checks.
Two B1 moments show the handoff: the MLP runs continuously; measured contact changes the next LLM command.
Playback failed. Please open the MP4 link below.
Top tactile load drops 4.506 → 2.138 → 0 N while the door stays at 112.44 mm. The LLM then moves that hand to the tray.
Handoff: measured contact supports the next high-level action.
Measurement dataPlayback failed. Please open the MP4 link below.
Displacement increments: 0.91 mm with one hand at 8 N, 0.07 mm with two hands at 6 N, then 9.27 mm at 7.5 N each.
Handoff: motion and force feedback change hand allocation and force caps.
B1 action logThe LLM supplies task-level intent; the MLP supplies high-frequency contact control. The traces expose this interaction in two B1 sequences.
| Run | Original criterion | Travel (mm) | Fully extracted | Peak hand contact (N) | API route |
|---|---|---|---|---|---|
| A1 | Passed | 303.1 | No | 7.505 | openai |
| A2 | Passed | 333.1 | No | 7.802 | openai |
| A3 | Rejected target | 0.0 | No | 5.576 | cctq_codex |
| B1 | Passed | 374.1 | Yes | 8.000 | openai |
| B2 | Passed | 324.6 | No | 7.897 | openai |
| B3 | Passed | 357.4 | No | 8.002 | cctq_codex |
| C1 | Passed | 335.7 | No | 8.008 | openai |
| C2 | Passed | 380.6 | Yes | 8.000 | openai |
| C3 | Passed | 322.0 | No | 8.000 | cctq_codex |
The failure. A3 requests z = 0.32 m in its first command, above the executor’s 0.30 m limit. The action is rejected. The prompt leaves this bound implicit; the run remains a failure.
The scope. Nine layout variations across two API routes. An additional A-class full-extraction run is reported separately. The evaluation setting is simulated and task-specific.
Experiment snapshot: September 29, 2026. Linked raw records are preserved in their original language.