A MUJOCO MANIPULATION EXPERIMENT

Third Hand

An LLM plans.
A small tactile model controls.

Two hands must push a tray through a gravity-loaded door. The LLM plans how to use a block to hold the door open; a 0.52M-parameter model uses touch to adjust each hand every 10 milliseconds. Measured contact and motion return to the LLM for its next decision.

LLM task planning0.52M tactile parameters100 Hz local control8 N force cap per hand
01

Hold the door

02

Slide a block underneath

03

Unload and check support

04

Push with both hands

01 / MOTIVATION

Physical intelligence beyond vision.

This project asks whether a general-purpose LLM can use non-visual physical feedback—force, contact, and motion—to revise manipulation decisions. The system adopts a human-inspired rate separation: a multimodal high-level controller updates task-level intent at a lower rate, while a restricted tactile controller updates local hand motion at a higher rate.

QUESTION 01

Can an LLM act on contact?

We test whether structured tactile and force feedback changes hand roles, support decisions, and force limits during manipulation.

QUESTION 02

Can rate separation reduce repeated planning?

The LLM supplies subgoals; a small tactile MLP maintains local contact between LLM calls. This rate separation targets a faster local control loop.

Study scope. This study evaluates information flow and rate separation in one two-hand MuJoCo task.

02 / METHOD

One task, two control loops.

The LLM handles what to do next. The small model handles how to keep contact stable. They exchange a compact goal and measured outcome at different rates.

Two-rate control architecture The LLM sends a goal to a tactile MLP. The MLP sends incremental hand commands to MuJoCo. Raw touch returns to the MLP every 10 milliseconds, while summarized observations return to the LLM after each planning call. HIGH LEVELLLM plannerReads images, geometry,and summarized feedbackGOAL · FORCE · DURATION FAST LOCAL CONTROLTactile MLPTarget error + touch + state→ bounded correction ΔXYZ100 Hz · 0.2 mm MAX STEP ACTUATION + MEASUREMENTMuJoCo + hands500 Hz servo and physicsTouch · pose · motion8 N CAP / HAND goalΔXYZraw touch · 10 mssummary · each LLM call
FIGURE 01 The LLM sets intent; the small model closes the contact loop; the simulator produces the measurements used by both.

Observation summary Images, geometry, contact and motion measurements

Per LLM call
HIGH LEVELLLM

Plan a subtask

Choose the next hand role and movement.

READS

Two views + simulator state

RETURNS

Target XYZ · force cap · duration

Goalvalidated
workspace
FAST LOOPMLP · 100 Hz

Correct from touch

Convert the goal and the latest contact history into a tiny hand motion.

READS

4 tactile frames + 13 state values

RETURNS

ΔXYZ servo-target offset

ΔXYZ≤ 0.2 mm
per step
PHYSICS500 Hz

Act and measure

Servo the two contact pads, then measure touch, pose, and motion.

STEP

Every 2 ms under an 8 N cap

FEEDS BACK

Raw touch to MLP; summary to LLM

Immediate tactile loop Four recent frames and current hand state

Every 10 ms
THE SMALL MODEL

A goal-conditioned local policy.

The policy receives target error, touch, velocity, and hand position, then produces a bounded 3D servo offset while the LLM is thinking.

INPUT1,777 values

4 × 441 tactile values
+ target error and hand state

NETWORK1,777 → 256 → 256 → 3

ReLU hidden layers · tanh output
521,731 shared parameters

OUTPUTCorrection ΔXYZ every 10 ms

Clamped to 0.2 mm per step
same weights for left and right hands

Concrete example. For a target of 0.20 m and a hand position of 0.18 m, the network combines target error, touch, and motion to produce a bounded correction such as +0.2 mm; the servo target accumulates these corrections.

RATE SEPARATION

High-level planning and local control operate at different rates.

In B1, the local policy runs at 99.84 Hz with a 1.203–1.426 ms inference benchmark, whereas the first LLM response takes 25.06 s. These measurements characterize the local control loop. An end-to-end speed comparison requires a matched direct-LLM ΔXYZ baseline.

SYSTEMS ANALOGY

The rate split follows a human movement analogy: task-level intention is updated relatively slowly, while tactile and proprioceptive feedback support faster corrections during contact. We implement this analogy as an engineering hierarchy with multimodal LLM planning above a restricted tactile policy.

01

LLM sets a target
“Move behind the tray; use up to 7.5 N.”

02

MLP feels the contact
It sees recent touch, position, velocity, and target error.

03

Servo makes a small move
The next offset is applied; physics measures the result.

04

LLM plans again
Runtime sends a structured summary after the action interval.

Training and scope. The MLP is behavior-cloned from a local compliant teacher (180,900 training samples; 36,180 validation samples). Training covers local contact behavior; task-level sequencing remains with the LLM.

03 / EXPERIMENTS

Nine runs. All shown.

Three layout variants in each of A, B, and C. Every full-run replay lasts 15 seconds, including a 1.5-second hold on the final state. Eight runs pass the original criterion; A3 stops after a rejected target.

Rendered from saved states in their original time order. Sources, timing, and playback rates

AOne support block

Small changes to block position, block height, and tray friction.

A1

Original criterion passed

303.1 mm displacement · 19 LLM calls

15 s · Full-run replayOpen MP4Run data

A2

Original criterion passed

333.1 mm displacement · 20 LLM calls

15 s · Full-run replayOpen MP4Run data

A3

Command rejected

First target: z = 0.32 m, above the 0.30 m workspace limit.

15 s · Full-run replayOpen MP4Run data

BTwo candidate blocks

One block is too short; the other can support the door.

B1

Fully extracted

374.1 mm displacement · 21 LLM calls

15 s · Full-run replayOpen MP4Run data

B2

Original criterion passed

324.6 mm displacement · 22 LLM calls

15 s · Full-run replayOpen MP4Run data

B3

Original criterion passed

357.4 mm displacement · 22 LLM calls

15 s · Full-run replayOpen MP4Run data

CA shifted exit and a door lip

The exit moves sideways, with an added lip on the support side.

C1

Original criterion passed

335.7 mm displacement · 20 LLM calls

15 s · Full-run replayOpen MP4Run data

C2

Fully extracted

380.6 mm displacement · 19 LLM calls

15 s · Full-run replayOpen MP4Run data

C3

Original criterion passed

322.0 mm displacement · 22 LLM calls

15 s · Full-run replayOpen MP4Run data
Two completion criteria

8/9 pass the original criterion: at least 300 mm of tray travel plus all safety checks. 2/9 fully extract the tray: B1 and C2 exceed 365 mm, clearing the rear edge through the doorway, and pass safety checks.

04 / ANALYSIS

What contact feedback tells the planner.

Two B1 moments show the handoff: the MLP runs continuously; measured contact changes the next LLM command.

B1 · 177.6–224.4 s · About 4.5×Open MP4
01 / SUPPORT TRANSFER

Unload, then release.

Top tactile load drops 4.506 → 2.138 → 0 N while the door stays at 112.44 mm. The LLM then moves that hand to the tray.

Handoff: measured contact supports the next high-level action.

Measurement data
B1 · 312.0–478.2 s · About 12.3×Open MP4
02 / PUSHING RESISTANCE

One hand stalls.
Two hands move.

Displacement increments: 0.91 mm with one hand at 8 N, 0.07 mm with two hands at 6 N, then 9.27 mm at 7.5 N each.

One hand · 8 N0.91 mm
Two hands · 6 N each0.07 mm
Two hands · 7.5 N each9.27 mm

Handoff: motion and force feedback change hand allocation and force caps.

B1 action log

Interpretation

The LLM supplies task-level intent; the MLP supplies high-frequency contact control. The traces expose this interaction in two B1 sequences.

05 / RESULTS & EVIDENCE

Results and scope.

The original criterion also checks block support, hand withdrawal, clearance, collision, force peaks, and penetration. Full extraction adds a rear-edge position check.
RunOriginal criterionTravel (mm)Fully extractedPeak hand contact (N)API route
A1Passed303.1No7.505openai
A2Passed333.1No7.802openai
A3Rejected target0.0No5.576cctq_codex
B1Passed374.1Yes8.000openai
B2Passed324.6No7.897openai
B3Passed357.4No8.002cctq_codex
C1Passed335.7No8.008openai
C2Passed380.6Yes8.000openai
C3Passed322.0No8.000cctq_codex

The failure. A3 requests z = 0.32 m in its first command, above the executor’s 0.30 m limit. The action is rejected. The prompt leaves this bound implicit; the run remains a failure.

The scope. Nine layout variations across two API routes. An additional A-class full-extraction run is reported separately. The evaluation setting is simulated and task-specific.

Experiment snapshot: September 29, 2026. Linked raw records are preserved in their original language.