All writing

G1 Bottle Pick-and-Lift

G1 Bottle Pick-and-Lift: Object-Relative Corrections and Real-Robot Integration

Week 3: improving integrated success with object-relative control and moving the learned policy toward real G1 deployment.

EnglishPhysical AIRoboticsReinforcement LearningIsaac LabSim-to-Real

Introduction

In the previous article, I covered the following parts of the bottle grasping and lifting reinforcement learning project for the Unitree G1 Inspire Hand.

  • Mass Curriculum from 0.08 kg to 0.62 kg
  • Stage 2 grasp / lift policy adapted to 0.62 kg conditions
  • Stage 1 Safe Approach Policy to approach without moving the dynamic bottle
  • controlled ingress connecting Stage 1 and Stage 2
  • Phase timing issue due to Isaac Lab's decimation
  • Confirmation of checkpoint architecture and observation contract
  • Termination-based formal deterministic evaluation

The main simulation results up until now were as follows.

ModuleConditionResult
Stage 1Dynamic 0.62 kg, x/y ±1.5 cm768 / 768
Stage 2Fixed nominal, 0.62 kg766 / 768
Zero-residual bridgeFixed nominalReaching Stage 2 handover

On the other hand, at the time of the previous article, the integrated evaluation of operating Stage 2 learned residual from the actual Stage 1 handover state had not been completed.

This week, after completing this integration part, I proceeded with failure analysis of randomized workspaces, object-relative trajectory correction, and expansion to wide-area workspaces.

Additionally, I migrated the trained checkpoint to NVIDIA DGX Spark and built a standalone actor that does not depend on Isaac Lab. I then validated state acquisition from the real Unitree G1 and Inspire Hand, the right-arm transition to a task-ready pose, open-air hand motion, and the first bounded model-derived command.

This article summarizes the following contents.

  • Nominal integration with Learned Stage 2 connected
  • Coordinate semantics problem encountered with Integrated RandPos
  • Object-relative ingress / lift correction
  • Improvement from 55.7% to 97.4% success rate in ±1.5 cm condition
  • A ±5 cm workspace and Stage 0 coarse bias
  • 84.4% success in practical asymmetric workspace
  • Checkpoint migration to DGX Spark
  • Reconstruction of Standalone PyTorch actor
  • Inference matching between Brev x86 and DGX Spark ARM64
  • Initial integration with actual G1 / Inspire Hand
  • Failures encountered this week and notes on implementation
  • Current achievements and future challenges

The success rate in this article is a simulation result using simplified rigid-body bottle on Isaac Lab.

The real-robot work described here was conducted incrementally in a fixed standing position, under supervision, at low speed, and with either no object or a lightweight empty bottle. It does not demonstrate autonomous grasping and lifting of a filled beverage bottle, complete sim-to-real transfer, human handover, or safety certification.


Connect Learned Stage 2 to the integrated Pipeline

Last time, I confirmed that it was possible to reach Stage 2 from Stage 1 via controlled ingress.

In the first bridge diagnostic, Stage 2 residual was set to strictly zero.

learned Stage 1
  ↓
controlled ingress
  ↓
scripted Stage 2 base trajectory
  + zero residual

Although this configuration was successful in safe handover, it was not possible to lift the 0.62 kg bottle to any meaningful height.

The typical results were as follows.

MetricResult
Maximum center-height increaseapproximately 4.8 mm
Final center-height increaseapproximately 0.04 mm
Maximum horizontal displacementapproximately 43 mm
Strict retention / liftFalse

In other words, the geometry bridge from Stage 1 to Stage 2 was established, but a learned residual was required for grasp/retention under the 0.62 kg condition.

Therefore, I connected the selected Stage 2 checkpoint to the integrated environment.

Stage 1:
  frozen deterministic policy

controlled ingress:
  deterministic arm interpolation

Stage 2:
  scripted base trajectory
  + frozen deterministic 12D hand residual

Stage 2 actor inference runs only once per environment step, not per physics substep.

environment step:
  actor inference
  previous residual update
  target calculation

remaining physics substep:
  hold the same target

This is as important as the decimation problem I discovered last time.

If actor inference or previous residual update is executed for each physics substep, the policy will operate on a different time scale than during training.

After modification, Stage 2 actor inference count matched the expected value.

actual inference count:
  742

expected inference count:
  742

Under nominal conditions, I confirmed a meaningful lift of approximately 3.58 cm by enabling Stage 2 learned residual.

However, if the policy target continues to update after a successful lift, the bottle may gradually tilt and eventually slip.


Target Latch after successful lifting

For formal success determination, the following must be met continuously for a certain period of time.

  • bottle height
  • retention
  • workspace
  • uprightness
  • safety condition

In this configuration, I used 24 consecutive valid steps as the formal success condition.

After success, I saved the last commanded hand target and added a supervisory latch to hold that target.

retention / lift success
  ↓
copy current commanded hand target
  ↓
stop updating hand target
  ↓
continue normal physics simulation

This latch fixes only the command target.

The following operations have not been performed.

  • bottle attachment
  • bottle freeze
  • kinematic override
  • teleport
  • stopping physics
  • Direct change of object pose

Typical results after Latch were as follows.

MetricResult
Commanded-target drift0.000000000
Final retention count423
Final center-height increaseapproximately 0.0358 m
Final uprightnessapproximately 0.969

As a result, I was able to confirm the following sequence of operations under the nominal condition.

Stage 1 outer approach
  ↓
controlled ingress
  ↓
Stage 2 grasp
  ↓
lift
  ↓
retention
  ↓
stable extended hold

However, this result is a hybrid control.

learned Stage 1
  + deterministic geometric bridge
  + learned Stage 2
  + supervisory target latch

Not pure end-to-end RL.


Initial result of Integrated RandPos is 55.7%

After nominal integration was established, integrated RandPos evaluation was performed by changing the bottle position by ±1.5 cm in the x/y direction.

The conditions are below.

bottle:
  dynamic simplified rigid body
  0.62 kg

randomization:
  x/y ±0.015 m

Stage 1:
  frozen model_149

ingress / lift:
  fixed nominal joint-space targets

Stage 2:
  frozen model_100

success:
  evaluated by termination_manager
  24-step retention / lift success

The results were as follows.

SeedSuccess
42145 / 256
43283 / 512
Combined428 / 768

The combined success rate was 55.73%.

The failure distribution was as follows.

OutcomeCountRate
Success42855.73%
Unsafe bottle disturbance19725.65%
Object out of workspace435.60%
Timeout10013.02%

Stage 1 was 768/768 for individual evaluation, and Stage 2 was 766/768 for fixed conditions.

Nevertheless, the combined result was 428/768.

This shows that even if you multiply the success rates of individual policies, you will not get the integrated success rate.


The Problem Was Coordinate Semantics, Not Policy Performance

When I checked the failure map, I found that the failures were not random, but had a direction depending on the bottle position.

In the failure analysis of 1024 episodes, 233 out of 234 unsafe failures occurred in the area where the bottle's y offset was negative.

Also, timeout and workspace failure have increased on the positive x side.

From this result, I suspected a geometry inconsistency rather than simple PPO noise.

Stage 1 was Bottle-Relative

Stage 1 target is calculated from the bottle position.

outer_target
  = bottle_position + outer_offset

Therefore, even if the bottle moves, Stage 1 can move to the same relative position with respect to the bottle.

Ingress and lift were Nominal Joint Pose

On the other hand, ingress and lift after Stage 1 were fixed joint-space targets.

Stage 1:
  follow the displaced bottle

ingress:
  return to the nominal joint pose

lift:
  return to the nominal joint pose

In other words, it was object-relative up to Stage 1, but object-relative semantics were lost in subsequent trajectories.

bottle-relative outer pose
  ↓
fixed nominal ingress
  ↓
fixed nominal lift

This causes the hand to return to the position away from the bottle even if the bottle position changes by a few centimeters.

An important point is that before relearning checkpoints or PPOs, you need to make sure that the coordinate semantics are consistent across controllers.

Even if the learned policy is object-relative, if the subsequent scripted controller returns to nominal joint-space, the entire system is not object-relative.


Object-Relative correction using Local Jacobian

To solve this problem, I numerically measured the hand-center Jacobian around the nominal ingress pose.

The target joint is the following 4D.

  • right shoulder pitch
  • right shoulder roll
  • right shoulder yaw
  • right elbow

Hand center is the average position of the following four proximal links.

  • index proximal
  • middle proximal
  • ring proximal
  • pinky proximal

For conversion from Cartesian displacement to joint correction, I used measured Jacobian's pseudoinverse.

bottle_offset = settled_bottle_position - nominal_bottle_position

desired_hand_translation = gain × [offset_x, offset_y, 0]

joint_correction = Cartesian-to-arm pseudoinverse × desired_hand_translation

The target after correction is as follows.

corrected_ingress_target = nominal_ingress_target + joint_correction

corrected_lift_target = nominal_lift_target + joint_correction

By applying the same correction to both ingress and lift, the object-relative geometry is maintained even after starting the grasp.

Measured matrix generally had the following characteristics.

Jointx directiony direction
Shoulder pitch-5.28 rad/m-1.08 rad/m
Shoulder roll+0.32 rad/m+1.21 rad/m
Shoulder yaw-1.79 rad/m+1.66 rad/m
Elbow+5.40 rad/m+1.42 rad/m

The RMSE of the Jacobian fit was approximately 0.036 mm.

However, this is a local model around the nominal ingress pose.

This does not mean that you can directly extrapolate to ±10 cm or ±15 cm.


Comparison of Correction Gain

Correction gain was compared under x/y ±1.5 cm conditions.

GainSuccess
0.0082 / 128
0.50116 / 128
0.75124 / 128
1.00126 / 128

I selected Gain 1.0 and conducted a formal evaluation.

The following remains unchanged.

  • Stage 1 checkpoint
  • Stage 2 checkpoint
  • PPO weight
  • observation
  • residual scale
  • self-collision
  • bottle physics
  • formal success condition
  • unsafe threshold

No additional training is provided.


Success rate 97.4% at ±1.5 cm

The formal evaluation results with Object-relative ingress/lift correction enabled are shown below.

Seed 42

OutcomeResult
Success253 / 256
Unsafe0
Workspace0
Timeout3

Seed 43

OutcomeResult
Success495 / 512
Unsafe0
Workspace3
Timeout14

Combined

OutcomeResultRate
Success748 / 76897.40%
Unsafe0 / 7680%
Workspace3 / 7680.39%
Timeout17 / 7682.21%

A comparison before and after correction is as follows.

MetricBeforeAfter
Success428748
Success rate55.73%97.40%
Unsafe1970
Total failures34020

Success increased by 320 episodes.

Failures decreased from 340 to 20, a decrease of approximately 94.1%.

The success rate improved by 41.67 percentage points without any additional PPO and just by modifying the coordinate semantics.

This was the most important result of the week.

Before scaling or retraining the policy, verify that frame and target semantics remain consistent between learned and geometric control.


Extending the Workspace to ±5 cm

Next, I extended the bottle position range to x/y ±5 cm.

First, I evaluated it using only the same object-relative correction and without adding Stage 0.

The results were as follows.

OutcomeCountRate
Success536 / 76869.79%
Unsafe20 / 7682.60%
Workspace2 / 7680.26%
Timeout210 / 76827.34%

The conditional success after reaching Stage 2 was approximately 91.94%.

Stage 2 reached:
  583 / 768

success after Stage 2:
  536 / 583

In other words, if I could reach Stage 2, I would have succeeded in grasping/lifting in many episodes.

The main bottleneck was that Stage 1 had a large workspace and could not reach the outer gate.

In particular, the positive y direction was weak, and the following results were obtained for the bin at y=+4 to +5 cm.

samples:
  51

outer gate:
  0 / 51

success:
  0 / 51

In this area, even if the episode duration was extended, the target could not be reached.


Stage 0 Coarse Bias

The Stage 1 policy was originally learned with a local capture range of ±1.5 cm.

Therefore, I added a deterministic Stage 0 before Stage 1.

The purpose of Stage 0 is to remove only the portion of the wide workspace offset that exceeds the Stage 1 local range.

bottle_offset
  ↓
local_offset 
  = clamp(bottle_offset, -0.015, +0.015)

coarse_translation 
  = bottle_offset - local_offset

Stage 0 bias = safe-pose Jacobian pseudoinverse × coarse_translation

The execution order is as follows.

Stage 0:
  60-step smooth coarse motion
  ↓
Stage 1:
  frozen local safe approach
  ↓
object-relative ingress / lift
  ↓
Stage 2

Stage 0 is a geometric controller that uses measured local Jacobian around safe pose instead of learned policy.

The selected settings are below.

local capture half-range:
  0.015 m

coarse steps:
  60

correction bound:
  0.24 rad

Results for Full ±5 cm Square

The formal evaluation results for full x/y ±5 cm square with Stage 0 enabled are as follows.

Seed 42

success:
  193 / 256
  75.39%

Seed 43

success:
  376 / 512
  73.44%

Combined

OutcomeCountRate
Success569 / 76874.09%
Unsafe16 / 7682.08%
Workspace21 / 7682.73%
Timeout162 / 76821.09%

Comparison with without Stage 0 is below.

MetricWithout Stage 0With Stage 0
Outer gate reached601750
Stage 2 reached583737
Success536569

Stage 0 has improved the majority of Stage 1 reach failures.

On the other hand, the success rate after reaching Stage 2 decreased to approximately 77.2%.

So, after being "reachable" at Stage 0, the next bottleneck was handover geometry at wide-area locations and Stage 2 compatibility.


Failure Analysis in the Positive-y region

When analyzing 162 Full square timeout cases, 147 of them reached Stage 2.

The initial position averages of successful episodes and timeout episodes had the following trends.

OutcomeMean xMean y
Success-0.0014 m-0.0106 m
Timeout after Stage 2+0.0088 m+0.0248 m
Workspace failure+0.0244 m+0.0093 m

On the positive y side, I thought that the grasp geometry would easily deviate from the Stage 2 training distribution.

So I tried the following.

  • global ingress gain
  • positive-y trim
  • pre-contact alignment

Global Gain 1.04

Fixed y=+5 cm could be successful with gain 1.04.

However, it was worse in paired wide evaluation.

GainSuccess
1.00204 / 256
1.04194 / 256

Parameters that are valid for Fixed-point are not necessarily valid for the entire workspace.

Positive-y Trim

In the first-episode-per-environment paired comparison, no stable improvement was obtained by trim.

TrimSuccess / 128
0 mm97
1 mm91
2 mm96
3 mm98

I did not adopt it because there were conditions that would increase Unsafe disturbance.

Pre-contact Alignment

I also tried additional alignment using hand center and bottle center errors.

Fixed y=+5 cm succeeded in some gains, but wide paired evaluation did not improve.

alignment off:
  97 / 128

y-only gain 0.25:
  95 / 128

Based on this result, I decided not to use fixed-point improvement as a global correction.


Choose Practical Workspace over Symmetric Square

Full ±5 cm square is a valid comparison condition for research.

However, it is not necessarily necessary to stick to a perfect square including the edges of positive y.

In post-hoc analysis of full-square evaluation, the following workspaces showed relatively good results.

x:
  [-0.05, +0.05] m

y:
  [-0.05, +0.03] m

This is an asymmetric workspace with a width of 10 cm and a depth of 8 cm.

However, a post-hoc mask alone cannot be called a formal result.

Therefore, I implemented this range as an explicit randomization distribution and conducted prospective evaluation.


84.4% in Practical Asymmetric Workspace

The selected workspaces are:

bottle x:
  [-0.05, +0.05] m

bottle y:
  [-0.05, +0.03] m

size:
  10 cm × 8 cm

The Controller configuration is as follows.

Stage 0:
  enabled

Stage 1:
  frozen model_149

ingress / lift:
  object-relative correction
  gain 1.0

Stage 2:
  frozen model_100

positive-y trim:
  disabled

pre-contact alignment:
  disabled

The formal evaluation results were as follows.

Seed 42

OutcomeResult
Success209 / 256
Success rate81.64%
Unsafe2
Workspace3
Timeout42

Seed 43

OutcomeResult
Success439 / 512
Success rate85.74%
Unsafe2
Workspace11
Timeout60

Combined

OutcomeCountRate
Success648 / 76884.375%
Unsafe4 / 7680.52%
Workspace14 / 7681.82%
Timeout102 / 76813.28%

Phase results are below.

PhaseCountRate
Outer gate reached764 / 76899.48%
Ingress started764 / 76899.48%
Stage 2 reached763 / 76899.35%
Success after Stage 2648 / 76384.93%

With this result, I was able to exceed my initial goal of 80% integrated success with an explicitly defined practical workspace.


Separate the three types of results

The results of this time need to be recorded separately depending on the workspace conditions.

LevelWorkspaceResult
High-reliability localx/y ±1.5 cm748 / 768, 97.40%
Full symmetric squarex/y ±5 cm569 / 768, 74.09%
Practical asymmetricx ±5 cm, y -5 to +3 cm648 / 768, 84.375%

These are not the same results.

  • ±1.5 cm is the most reliable local result
  • Full ±5 cm square is a research result of less than 80%
  • 10 cm × 8 cm is a practical result of more than 80% in prospective evaluation

It is important that the weak areas of Full square were not only excluded post-hoc, but also re-evaluated as a separate distribution.


Why additional PPO was not performed

In this improvement, the checkpoints of Stage 1 and Stage 2 were not changed.

selected Stage 1:
  model_149

selected Stage 2:
  model_100

The reason why I did not perform additional PPO is that the main failure was not a lack of policy capacity, but the following.

  • Mismatch in object-relative semantics
  • Wide offset input to local policy
  • Stage 1 reach range
  • Handover difference with Stage 2 training distribution
  • phase / timing implementation

If you continue PPO, it is possible to absorb some failures.

However, if you add training when the controller's frame semantics are incorrect, the policy will learn inconsistencies in the system design.

This large improvement is due to system integration fixes, not additional training.

frozen learned Stage 1
  + measured geometric correction
  + optional Stage 0
  + frozen learned Stage 2
  + supervisory hold

This is a hybrid learned-plus-geometric control system.


Migrate checkpoints to DGX Spark

After confirming the simulation milestone, I migrated the trained checkpoints to NVIDIA DGX Spark.

The purpose is not to run Isaac Sim or Isaac Lab on DGX Spark.

simulator:
  run in the Isaac Lab environment

DGX Spark:
  standalone actor inference
  robot state adapter
  safety supervisor
  logging
  future perception runtime

The selected checkpoints are as follows.

Stage 1

input:
  24D

network:
  24 → 128 → 64 → 4

activation:
  ELU

Stage 2

input:
  53D

network:
  53 → 256 → 128 → 64 → 12

activation:
  ELU

I extracted only the actor part from Checkpoint and rebuilt it as a standalone PyTorch network.

This runtime does not require:

  • Isaac Sim
  • Isaac Lab
  • Omniverse
  • ManagerBasedRLEnv
  • RSL-RL runtime

Comparing inference results between x86 and ARM64

Even if Checkpoint binary can be copied, it does not necessarily mean that the inference results will be the same with a different architecture.

Therefore, I compared the actor output for the same input tensor in the x86 environment on Brev and the ARM64 environment on DGX Spark.

The following was used for comparison:

  • zero observation
  • deterministic synthetic probe
  • actor mean output
  • tolerance 1e-6

The results were as follows.

ActorInputMaximum absolute difference
Stage 1Zero2.98e-8
Stage 1Probe1.49e-8
Stage 2Zero2.98e-8
Stage 2Probe1.19e-7

All were within 1e-6.

This allowed us to confirm the following:

  • checkpoint integrity
  • actor architecture reconstruction
  • layer order
  • activation function
  • deterministic inference
  • x86 / ARM64 numerical parity

However, this parity only proves that if you input the same tensor, the actor outputs will match.

The following is not proven.

  • real-world observation equivalence
  • real robot action mapping
  • real camera accuracy
  • real grasp success
  • real Jacobian validity

Observation Contract is important in Real Runtime

If you only need to run a standalone actor, checkpoint and network architecture are sufficient.

However, to connect policy to actual robot state, I need to fully reproduce the meaning of observation.

Stage 1 Observation

Stage 1 is 24D.

TermDimension
Right-arm absolute joint position4
Right-arm joint velocity4
Hand-center position3
Hand-center linear velocity3
Bottle position3
Hand-center to outer-target vector3
Previous raw action4
Total24

Stage 2 Observation

Stage 2 is 53D.

TermDimension
Right-hand articulated joint position12
Right-hand joint velocity12
Hand-center position3
Bottle position3
Bottle quaternion wxyz4
Bottle linear velocity3
Bottle minus hand-center3
Retention phase1
Previous raw residual12
Total53

Even if the Observation dimensions match, they are not compatible if the following differ.

  • term order
  • coordinate frame
  • quaternion order
  • unit
  • normalization
  • timing
  • previous action semantics

Similar to the Checkpoint architecture, the observation config during training must be treated as a source of truth.


Connecting Inspire Hand's 6D Motor Space and Simulator's 12D representation

The right hand used in Simulation is represented by 12 articulated joints.

On the other hand, in the real robot RH56DFQ state obtained this time, the right hand is treated as a 6-dimensional motor-space state.

When I checked the official DFQ URDF, 6 out of 12 articulated joints had mimic relations defined.

Typical relationships are as follows.

index intermediate
  = index proximal

middle intermediate
  = middle proximal

ring intermediate
  = ring proximal

pinky intermediate
  = pinky proximal

thumb intermediate
  = 1.6 × thumb proximal pitch

thumb distal
  = 2.4 × thumb proximal pitch

Based on this mimic structure, I created an adapter that expands the 6D motor-space state into a 12D articulated representation that can be input to Stage 2 policy, and projects the Stage 2 output back to the 6D motor-space candidate.

real motor-space q6
  ↓
initial motor-to-joint range model
  ↓
DFQ mimic expansion
  ↓
simulator-side 12D articulated representation
  ↓
Stage 2 actor
  ↓
unilateral target q12
  ↓
least-squares projection onto the DFQ manifold
  ↓
bounded motor-space q6 candidate

I confirmed roundtrip error 0 for 6D → 12D → 6D within the adopted conversion model.

However, what this roundtrip result shows is that the adapter's forward/inverse equation is internally consistent. This does not mean that robot-specific calibration of the actual motor coordinate and URDF joint coordinate has been completed.

At this time, additional real robot confirmation is required for motor order, command/state index correspondence, range, sign, zero offset, motor position, and scale of URDF joint angle.

Also, although I was able to confirm visual hand motion using the real robot command, I was unable to confirm a clear proportional relationship between the commanded motor change and the expected index response of rt/inspire/state.

Therefore, the 12D state used in this article is not a fully calibrated actual joint angle, but a simulator-side articulated representation based on the official DFQ mimic structure and initial range model.

Furthermore, the simulator's SIDE_OPEN/CONTACT/SQUEEZE pose does not completely exist on the actual DFQ's mimic manifold.

Therefore, I recorded a projection error when converting Stage 2 target to motor space, and explicitly limited the closing direction and maximum delta in the real robot command.


In the real robot, start from Normal Standing

The safe pose of the simulation and the native normal standing pose of the G1 were very different.

As a result of comparing the hand center with Official-model FK, the distance from Normal standing to simulator task-ready pose was approximately 51 cm.

hand-center translation:
  x approximately +0.35 m
  y approximately +0.15 m
  z approximately +0.34 m

total distance:
  approximately 0.51 m

This is no small policy residual.

Therefore, Real Stage -1 was added before Stage 1 on the real robot.

Real Stage -1:
  Normal standing
  → task-ready right-arm pose

Stage 1:
  local task approach

Rather than sending the simulation reset pose as is to the real robot, I recorded the right-arm pose actually generated by the robot's native controller as read-only and used it as a task-ready candidate.


Real Stage -1 step-by-step verification

Rather than increasing the control weight of Arm SDK all at once, I verified the following step by step.

  1. Current-state hold
  2. Low-weight bumpless takeover
  3. Small shoulder / elbow waypoint
  4. Full-weight fixed-anchor takeover
  5. transition to captured task-ready pose
  6. continuous hold in captured pose
  7. Controlled return
  8. Release to Normal controller

The first exact-state hold continued to command the fixed pose recorded at the start, which conflicted with the native controller's small pose adjustments and exceeded the deviation threshold.

Next, I tried bumpless takeover, which follows the live state during takeover and gradually transfers control ownership at a low weight.

With this method, the worst deviation during frozen hold was approximately 0.00045 rad, confirming stable takeover.

After that, in the full-weight takeover, the right-arm pose obtained just before the takeover was used as a fixed anchor.

The typical results were as follows.

MetricResult
Maximum weight1.0
Worst velocityapproximately 0.244 rad/s
Worst deviationapproximately 0.024 rad
Final deviationapproximately 0.017 rad
ResultPASS

Velocity Spike caused by Measured-State Latch

After reaching the task-ready pose, I tried saving the measured joint state as a new hold target.

However, a tracking error of approximately 0.07 rad remained between the commanded target and the actual measured pose.

A velocity spike occurred on the right elbow because the command target was instantaneously changed to the measured state when the weight was 1.0.

previous command:
  captured target

new command:
  measured state

command discontinuity:
  approximately 0.07 rad

observed elbow velocity:
  approximately 0.44 rad/s

As a result, the software safety gate was activated and the arm weight was released.

After the modification, the method was changed to continuously publish the same captured target during task-ready hold.

HOLD_TARGET_POLICY:
  CONTINUOUS_CAPTURED_TARGET

MEASURED_STATE_LATCH:
  DISABLED

This fix resulted in the following series of successful dry runs.

Normal standing
  ↓
captured task-ready pose
  ↓
continuous arm hold
  ↓
scripted empty-air hand close
  ↓
hand restore
  ↓
controlled arm return
  ↓
normal weight release

From this failure, I learned that it is not always safe to set the current value as the target in the actual control.

When control weight is high, switching targets becomes a step command in itself.


Open-Hand Proximity with an empty bottle

Next, I performed a proximity test in which I manually placed an empty lightweight plastic bottle and moved it to a task-ready pose without closing my hands.

The conditions are below.

bottle:
  empty lightweight plastic bottle

pose source:
  manual measurement

hand:
  open

hand command:
  none

contact:
  not intended

perception:
  not used

The typical results were as follows.

MetricResult
Worst arm velocityapproximately 0.313 rad/s
Worst tracking errorapproximately 0.085 rad
Proximity readyYes
Hold completedYes
Return to startPass
Hard faultNo
Soft faultNo
Final arm errorapproximately 0.0010 rad

This is an open-hand proximity result for manual nominal placement.

It does not imply any of the following:

  • camera-based approach
  • contact grasp
  • object retention
  • object lift
  • autonomous manipulation

First Model-Derived actual command

Stage 2 model output was first verified as a shadow target.

After that, a closing-only command limited to a maximum of 0.02 normalized motor units was sent to the actual hand in a no-object condition.

Stage 2 model:
  model_100

observation:
  live hand state
  + explicitly labeled nominal geometry

controlled coordinates:
  four fingers

thumb:
  held current

direction:
  closing only

maximum delta:
  0.02

object:
  none

The generated bounded delta was as follows.

[-0.0141,
 -0.0200,
 -0.0200,
 -0.0200,
  0.0000,
  0.0000]

Command reader match, ramp, hold, and restore have been completed.

However, the observed state change was approximately 0.001 to 0.002, and the finger motion was not large enough to be clearly seen with the naked eye.

Therefore, I treat this result as follows.

model inference:
  passed

DFQ projection:
  passed

bounded command publication:
  passed

restore:
  passed

effective grasp:
  not established

Passing the command path and achieving a practical grasp are two different results.


Integration of Real Stage -1 and Stage 2 Model Command

Finally, I executed the validated Real Stage -1 arm hold and bounded Stage 2 model-derived hand command in the same sequence.

Normal standing
  ↓
captured Real Stage -1
  ↓
continuous captured-target hold
  ↓
bounded Stage 2 model hand command
  ↓
hand restore
  ↓
controlled arm return
  ↓
normal arm weight release

The typical results were as follows.

MetricResult
Arm command readerMatched
Hand command readerMatched
Arm readyYes
Stage 2 command completedYes
Hand restore completedYes
Return to startPass
Hard faultNo
Soft faultNo
Final arm errorapproximately 0.0005 rad
Final hand restore error0

This is the result of the Stage 2 model-derived command being sent to the real robot.

On the other hand, the Stage 1 actor does not directly control the actual arm yet.

Real Stage -1 is a reviewed scripted transition using captured native pose.


Offline design of +3 cm Lift Target

As the next prototype, I designed an offline target that vertically raises the hand center by 3 cm from the captured task-ready pose.

I used Official URDF, numerical Jacobian, and damped least squares IK to correct only the 4 joints of the right shoulder/elbow.

Wrist joint is fixed at the captured pose value.

desired translation:
  x 0
  y 0
  z +0.03 m

The results were as follows.

MetricResult
Requested lift0.030000 m
Actual lift0.029994 m
Endpoint position errorapproximately 0.015 mm
Z errorapproximately 0.006 mm
Maximum XY driftapproximately 0.014 mm
Jacobian rank3
Condition numberapproximately 4.13
Maximum joint correctionapproximately 0.0633 rad
Predicted maximum velocityapproximately 0.0162 rad/s
Joint limitsPass
Waypoint convergencePass

The main joint correction was about -0.0633 rad on the right elbow.

This result is an offline kinematic candidate.

The following have not been confirmed yet.

  • real actuator tracking
  • self-collision
  • support-rack clearance
  • table clearance
  • cable interference
  • object load

Before using it on an actual device, you must first perform a dry test with no-object / open-hand conditions.


Main Pitfalls I encountered this week

1. Integration fails even with Strong Modular Policies

Even though Stage 1 was 768/768 and Stage 2 was 766/768, the initial integrated RandPos was 428/768.

Success of individual policy does not guarantee compatibility of handover state distribution.

2. Don't lose Coordinate Semantics midway

Even if Stage 1 is object-relative, if the subsequent ingress returns to the nominal joint-space, the entire system is not object-relative.

3. Don't extrapolate Local Jacobian to a wide range

Local Jacobian is reliable only around its pose.

A correction that was valid for ±1.5 cm cannot be used for ±10 cm without verification.

4. Fixed-Point Improvement should not be called Global Improvement

Gain and alignment that were successful at a specific y=+5 cm position sometimes deteriorated in wide paired evaluation.

5. Same Seed alone does not result in Paired Comparison

In batch evaluation that includes episodes after auto-reset, the episode sequences may not match even with the same seed.

True paired comparison requires explicit design, such as first episode per environment.

6. Exit Code alone cannot determine success

Even if the simulation application exits with code 0, there may be a traceback or missing completion marker.

You need to check the result marker, completion marker, JSON, and episode count together.

7. Switching the target on the actual device is itself a command

If you switch the target to the measured state when the weight is high, a velocity spike may occur even if the difference is small.

Continuous targets and controlled returns must be explicitly designed.


Current achievements

The simulation results are below.

Result levelWorkspaceSuccess
Stage 1 modularx/y ±1.5 cm768 / 768
Stage 2 modularFixed nominal766 / 768
Initial integrated baselinex/y ±1.5 cm428 / 768
Object-relative correctedx/y ±1.5 cm748 / 768
Full wide squarex/y ±5 cm569 / 768
Practical asymmetricx ±5 cm, y -5 to +3 cm648 / 768

Deployment/hardware integration is below.

ItemStatus
Checkpoint migration to DGX SparkCompleted
x86 / ARM64 actor parityPassed
Standalone actor inferencePassed
Real G1 state acquisitionPassed
Real Inspire Hand state acquisitionPassed
Official-model hand-center FKPassed
Real Stage -1 right-arm transitionPassed
Scripted empty-air hand motionPassed
Manual bottle open-hand proximityPassed
First bounded Stage 2 real commandPassed
Combined Stage -1 + Stage 2 dry runPassed
+3 cm lift target offline designPassed
Real no-object +3 cm liftNot yet tested
Real object grasp / liftNot yet tested

Future plans

In the next real robot verification, I will first check the +3 cm lift target designed offline under a no-object condition.

Normal standing
  ↓
captured task-ready pose
  ↓
open-hand +3 cm lift
  ↓
captured task-ready pose
  ↓
Normal standing

If this dry test is successful, I will next consider minimizing gripping and lifting using an empty sealed lightweight plastic bottle.

manual central placement
  ↓
captured pregrasp
  ↓
bounded scripted closure
  ↓
short hold
  ↓
maximum +3 cm lift
  ↓
return bottle to table
  ↓
hand restore
  ↓
arm return

The first object-lifting step should use a lightweight empty plastic bottle, not a filled bottle or a 0.62 kg payload.


Summary

In the previous article, I built Stage 1 Safe Approach and Stage 2 Grasp/Lift for 0.62 kg bottle separately and confirmed that they can be connected by controlled ingress.

This week, I integrated learned Stage 2 and made it work up to grasp/lift/retention under nominal conditions.

Then, when I performed integrated RandPos evaluation of x/y ±1.5 cm, the initial success rate was 55.73%.

The cause was not policy performance, but Stage 1 was object-relative, while ingress and lift returned to fixed nominal joint-space target.

By using Measured local Jacobian and correcting both ingress and lift according to bottle offset, the success rate improved to 97.40% without additional PPO.

Additionally, I expanded the workspace with Stage 0 coarse bias.

Full x/y ±5 cm square was 74.09%, but as a result of prospective evaluation of a practical asymmetric workspace, x [-5,+5] cm, y [-5,+3] cm, I achieved 648/768, 84.375%.

After the simulation milestone was confirmed, I migrated the checkpoint to DGX Spark and confirmed that the deterministic actor outputs of Brev x86 and DGX Spark ARM64 matched within 1e-6.

In the real robot, I confirmed the official-model FK, Real Stage -1 from Normal standing to task-ready pose, continuous arm hold, scripted empty-air grasp, manual bottle proximity, and the first bounded Stage 2 model-derived command.

However, I have not yet performed object grasping or lifting on a real robot.

The current milestone is as follows.

By combining object-relative geometric correction and frozen PPO policy, I achieved 84.375% formal integration success in bounded randomized workspace. Furthermore, I migrated the checkpoint to DGX Spark and confirmed the task-ready transition and the first bounded Stage 2 model-derived hand command on the real G1.

This is not pure end-to-end RL, but a hybrid system that combines learned policies, geometric control, and supervisory control.

In the future, after verifying the +3 cm lift target designed offline without an object, I plan to move on to minimum pick-and-lift using a lightweight empty bottle in a fixed standing position and under supervision.

I will continue by validating observation semantics, action mapping, contact conditions, and safe return behavior one step at a time, without treating a high simulation success rate as evidence of real-robot performance.