Contact Deflection is a MuJoCo benchmark for goal-conditioned projectile redirection with a free-floating spacecraft manipulator. A six-joint UR5 carries a shield and must intercept an incoming projectile from noisy position measurements while arm motion reacts on the uncontrolled spacecraft base.
The central research question is:
Can a low-frequency learned policy choose where, when, and how to contact a projectile while a model-based stack realizes those decisions at control rate—and can the resulting contact achieve a desired outgoing velocity without excessive spacecraft disturbance?
The current approach uses Soft Actor-Critic (SAC) to choose a structured end-effector contact intent, while estimation, inverse kinematics, trajectory generation, and torque control remain model-based. This separates learning contact strategy from controlling joint-level robot motion.
Example contact trajectories after a short SAC training run. These rollouts are qualitative demonstrations rather than converged-policy results.
The learned policy operates at 4 Hz and chooses a structured 8D contact action. A model-based stack maps this action to a reachable end-effector contact goal, solves terminal IK and twist matching, generates a joint trajectory, and tracks it at 200 Hz while MuJoCo integrates free-floating contact dynamics at 1 kHz.
The main responsibilities are:
| Component | Responsibility |
|---|---|
| Environment | Sample randomized tasks, simulate free-floating dynamics and contact, and compute rewards |
| Estimator | Maintain a Gaussian belief over projectile position and velocity |
| SAC policy | Choose an 8D end-effector-centric contact intent every 250 ms |
| Decoder | Map the latent action to contact time, pose, and terminal shield twist |
| Mink IK | Find a collision-aware reachable terminal configuration |
| Trajectory generator | Connect the current state to the requested terminal joint state |
| Controller | Track joint references with bounded torque |
| MuJoCo | Integrate multibody and contact dynamics at 1 ms |
The repository currently includes:
- executable MuJoCo/Gymnasium environment
- randomized free-floating spacecraft/contact dynamics
- Gaussian projectile state estimator
- structured 8D contact-action decoder
- collision-aware terminal IK using Mink
- bounded terminal-twist fitting
- quintic joint-space references
- torque control
- Stable-Baselines3 SAC training and evaluation
- deterministic evaluation and video rendering
- automated scientific checks
This is research code, not a flight-dynamics or hardware-fidelity simulator. Important approximations are documented under Known limitations.
Create the Conda environment:
conda env create -f environment.yml
conda activate contact-deflectionFor an existing environment:
conda env update -f environment.yml --prune
conda activate contact-deflectionAlternatively:
python -m pip install -e '.[dev,render,rl]'Run the deterministic checks:
pytest
ruff check src tests scripts experiments
mypy srcSave a workspace visualization:
python scripts/visualize_reachable_workspace.py \
--config configs/decoder.yaml \
--visualization videos/reachable_workspace.pngOr inspect it interactively:
python scripts/visualize_reachable_workspace.py \
--config configs/decoder.yaml \
--liveRun a short integration test before launching a full experiment:
python experiments/train_sac.py \
--timesteps 10000 \
--seed 0 \
--run-dir outputs/sac/smokeExample full training run:
python experiments/train_sac.py \
--timesteps 1000000 \
--seed 0 \
--ent-coef auto_0.3 \
--eval-episodes 20 \
--render-episodes 3 \
--post-contact-seconds 4 \
--run-dir outputs/sac/full_runTraining saves periodic evaluation data, the best periodic checkpoint, the final checkpoint, deterministic held-out evaluation, and rendered episodes:
outputs/sac/full_run/
├── final_model.zip
├── evaluation_summary.json
├── best/
│ └── best_model.zip
├── evaluations/
│ └── evaluations.npz
└── renders/
├── episode_00_seed_10000.mp4
└── ...
Evaluate an existing run without training:
python experiments/train_sac.py \
--eval-only \
--run-dir outputs/sac/full_run \
--eval-episodes 20 \
--render-episodes 3Or select a checkpoint explicitly:
python experiments/train_sac.py \
--eval-only \
--run-dir outputs/sac/full_run \
--checkpoint outputs/sac/full_run/best/best_model.zipEval-only outputs are written separately and do not overwrite the original training report.
Each episode samples physical and initial-state parameters
including projectile and shield mass, MuJoCo contact parameters, arm configuration, spacecraft rates, and projectile initial state.
The simulator state contains spacecraft pose and twist, arm state, and projectile state:
MuJoCo integrates nonlinear free-floating dynamics and contact at 1 kHz:
The simulator has access to ground-truth state for physics, reward computation, and evaluation. The policy does not receive ground-truth projectile state or contact time.
The projectile is observed through noisy position measurements in spacecraft coordinates:
A Kalman filter maintains a Gaussian belief over world-frame projectile position and velocity using a constant-velocity prior:
The posterior is transformed to spacecraft-relative coordinates before being passed to the policy.
The Gymnasium observation contains observation desired_goal
The 49D observation contains the projectile belief mean and covariance diagonal, arm state, spacecraft pose and twist, previous action, interception-corridor features, and episode progress.
The goal is the desired outgoing projectile velocity expressed in current spacecraft coordinates.
SAC acts every 250 ms:
Each action specifies a complete contact goal and joint trajectory. Only the first 250 ms is executed before observing and replanning.
| Layer | Period | Frequency |
|---|---|---|
| SAC / replanning | 250 ms | 4 Hz |
| Torque control | 5 ms | 200 Hz |
| MuJoCo integration | 1 ms | 1 kHz |
This gives SAC a short receding-horizon decision sequence while retaining fine contact simulation.
For projectile belief mean
This trajectory is intersected with a spacecraft-attached ellipsoidal candidate workspace
If the feasible time interval is
The 8D action has the following semantics:
| Coordinate | Meaning |
|---|---|
| 0 | Coupled contact position/time along the estimated projectile trajectory |
| 1–2 | Shield-normal tilt about the contact tangent axes |
| 3–5 | Shield-center linear velocity in normal/tangent coordinates |
| 6–7 | Angular velocity about tangent axes; currently zero-scaled by default |
The contact frame is
Mink IK constrains shield-center position and alignment of the shield's local
IK freezes the measured spacecraft pose and enforces configured joint and collision constraints.
At the terminal configuration, desired joint velocity is obtained from a bounded regularized twist fit:
The current trajectory backend fits independent quintics
between the initial and requested terminal
Current implementation: this is endpoint interpolation, not constrained trajectory optimization. Position, velocity, acceleration, and jerk limits are checked for diagnostics but are not currently enforced by the trajectory solver.
Torque tracking uses bounded PD control:
Rewards are zero before termination.
Following shield contact, the environment executes a short follow-through and braking phase, waits for separation, and measures the outgoing projectile velocity.
The successful-contact reward is
where
Misses receive
Reward parameters are configured in configs/env.yaml.
SAC optimizes the standard maximum-entropy objective
The current training default uses adaptive entropy tuning initialized at auto_0.3). Short preliminary experiments favored adaptive entropy over fixed
Each reset samples Gaussian dynamics and initial-state parameters from configs/env.yaml.
Projectile trajectories use broad transverse variation around the candidate workspace. Outliers are rejection-sampled unless the true future trajectory crosses a 95%-scale inner workspace. This maintains a meaningful interception task without projecting samples onto the workspace boundary.
The desired outgoing velocity is sampled in the incoming contact frame with a positive normal component and Gaussian tangent components.
Episode parameters are available through:
info["episode_parameters"]configs/env.yaml— clocks, sensing, torque control, rewards, goal distribution, and randomized episode parametersconfigs/decoder.yaml— candidate workspace, IK, action scales, joint velocity limits, and trajectory derivative limitsenvironment.yml— reproducible Conda environment
Current velocity, acceleration, jerk, and torque limits are intentionally aggressive simulation settings and should not be interpreted as verified UR5 hardware limits.
assets/mjcf/ MuJoCo spacecraft/contact model
configs/ Environment and decoder parameters
experiments/train_sac.py Training/evaluation orchestration
scripts/visualize_reachable_workspace.py
src/contact_deflection/
├── envs/ Gymnasium environment + simulator
├── estimation/ Projectile belief + interception corridor
├── geometry/ Contact-frame construction
├── kinematics/ Workspace + collision-aware Mink IK
├── control/ Decoder, twist fit, trajectory, controller
├── rl/ SAC training + evaluation
└── visualization/ Offscreen + interactive visualization
tests/ Deterministic scientific checks
third_party/ Pinned upstream assets + licenses
scratch/ Ignored one-off investigations
Please preserve the separation between environment, estimation, contact decoding, kinematics, trajectory generation, control, and learning when extending the repository.
Use scratch/ for one-off debugging and investigations rather than adding temporary scripts to the tracked source tree.
Before committing:
pytest
ruff check src tests scripts experiments
mypy srcChanges that alter the benchmark definition, observation/action semantics, reward, randomization distribution, or model-based decoder should be reflected in the relevant configuration and documentation.
- Projectile motion is linear in the inertial frame before contact; CW orbital acceleration is intentionally out of scope.
- Spacecraft pose and twist are treated as known by the estimator.
- IK and terminal Jacobians freeze the current spacecraft pose rather than predicting generalized-Jacobian momentum coupling.
- MuJoCo still executes the resulting arm/base reaction dynamics.
- Collision constraints apply to terminal IK, not the full joint-space trajectory.
- The quintic backend is endpoint interpolation rather than collision-aware or dynamically constrained trajectory optimization.
- Trajectory derivative violations are diagnostic and do not currently prevent execution.
- Action coordinates 6–7 are reserved for shield angular velocity but are zero-scaled by default.
- Spacecraft geometry, contact parameters, and actuator limits are simplified research models rather than flight-qualified hardware.
- Terminal-reward design and SAC hyperparameters require validation over longer runs and multiple seeds.
UR5 meshes and the adapted arm model are derived from SpaceRobotEnv at the commit pinned in THIRD_PARTY.md.
The upstream Apache-2.0 license is reproduced under third_party/space_robot_env/.


