Advanced Metal Research
GitHub Contact AMR

Run the weld planner

On this page
  1. Before you start
  2. Install
  3. Environments
  4. Check the GPU
  5. Serve the motion planner
  6. Bind it safely
  7. Server options
  8. Connect OLP to it
  9. Plan from the command line
  10. Inspect inputs
  11. Run the tests
  12. When something goes wrong
  13. Related pages

The weld planner lives in weld_planner/v1. It has two halves with different needs:

  • The seam worker (CAD topology, weld joints, torch angle search, packing .weldplan files) runs on the CPU in the default environment. OLP starts it as a subprocess for every request; you never run it as a server.
  • The motion planner (seam search, trajectory optimisation, the verifier and the dense trajectory encoder) needs an NVIDIA GPU and runs in the motion environment, as an HTTP server on port 8796.

For what the planner does and what its verification proves, see Weld planning and verification.

Before you start#

  • Linux on x86-64 for the motion environment. For an NVIDIA Jetson Thor, see Environments. The default environment also runs on Windows and macOS.
  • An NVIDIA GPU and a driver that supports CUDA 12 or later. The CUDA runtime comes from the environment; only the driver is the system's.
  • pixi. Nothing else: pixi installs Python, CadQuery, PyTorch and the rest.
  • A clone of RosieOS with its Git LFS files, because the planner reads the robot meshes from robot_description/.

Install#

cd weld_planner/v1
pixi install -e default   # seam worker and authoring tools (CPU)
pixi install -e motion    # motion planner (CUDA PyTorch, several GB)

Always pass -e to pixi run for motion tasks. Some task names, such as motion-test, exist in both environments.

Environments#

EnvironmentPlatformsUsed for
defaultlinux-64, linux-aarch64, win-64, osx-64, osx-arm64The seam worker, CAD inspection tools and the fast self-tests
motionlinux-64The motion planner and its tests, with CUDA PyTorch
motion-thorlinux-aarch64The motion planner on an NVIDIA Jetson Thor. pixi provides everything except PyTorch, which comes from NVIDIA's own image.
benchlinux-64Benchmarks against reference libraries. Not needed to plan.

The motion environment pins TORCH_ALLOW_TF32_CUBLAS_OVERRIDE=0, so matrix maths keeps full float32 precision on every GPU. The planner was qualified that way.

Check the GPU#

pixi run -e motion motion-gpu         # the device, its capability, and whether this PyTorch has kernels for it
pixi run -e motion motion-self-test   # every motion worker's smoke check
pixi run -e motion motion-verify      # the verifier's self-test

Serve the motion planner#

cd weld_planner/v1
pixi run -e motion motion-serve

motion-serve starts the FastAPI server on 0.0.0.0:8796 and sets AMR_WELD_PLANNER_SOURCE_REVISION to the checkout's git rev-parse HEAD, so every result names the code that planned it. The server logs one line per plan to stderr: MiB in, how many seams were crossed, moves, MiB out, time queued and time served.

Check it:

curl -s http://localhost:8796/api/motion/health
# {"cuda": true, "device": "…"}

"cuda": false means PyTorch cannot see the GPU. Planning will not work until it can.

Bind it safely#

Warning

The planner listens on every interface by default, with no authentication. Anyone who can reach port 8796 can queue plans on your GPU and download every stored trajectory. Bind it to localhost when OLP runs on the same machine, or allow only the hosts that need it through a firewall.

When OLP runs on the same machine, bind to localhost:

cd weld_planner/v1
AMR_WELD_PLANNER_SOURCE_REVISION=$(git rev-parse HEAD) \
  pixi run -e motion python -m weld_motion_planner.server.motion_planner_server --host 127.0.0.1

A Steam Deck pendant plans through a workstation's planner, so there the planner must listen on the network. Firewall port 8796 so that only the Deck (and your own machine) can reach it. For example, with ufw, with the Deck's address in place of the placeholder:

sudo ufw allow from <deck-address> to any port 8796 proto tcp
sudo ufw deny 8796/tcp

The dev stack's motion service runs pixi run -e motion motion-serve, so it too listens on every interface. See Run everything in simulation.

Server options#

FlagEnvironment variableDefaultDescription
--host0.0.0.0Listen address
--port8796Listen port
--dense-store-dirWELD_PLANNER_DENSE_STORE_DIR~/.cache/rosieos-olp/weld-planner-denseWhere verified .rdt files are stored, by digest. Refused requests are kept under refused/ in the same directory.
AMR_WELD_PLANNER_SOURCE_REVISIONSet by motion-serveA full 40-character lowercase Git commit, stamped into results. Any other non-empty value makes every plan fail.

The store sits outside the checkout on purpose, so resetting the workspace does not delete a trajectory a program still plays. It is still a cache: OLP keeps its own copy of each trajectory, and a Load after the file is gone answers dense_blob_not_found. Plan again.

Connect OLP to it#

OLP reaches the planner through one variable, which the OLP launcher reads:

export OFFLINE_PROGRAMMING_WELD_PLANNER_MOTION_ORIGIN=http://127.0.0.1:8796
bash offline-programming/v1/start-offline-programming.sh

The launcher prints a motion: line. motion: DISABLED means the variable is not set. Without it, Plan answers motion_origin_unavailable and Load answers dense_blob_source_unavailable. The dev stack sets it for you. See the OLP server's variables.

OLP also runs the seam worker from weld_planner/v1 for every seam request and every plan. By default it launches pixi run -e default python -m seam_worker.workers --stdin, so the default environment must be installed on the machine that runs the OLP server. Two variables change that:

VariableDefaultDescription
SEAM_WORKER_PYTHONunset (use pixi run)A Python executable to run the worker directly
SEAM_WORKER_SPARES2Workers started ahead of their request. 0 starts every request cold.

Plan from the command line#

You can plan a .weldplan without the server. This is useful for looking at one part in detail:

cd weld_planner/v1
pixi run -e motion motion-plan data/motion/bracket_a_2x.weldplan --trajectories --output result.json

It prints a summary per seam and writes the full result document to --output. The command line never writes a .rdt. Only the server does, with dense=true.

FlagDefaultDescription
requestrequiredPath to a .weldplan
--k5Candidate paths per seam
--samples-per-seamby timeSpace the lattice by sample count
--no-collideoffSkip collision screening and avoidance. The result says it is unscreened.
--trajectoriesoffAlso run M5 and the connecting moves (minutes rather than seconds)
--no-movesoffWith --trajectories, skip the connecting moves
--no-verifyoffSkip the verifier. Nothing planned this way can become a .rdt.
--output PATHnoneWrite the result document as JSON
--source-revisionAMR_WELD_PLANNER_SOURCE_REVISIONThe full Git commit to record

The stages also run alone: motion-seams (M4), motion-weld-trajopt (M5) and motion-link (M6), each with a .weldplan argument.

The committed example requests are in weld_planner/v1/data/motion/. pixi run -e default motion-fixtures rebuilds them from the STEP files in data/fixtures/.

Inspect inputs#

TaskWhat it shows
pixi run -e motion motion-request --read <file.weldplan>A .weldplan as the planner sees it
pixi run -e default plan-request --read <file.weldplan>The container's manifest, with every digest verified
pixi run -e default program-v2 <program.json>Problems in a robot.v4.program.v2 document
pixi run -e motion motion-cellWhich axes are the arm, which move the work, which belong to neither
pixi run -e motion motion-profileThe cell profile: axis roles, rates, reset pose, and what is missing
pixi run -e motion motion-spheresThe arm's sphere model and the torch built from tooling.json

Run any task with --help for its arguments.

Run the tests#

cd weld_planner/v1
pixi run -e motion motion-test      # the motion planner's tests, on the GPU
pixi run -e default test            # the whole suite in the authoring environment
pixi run -e default self-test       # fast smoke checks, no GPU

CI runs only a subset of the planner's tests, on a CPU build of PyTorch. Run motion-test on a GPU before you rely on a change.

When something goes wrong#

SymptomCause
/api/motion/health says "cuda": falsePyTorch cannot see the GPU. Check the driver with nvidia-smi, then pixi run -e motion motion-gpu.
A plan fails with no cell meshes at …The robot's meshes are missing. Fetch the Git LFS files. The verifier refuses to judge a cell it cannot see.
A plan fails at once with a source_revision errorAMR_WELD_PLANNER_SOURCE_REVISION is set to something other than a full commit hash. Unset it, or use motion-serve.
A plan answers but has no trajectoryRead dense_error. The request and result are kept under <dense store>/refused/. See A plan without a trajectory.
A second plan waitsOnly one plan runs at a time. GET /api/motion/progress shows the running stage.