Skip to content

How verification works

Before any code reaches your robot, Quatern runs it against a recording of your robot's real sensors and checks the result itself. This page explains what a recording is, which checks run, and what a READY verdict does and doesn't promise.

Recordings

A recording (a capture) is a sensor session from your robot: every source in the robot's config, time-stamped, for up to five minutes.

quatern capture --robot c101 --seconds 30 --label hallway

During a capture Quatern only records; it never commands motion. You drive or jog the robot yourself. The CLI tells you what to do:

What to do: move the robot through the space it will work in — drive or jog it around
for the full 30 s and bring it back to where it started. Quatern only records; it never
commands motion during a capture.

End where you started

Bring the robot back to its starting spot. Quatern measures drift as the distance between where the run started and where it ended, which this protocol makes a known zero.

Setting Default Range
Length (--seconds) 30 s 1–300 s
Label (--label) unlabeled up to 64 characters

Each capture gets an id like c101.default_2026-10-02_hallway_v1 (<robot>.<instance>_<date>_<label>_v<n>) and a stream-health table:

recorded c101.default_2026-10-02_hallway_v1 (10.0 s)
  stream       verdict  health
  wheel_odom   ok       50.0 Hz, dropout 0%, noise 0.0040

Catalog robots ship with a sample capture, so quatern verify works offline right after quatern init --robot <name>. On the simulator, quatern quickstart records one for you. On a ROS 2 robot, captures are recorded with ros2 bag record.

Replay in a sandbox

quatern verify replays the capture as a stream of time-stamped sensor frames through your localizer module, which runs as a separate process in a sandbox. The harness computes every signal itself, so no localizer grades its own work.

quatern verify --robot c101 --goal "Navigate around the island to reach (1.2, 2.0)"
  • Sandbox. The module runs under resource limits (memory, CPU time, open files, processes; by default 1 GB of memory and a 60 s run). On Linux with bubblewrap it also gets its own filesystem and network namespace. On macOS the isolation tier is none: resource limits only. quatern doctor reports the tier.
  • Modes. In mapping mode the localizer estimates the trajectory and the harness builds an occupancy map from it. In localize mode it estimates its pose against a saved map.
  • Speed. A module slower than its input fails. The report shows how fast it ran on this computer (HOST) against the input rate.
  • Planning. If localization didn't fail, the planner plans to your goal on the map built from the capture. Quatern doesn't plan on a broken world model.
  • Retries. When signed in, the agent gets up to 3 verification rounds to fix problems before it stops and asks for manual review.

To verify code you wrote yourself instead of generated code:

quatern verify --robot c101 --goal "reach (1.2, 2.0)" --module ./my_localizer

Cross-checks

Every pair of sensors that estimate the same part of the robot is compared. If a robot has wheel odometry, visual odometry and depth odometry all estimating its base, that's three pairs. Quatern measures how far apart the pair's motion is over 1-second windows, as a fraction of the distance travelled: a sensor with a 10% scale error reads as about 10%.

Kind of motion Warn above Fail above
Planar (mobile base) 15% 30%
Joints (arms) 8% 20%

You can override these per pair in the robot config, or for every robot in thresholds.cross_check_defaults in ~/.quatern/config.json.

Other checks in the report:

Check Warn Fail
Drift at end (start-to-end distance) over 0.25 over 1.0
Mean loop-closure error over 0.20 over 0.75
No loop closures in a run of 60 s or more warning –
Plan violates a constraint – always; cannot be waived

Streams are time-aligned first. A pair whose streams are skewed by more than 0.10 s, or that go silent for more than 0.50 s, is excluded and shown as SKIP, never silently compared. Exclusions are warnings. With three or more sensors, Quatern names the outlier when every pair it's in disagrees and every pair without it agrees.

One sensor means nothing to cross-check

If only one source estimates a part of the robot (for example a TurtleBot with only wheel odometry estimating its base), there's no pair to compare. The report says so: only wheel_odom estimates base on this capture; single source, no cross-check is possible for it.

Reading the report

Verification of c101.default — goal: Navigate around the island to reach (1.2, 2.0)
  localizer: python module, build none, run `python3 node.py`, sandbox tier none
  planner: python module, build none, run `python3 node.py`, sandbox tier none
  calibration: fresh (1 min old, cal_c101.default_20261002T033903884123)
  capture: c101.default_2026-10-02_sample_v1
  localization (mapping mode), 1 iteration(s):
    base:wheel_odom vs visual_odom  6.7%  pass  under 15%
    base:wheel_odom vs depth_odom   3.8%  pass  under 15%
    base:visual_odom vs depth_odom  3.1%  pass  under 15%
    drift at end: 0.067 (pass; warn > 0.25, critical > 1.0)
    loop closures: 4, mean error 0.288
    performance (HOST): 1544 Hz sustainable vs 115 Hz input, CPU 6%, sandbox tier none — keeps up
    (warning) loop_closure: Mean loop closure error is 0.29 across 4 closures
  plan: 36 waypoints over 2.54 in grid2d, 10.8 s, no violations
  next: proceed_to_deploy — Offline verification passed: ...
  verdict: READY for the deploy gate (the code executed in the sandbox against the recorded stream)
  stack: stk_c101.default_20261002T033939643180 (pin it with `quatern pin stk_c101.default_20261002T033939643180`)
  • Each cross-check row is pass, WARN, FAIL or SKIP.
  • performance is keeps up or TOO SLOW.
  • next: is what Quatern suggests: proceed_to_deploy, calibrate, recapture, retry_slam, retry_plan or escalate.
  • verify exits 0 when the verdict is READY and 1 when it's NOT READY.

A NOT READY report lists the blockers:

  plan: FAILED — Goal (3.00, 3.00) is inside an inflated obstacle or outside the limits
  next: retry_plan — The motion plan violates constraints: ...
  verdict: NOT READY
    - motion plan has 1 constraint violation(s): Goal (3.00, 3.00) is inside an inflated obstacle or outside the limits

What READY means

A stack is READY when all three of these hold:

  1. Localization didn't fail (warnings are allowed).
  2. A motion plan exists, succeeded, and has no constraint violations.
  3. The robot has a valid actuator calibration less than an hour old (quatern calibrate).

READY means the code ran in the sandbox against your recorded sensor stream, produced the numbers in the report, and the planner found a clean plan on the map built from that recording.

READY does not mean:

  • It has run on the robot. Verification is offline. The first live run happens behind the gate and the watchdog.
  • It's cleared for hardware. Hardware-only requirements (a fresh stop-on-silence check, measured stopping inputs, sandbox isolation for heavier robots, a map under 24 hours old) are checked at the deploy gate, not here.
  • There were no warnings. Pairs in the warn band, skipped pairs, and drift or loop-closure warnings can all still be READY. Read the report.
  • Every sensor was cross-checked. A part of the robot with a single sensor has nothing to compare against.
  • Arm motions are collision-checked. For arms, joint-space plans are checked against joint limits only.

Stacks and pinning

Every verification run is saved as a stack: one software configuration (code, parameters, verdict, plan), ready or not.

$ quatern stacks --robot c101
Stacks for c101 (instance default)
    stack id                                created           instance  verdict  capture
    stk_c101.default_20261002T033736628654  2026-10-02 03:37  default   READY    c101.default_2026-10-02_sample_v1
  pinned: none (deploy needs `quatern pin <stack-id>`)

$ quatern pin stk_c101.default_20261002T033736628654
pinned stk_c101.default_20261002T033736628654 as last-known-good for c101 (verified on instance default)
  • quatern pin refuses a stack that isn't READY.
  • A failed re-verification saves a new stack and leaves the pin alone, so your last-known-good stays deployable.
  • quatern deploy uses the pinned stack unless you pass --stack.