Skip to main content
A Toolkit running in a container can update itself to new releases. A small supervisor built into the image downloads the new version, checks it, and switches to it. The container keeps running: it isn’t recreated, and on Kubernetes the pod isn’t restarted. Only the Toolkit process inside restarts, so the host disconnects from AO briefly and reconnects on the new version. It keeps its identity, so it doesn’t enroll again. You decide whether that happens. Each organization has an update policy, each toolkit can override it, and each host has a local off switch that AO can’t override.
Docker, Docker Compose, and Kubernetes DaemonSet installs update themselves. Hosts installed with the install script or as a standalone binary, and the Kubernetes cluster controller, don’t: AO shows them when a newer release is available, and you update them yourself.

Update policies

Pin to a published release. AO refuses a pin newer than the latest release, but it doesn’t check that an older version string was ever published, and hosts can’t download a version that wasn’t: they report Download failed. Pinning never downgrades a host, and if AO has paused the release you pinned to, the pin waits until AO resumes it. Until your organization saves a policy, it follows AO’s default, which is Notify only today. While it does, the settings page marks that policy (default). AO may change its default in future, so if you rely on Notify only, select it and save it explicitly.

Set the organization policy

Go to Settings → Toolkit updates and choose Automatic, Notify only, or Pinned. For Pinned, enter the version under Version to hold at (for example 0.0.1-rc50), then Save. The policy applies to every toolkit in the organization that doesn’t have its own.

Override it for one toolkit

Open the toolkit from Toolkits. In the Updates card, under This toolkit’s policy, choose a policy and Save. Choose Use organization policy to remove the override. The toolkit’s own policy always wins over the organization’s. The Update policy line on the toolkit shows which one applies: Set on this toolkit, Organization policy, or AO default. Changing either policy needs the Edit toolkit upgrade policy permission. The Owner, Admin, and Toolkits admin roles have it. For a single toolkit’s override, the permission can be limited to toolkits with particular tags. Changes reach connected toolkits without a restart.

How an update is kept safe

Every release is signed by AO and verified on the host. The Toolkit checks the release’s signature, platform, and checksum against keys built into its image before it runs a new version, and again every time it starts. A download that doesn’t verify is never run. A new version must pass a health check, or it’s rolled back. The new version must connect to AO within three minutes of starting. It then runs a self-test that proves it can download and verify the next release. If it doesn’t connect in time, exits, or fails the self-test, the host goes back to the version it was running, automatically. For the 24 hours after an update, three crashes within ten minutes also roll it back (this needs a container restart policy; see Requirements). The host keeps the previous version on disk so it always has one to go back to. If the self-test can’t reach the registry, the new version keeps running while the self-test retries. After three unreachable attempts, or six hours without a pass, the host goes back to the previous version and tries the update again later. A version that failed its health check on a host is never offered to that host again. A failure caused by the network rather than the release, like the unreachable registry above, is retried later instead. AO never downgrades a host through a policy. Policies only ever move a host forward. The one exception is when AO withdraws a release: a host that updated to the withdrawn version moves back to the version it ran before. It does so within about an hour, as long as it’s connected to AO and can reach the registry. A withdrawn version that a host runs because it’s the version in the container image stays until you recreate the container from a different image. Releases roll out gradually, and AO can stop a rollout. Under Automatic, a new release reaches a growing share of toolkits over time, not all of them at once. If hosts start rolling back, AO stops the rollout. A pinned version skips the gradual rollout, because you chose it, but still waits while AO has stopped it. Running work is allowed to finish. Before it switches, the Toolkit stops accepting new actions and waits up to five minutes for actions in progress to finish. New actions sent during that wait are refused with a retryable error. If the work doesn’t finish in time, the Toolkit accepts actions again, postpones the update, and tries again two minutes later, with the same wait. A host that always runs long actions can therefore refuse new ones repeatedly. To avoid that, pin the host (or its organization) while that work runs, and move the pin forward when the host is quiet.

Requirements

  • Run the image’s own entrypoint. Don’t override the container’s entrypoint (--entrypoint, or command: in Kubernetes). The entrypoint starts the supervisor that installs updates. Without it, the Toolkit runs normally but never updates, and AO shows the host as having updates off. Passing flags as container arguments (args: in Kubernetes) is fine.
  • Keep /var/lib/ao-toolkit on a persistent volume. Updates are stored there, next to the host’s identity, so they survive the container being recreated. The install snippets in the dashboard already do this: a named volume for Docker and Compose, and a hostPath on the node for the Kubernetes DaemonSet. Without it, a recreated container starts again from the image’s version.
  • Let the host reach the container registry. The Toolkit downloads releases itself, over HTTPS, from ghcr.io and pkg-containers.githubusercontent.com (where ghcr.io serves downloads from). It honours HTTPS_PROXY and NO_PROXY. To use a registry mirror, set AO_UPDATE_REGISTRY to its host (or an https:// URL). The mirror must serve the Toolkit’s packages exactly as ghcr.io does, under the same names and paths, including the release package the update check reads, not just the container image. The signature check is the same whichever registry the bytes come from.
  • Allow programs to run from the update directory. Updates are run from /var/lib/ao-toolkit/bin. If that volume is mounted noexec (Container-Optimized OS on GKE mounts /var that way), the Toolkit refuses to update until you set AO_UPDATE_DIR to a directory that allows it. Put that directory on persistent storage too, or updates are downloaded again after the container is recreated.
  • Give the container a restart policy. If the Toolkit crashes, the container exits, and the crash-rollback check above only counts crashes when the container is restarted. The install snippets use --restart unless-stopped for Docker and restart: unless-stopped for Compose. Kubernetes restarts the pod’s container itself.

Docker and Docker Compose

  • docker ps and your Compose file keep showing the image you started with. After an update, the Toolkit runs a newer version than that image. AO shows both, for example running v0.0.1-rc52 (image v0.0.1-rc50).
  • Restarting the container, or recreating it with the same volume, keeps the update.
  • If you start the container from an image newer than the installed update, the image’s version runs and older updates are removed from the volume.

Kubernetes DaemonSet

  • The update happens inside the running pod. The pod spec doesn’t change, so kubectl keeps showing the image in your manifest, and the pod isn’t restarted.
  • Updates are stored in /var/lib/ao-toolkit on each node. A pod that restarts on the same node keeps its update. A pod on a new node starts from the image’s version and updates according to its policy.
  • GitOps tools such as Flux and Argo CD don’t see a change, so they neither report drift nor revert the update. If your pipeline should decide which version runs, use Notify only or Pinned, or set AO_AUTO_UPDATE=0 in your manifest. The DaemonSet snippet in the dashboard includes that line, commented out.

Kubernetes cluster controller

The cluster controller (the k8s-controller role) never updates itself. Update it by changing the image in its StatefulSet.

Turn off updates on one host

Set AO_AUTO_UPDATE=0 in the container’s environment:
The switch is read on the host, and nothing AO sends can turn it back on. Updates are on when AO_AUTO_UPDATE is unset, empty, 1, true, or on (ignoring case and surrounding spaces). Any other value turns them off, so a typo leaves updates off rather than on. Environment changes take effect when the container is recreated. With the switch off, the container runs exactly the version in its image, with no supervisor. If an update was installed earlier, it stays on the volume but doesn’t run, and AO shows it as downloaded, not running. Turning the switch back on returns to that update, as long as it still verifies. AO shows a host with the switch off as updates off. Whatever the policy, that host won’t install updates.
Toolkit images 0.0.1-rc49 and earlier turn this switch off in the image itself, so they show as updates off. To get automatic updates, recreate the container from a current image: docker pull the image, docker rm -f ao-toolkit, then run the same docker run command (or docker compose pull && docker compose up -d). Keep the same volume so the host keeps its identity.

What you see in AO

On Toolkits, a toolkit row can carry these markers: update failed appears only for Rolled back, Verification failed, and Download failed. Refused and Inconclusive don’t show a marker; open the toolkit to see them. On a toolkit’s page, Toolkit version shows what’s running, as running v… (image v…) after an update. The Updates card shows the policy and where it comes from, the next version, the last result with a short explanation, and Recent updates, which include the host’s own detail message for each result.

Troubleshooting

A host isn’t updating

Check, in order:
  1. The policy. The Update policy line must be Automatic, or Pinned to a version above the one running. Notify only never installs anything. Under Automatic, a new release reaches hosts gradually, so some hosts get it later than others.
  2. The local switch. If the host shows updates off, look for AO_AUTO_UPDATE in its environment, and check its image version: images 0.0.1-rc49 and earlier set it off.
  3. The entrypoint. A container started with --entrypoint or a Kubernetes command: never updates and shows updates off. Remove the override.
  4. The install type. Install-script and binary installs, and the cluster controller, don’t update themselves.
  5. The last update result. Download failed means the host can’t reach the registry (see Requirements). Refused carries a detail message, for example a noexec update directory or an image too old to verify the current release, which needs a recreate from a current image.
  6. Versions that failed here. A version listed under Not offered again on this host won’t be retried on that host. A later release will be.
A persistent volume doesn’t affect whether an update installs, but without one a recreated container goes back to its image’s version.

Read the logs

The Toolkit logs one JSON line per event. On Docker:
On Kubernetes, use kubectl logs on the DaemonSet pod. The messages to look for: docker exec ao-toolkit ao-toolkit version reports the image’s version, not an installed update. The Toolkit version in AO shows what is actually running.

Go back to an earlier version

Apart from moving off a release AO has withdrawn, a host never goes back to an older version on its own. A pin below the running version holds the host where it is rather than downgrading it. To run an earlier version on a host:
  1. Stop further updates: set the toolkit’s policy to Notify only or Pinned.
  2. Recreate the container from the image tag you want, with AO_AUTO_UPDATE=0 set, so the container runs exactly that image’s version.
Recreating from an older image without turning the switch off isn’t enough: an update installed on the volume is newer than the image, so it keeps running. Likewise, if you later remove AO_AUTO_UPDATE=0, the host immediately runs the newer update still on the volume, whatever its policy. To discard installed updates for good, also delete the update directory, /var/lib/ao-toolkit/bin (or the directory AO_UPDATE_DIR names, if you set it), on the volume or, on Kubernetes, on each node. Then restart the container. It starts from the image’s version and follows its policy from there.