Docker, Docker Compose, and Kubernetes DaemonSet installs update
themselves. Hosts installed with the install script or as a standalone
binary, and the Kubernetes cluster controller, don’t: AO shows them
when a newer release is available, and you update them yourself.
Update policies
Pin to a published release. AO refuses a pin newer than the latest
release, but it doesn’t check that an older version string was ever
published, and hosts can’t download a version that wasn’t: they report
Download failed. Pinning never downgrades a host, and if AO has
paused the release you pinned to, the pin waits until AO resumes it.
Until your organization saves a policy, it follows AO’s default, which
is Notify only today. While it does, the settings page marks that
policy (default). AO may change its default in future, so if you
rely on Notify only, select it and save it explicitly.
Set the organization policy
Go to Settings → Toolkit updates and choose Automatic, Notify only, or Pinned. For Pinned, enter the version under Version to hold at (for example0.0.1-rc50), then Save. The
policy applies to every toolkit in the organization that doesn’t have
its own.
Override it for one toolkit
Open the toolkit from Toolkits. In the Updates card, under This toolkit’s policy, choose a policy and Save. Choose Use organization policy to remove the override. The toolkit’s own policy always wins over the organization’s. The Update policy line on the toolkit shows which one applies: Set on this toolkit, Organization policy, or AO default. Changing either policy needs the Edit toolkit upgrade policy permission. The Owner, Admin, and Toolkits admin roles have it. For a single toolkit’s override, the permission can be limited to toolkits with particular tags. Changes reach connected toolkits without a restart.How an update is kept safe
Every release is signed by AO and verified on the host. The Toolkit checks the release’s signature, platform, and checksum against keys built into its image before it runs a new version, and again every time it starts. A download that doesn’t verify is never run. A new version must pass a health check, or it’s rolled back. The new version must connect to AO within three minutes of starting. It then runs a self-test that proves it can download and verify the next release. If it doesn’t connect in time, exits, or fails the self-test, the host goes back to the version it was running, automatically. For the 24 hours after an update, three crashes within ten minutes also roll it back (this needs a container restart policy; see Requirements). The host keeps the previous version on disk so it always has one to go back to. If the self-test can’t reach the registry, the new version keeps running while the self-test retries. After three unreachable attempts, or six hours without a pass, the host goes back to the previous version and tries the update again later. A version that failed its health check on a host is never offered to that host again. A failure caused by the network rather than the release, like the unreachable registry above, is retried later instead. AO never downgrades a host through a policy. Policies only ever move a host forward. The one exception is when AO withdraws a release: a host that updated to the withdrawn version moves back to the version it ran before. It does so within about an hour, as long as it’s connected to AO and can reach the registry. A withdrawn version that a host runs because it’s the version in the container image stays until you recreate the container from a different image. Releases roll out gradually, and AO can stop a rollout. Under Automatic, a new release reaches a growing share of toolkits over time, not all of them at once. If hosts start rolling back, AO stops the rollout. A pinned version skips the gradual rollout, because you chose it, but still waits while AO has stopped it. Running work is allowed to finish. Before it switches, the Toolkit stops accepting new actions and waits up to five minutes for actions in progress to finish. New actions sent during that wait are refused with a retryable error. If the work doesn’t finish in time, the Toolkit accepts actions again, postpones the update, and tries again two minutes later, with the same wait. A host that always runs long actions can therefore refuse new ones repeatedly. To avoid that, pin the host (or its organization) while that work runs, and move the pin forward when the host is quiet.Requirements
- Run the image’s own entrypoint. Don’t override the container’s
entrypoint(--entrypoint, orcommand:in Kubernetes). The entrypoint starts the supervisor that installs updates. Without it, the Toolkit runs normally but never updates, and AO shows the host as having updates off. Passing flags as container arguments (args:in Kubernetes) is fine. - Keep
/var/lib/ao-toolkiton a persistent volume. Updates are stored there, next to the host’s identity, so they survive the container being recreated. The install snippets in the dashboard already do this: a named volume for Docker and Compose, and ahostPathon the node for the Kubernetes DaemonSet. Without it, a recreated container starts again from the image’s version. - Let the host reach the container registry. The Toolkit downloads
releases itself, over HTTPS, from
ghcr.ioandpkg-containers.githubusercontent.com(whereghcr.ioserves downloads from). It honoursHTTPS_PROXYandNO_PROXY. To use a registry mirror, setAO_UPDATE_REGISTRYto its host (or anhttps://URL). The mirror must serve the Toolkit’s packages exactly asghcr.iodoes, under the same names and paths, including the release package the update check reads, not just the container image. The signature check is the same whichever registry the bytes come from. - Allow programs to run from the update directory. Updates are run
from
/var/lib/ao-toolkit/bin. If that volume is mountednoexec(Container-Optimized OS on GKE mounts/varthat way), the Toolkit refuses to update until you setAO_UPDATE_DIRto a directory that allows it. Put that directory on persistent storage too, or updates are downloaded again after the container is recreated. - Give the container a restart policy. If the Toolkit crashes, the
container exits, and the crash-rollback check above only counts
crashes when the container is restarted. The install snippets use
--restart unless-stoppedfor Docker andrestart: unless-stoppedfor Compose. Kubernetes restarts the pod’s container itself.
Docker and Docker Compose
docker psand your Compose file keep showing the image you started with. After an update, the Toolkit runs a newer version than that image. AO shows both, for examplerunning v0.0.1-rc52 (image v0.0.1-rc50).- Restarting the container, or recreating it with the same volume, keeps the update.
- If you start the container from an image newer than the installed update, the image’s version runs and older updates are removed from the volume.
Kubernetes DaemonSet
- The update happens inside the running pod. The pod spec doesn’t
change, so
kubectlkeeps showing the image in your manifest, and the pod isn’t restarted. - Updates are stored in
/var/lib/ao-toolkiton each node. A pod that restarts on the same node keeps its update. A pod on a new node starts from the image’s version and updates according to its policy. - GitOps tools such as Flux and Argo CD don’t see a change, so they
neither report drift nor revert the update. If your pipeline should
decide which version runs, use Notify only or Pinned, or set
AO_AUTO_UPDATE=0in your manifest. The DaemonSet snippet in the dashboard includes that line, commented out.
Kubernetes cluster controller
The cluster controller (thek8s-controller role) never updates itself.
Update it by changing the image in its StatefulSet.
Turn off updates on one host
SetAO_AUTO_UPDATE=0 in the container’s environment:
AO_AUTO_UPDATE is unset, empty, 1,
true, or on (ignoring case and surrounding spaces). Any other
value turns them off, so a typo leaves updates off rather than on.
Environment changes take effect when the container is recreated.
With the switch off, the container runs exactly the version in its
image, with no supervisor. If an update was installed earlier, it stays
on the volume but doesn’t run, and AO shows it as downloaded, not
running. Turning the switch back on returns to that update, as long as it
still verifies.
AO shows a host with the switch off as updates off. Whatever the
policy, that host won’t install updates.
Toolkit images 0.0.1-rc49 and earlier turn this switch off in the
image itself, so they show as updates off. To get automatic
updates, recreate the container from a current image:
docker pull
the image, docker rm -f ao-toolkit, then run the same docker run
command (or docker compose pull && docker compose up -d). Keep the
same volume so the host keeps its identity.What you see in AO
On Toolkits, a toolkit row can carry these markers:
update failed appears only for Rolled back, Verification
failed, and Download failed. Refused and Inconclusive don’t
show a marker; open the toolkit to see them.
On a toolkit’s page, Toolkit version shows what’s running, as
running v… (image v…) after an update. The Updates card shows the
policy and where it comes from, the next version, the last result with a
short explanation, and Recent updates, which include the host’s own
detail message for each result.
Troubleshooting
A host isn’t updating
Check, in order:- The policy. The Update policy line must be Automatic, or Pinned to a version above the one running. Notify only never installs anything. Under Automatic, a new release reaches hosts gradually, so some hosts get it later than others.
- The local switch. If the host shows updates off, look for
AO_AUTO_UPDATEin its environment, and check its image version: images 0.0.1-rc49 and earlier set it off. - The entrypoint. A container started with
--entrypointor a Kubernetescommand:never updates and shows updates off. Remove the override. - The install type. Install-script and binary installs, and the cluster controller, don’t update themselves.
- The last update result. Download failed means the host can’t
reach the registry (see Requirements). Refused
carries a detail message, for example a
noexecupdate directory or an image too old to verify the current release, which needs a recreate from a current image. - Versions that failed here. A version listed under Not offered again on this host won’t be retried on that host. A later release will be.
Read the logs
The Toolkit logs one JSON line per event. On Docker:kubectl logs on the DaemonSet pod. The messages to
look for:
docker exec ao-toolkit ao-toolkit version reports the image’s
version, not an installed update. The Toolkit version in AO shows
what is actually running.
Go back to an earlier version
Apart from moving off a release AO has withdrawn, a host never goes back to an older version on its own. A pin below the running version holds the host where it is rather than downgrading it. To run an earlier version on a host:- Stop further updates: set the toolkit’s policy to Notify only or Pinned.
- Recreate the container from the image tag you want, with
AO_AUTO_UPDATE=0set, so the container runs exactly that image’s version.
AO_AUTO_UPDATE=0, the
host immediately runs the newer update still on the volume, whatever
its policy.
To discard installed updates for good, also delete the update directory,
/var/lib/ao-toolkit/bin (or the directory AO_UPDATE_DIR names, if you
set it), on the volume or, on Kubernetes, on each node. Then restart the
container. It starts from the image’s version and follows its policy
from there.