> ## Documentation Index
> Fetch the complete documentation index at: https://docs.automatedoperations.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Automatic updates

> How container installs of the AO Toolkit update themselves: update policies, the safety checks every update passes, the per-host off switch, and troubleshooting.

A Toolkit running in a container can update itself to new releases. A
small supervisor built into the image downloads the new version, checks
it, and switches to it. The container keeps running: it isn't recreated,
and on Kubernetes the pod isn't restarted. Only the Toolkit process
inside restarts, so the host disconnects from AO briefly and reconnects
on the new version. It keeps its identity, so it doesn't enroll again.

You decide whether that happens. Each organization has an update
policy, each toolkit can override it, and each host has a local off
switch that AO can't override.

<Note>
  Docker, Docker Compose, and Kubernetes DaemonSet installs update
  themselves. Hosts installed with the install script or as a standalone
  binary, and the Kubernetes cluster controller, don't: AO shows them
  when a newer release is available, and you update them yourself.
</Note>

## Update policies

| Policy | What happens |
| - | - |
| **Notify only** (the default) | AO shows when a newer release is available. Nothing is installed. |
| **Automatic** | Each toolkit updates to a new release as AO rolls it out. |
| **Pinned** | Toolkits hold at the version you choose. Hosts below it update to it; hosts above it stay where they are. |

Pin to a published release. AO refuses a pin newer than the latest
release, but it doesn't check that an older version string was ever
published, and hosts can't download a version that wasn't: they report
**Download failed**. Pinning never downgrades a host, and if AO has
paused the release you pinned to, the pin waits until AO resumes it.

Until your organization saves a policy, it follows AO's default, which
is **Notify only** today. While it does, the settings page marks that
policy **(default)**. AO may change its default in future, so if you
rely on **Notify only**, select it and save it explicitly.

### Set the organization policy

Go to **Settings** → **Toolkit updates** and choose **Automatic**,
**Notify only**, or **Pinned**. For **Pinned**, enter the version under
**Version to hold at** (for example `0.0.1-rc50`), then **Save**. The
policy applies to every toolkit in the organization that doesn't have
its own.

### Override it for one toolkit

Open the toolkit from **Toolkits**. In the **Updates** card, under
**This toolkit's policy**, choose a policy and **Save**. Choose **Use
organization policy** to remove the override.

The toolkit's own policy always wins over the organization's. The
**Update policy** line on the toolkit shows which one applies: **Set on
this toolkit**, **Organization policy**, or **AO default**.

Changing either policy needs the **Edit toolkit upgrade policy**
permission. The Owner, Admin, and Toolkits admin roles have it. For a
single toolkit's override, the permission can be limited to toolkits
with particular tags. Changes reach connected toolkits without a
restart.

## How an update is kept safe

**Every release is signed by AO and verified on the host.** The Toolkit
checks the release's signature, platform, and checksum against keys
built into its image before it runs a new version, and again every time
it starts. A download that doesn't verify is never run.

**A new version must pass a health check, or it's rolled back.** The
new version must connect to AO within three minutes of starting. It then
runs a self-test that proves it can download and verify the next
release. If it doesn't connect in time, exits, or fails the self-test,
the host goes back to the version it was running, automatically. For
the 24 hours after an update, three crashes within ten minutes also
roll it back (this needs a container restart policy; see
[Requirements](#requirements)). The host keeps the previous version on
disk so it always has one to go back to.

If the self-test can't reach the registry, the new version keeps
running while the self-test retries. After three unreachable attempts,
or six hours without a pass, the host goes back to the previous version
and tries the update again later.

A version that failed its health check on a host is never offered to
that host again. A failure caused by the network rather than the
release, like the unreachable registry above, is retried later
instead.

**AO never downgrades a host through a policy.** Policies only ever move a host forward.
The one exception is when **AO withdraws a release**: a host that
updated to the withdrawn version moves back to the version it ran
before. It does so within about an hour, as long as it's connected to
AO and can reach the registry. A withdrawn version that a host runs
because it's the version in the container image stays until you
recreate the container from a different image.

**Releases roll out gradually, and AO can stop a rollout.** Under
**Automatic**, a new release reaches a growing share of toolkits over
time, not all of them at once. If hosts start rolling back, AO stops the
rollout. A pinned version skips the gradual rollout, because you chose
it, but still waits while AO has stopped it.

**Running work is allowed to finish.** Before it switches, the Toolkit
stops accepting new actions and waits up to five minutes for actions in
progress to finish. New actions sent during that wait are refused with
a retryable error. If the work doesn't finish in time, the Toolkit
accepts actions again, postpones the update, and tries again two
minutes later, with the same wait. A host that always runs long actions
can therefore refuse new ones repeatedly. To avoid that, pin the host
(or its organization) while that work runs, and move the pin forward
when the host is quiet.

## Requirements

* **Run the image's own entrypoint.** Don't override the container's
  `entrypoint` (`--entrypoint`, or `command:` in Kubernetes). The
  entrypoint starts the supervisor that installs updates. Without it, the
  Toolkit runs normally but never updates, and AO shows the host as
  having updates off. Passing flags as container arguments (`args:` in
  Kubernetes) is fine.
* **Keep `/var/lib/ao-toolkit` on a persistent volume.** Updates are
  stored there, next to the host's identity, so they survive the
  container being recreated. The install snippets in the dashboard already
  do this: a named volume for Docker and Compose, and a `hostPath` on the
  node for the Kubernetes DaemonSet. Without it, a recreated container
  starts again from the image's version.
* **Let the host reach the container registry.** The Toolkit downloads
  releases itself, over HTTPS, from `ghcr.io` and
  `pkg-containers.githubusercontent.com` (where `ghcr.io` serves
  downloads from). It honours `HTTPS_PROXY` and `NO_PROXY`. To use a
  registry mirror, set `AO_UPDATE_REGISTRY` to its host (or an
  `https://` URL). The mirror must serve the Toolkit's packages exactly
  as `ghcr.io` does, under the same names and paths, including the
  release package the update check reads, not just the container image.
  The signature check is the same whichever registry the bytes come from.
* **Allow programs to run from the update directory.** Updates are run
  from `/var/lib/ao-toolkit/bin`. If that volume is mounted `noexec`
  (Container-Optimized OS on GKE mounts `/var` that way), the Toolkit
  refuses to update until you set `AO_UPDATE_DIR` to a directory that
  allows it. Put that directory on persistent storage too, or updates are
  downloaded again after the container is recreated.
* **Give the container a restart policy.** If the Toolkit crashes, the
  container exits, and the crash-rollback check above only counts
  crashes when the container is restarted. The install snippets use
  `--restart unless-stopped` for Docker and `restart: unless-stopped`
  for Compose. Kubernetes restarts the pod's container itself.

### Docker and Docker Compose

* `docker ps` and your Compose file keep showing the image you started
  with. After an update, the Toolkit runs a newer version than that
  image. AO shows both, for example `running v0.0.1-rc52 (image
  v0.0.1-rc50)`.
* Restarting the container, or recreating it with the same volume, keeps
  the update.
* If you start the container from an image newer than the installed
  update, the image's version runs and older updates are removed from
  the volume.

### Kubernetes DaemonSet

* The update happens inside the running pod. The pod spec doesn't
  change, so `kubectl` keeps showing the image in your manifest, and the
  pod isn't restarted.
* Updates are stored in `/var/lib/ao-toolkit` on each node. A pod that
  restarts on the same node keeps its update. A pod on a new node starts
  from the image's version and updates according to its policy.
* GitOps tools such as Flux and Argo CD don't see a change, so they
  neither report drift nor revert the update. If your pipeline should
  decide which version runs, use **Notify only** or **Pinned**, or set
  `AO_AUTO_UPDATE=0` in your manifest. The DaemonSet snippet in the
  dashboard includes that line, commented out.

### Kubernetes cluster controller

The cluster controller (the `k8s-controller` role) never updates itself.
Update it by changing the image in its StatefulSet.

## Turn off updates on one host

Set `AO_AUTO_UPDATE=0` in the container's environment:

```bash theme={null}
docker run -d ... -e AO_AUTO_UPDATE=0 ... <image>
```

```yaml theme={null}
# Docker Compose
environment:
  AO_AUTO_UPDATE: "0"
```

```yaml theme={null}
# Kubernetes DaemonSet
- name: AO_AUTO_UPDATE
  value: "0"
```

The switch is read on the host, and nothing AO sends can turn it back
on. Updates are **on** when `AO_AUTO_UPDATE` is unset, empty, `1`,
`true`, or `on` (ignoring case and surrounding spaces). **Any other
value turns them off**, so a typo leaves updates off rather than on.
Environment changes take effect when the container is recreated.

With the switch off, the container runs exactly the version in its
image, with no supervisor. If an update was installed earlier, it stays
on the volume but doesn't run, and AO shows it as downloaded, not
running. Turning the switch back on returns to that update, as long as it
still verifies.

AO shows a host with the switch off as **updates off**. Whatever the
policy, that host won't install updates.

<Note>
  Toolkit images **0.0.1-rc49 and earlier** turn this switch off in the
  image itself, so they show as **updates off**. To get automatic
  updates, recreate the container from a current image: `docker pull`
  the image, `docker rm -f ao-toolkit`, then run the same `docker run`
  command (or `docker compose pull && docker compose up -d`). Keep the
  same volume so the host keeps its identity.
</Note>

## What you see in AO

On **Toolkits**, a toolkit row can carry these markers:

| Marker | Meaning |
| - | - |
| **updates off** | Automatic updates are off on this host: `AO_AUTO_UPDATE` is off there, or the container isn't running the image's own entrypoint, which installs updates. |
| **pinned v…** | The toolkit's policy pins it to that version. |
| **update failed** | The last update didn't succeed. Hover for the result and versions. |

**update failed** appears only for **Rolled back**, **Verification
failed**, and **Download failed**. **Refused** and **Inconclusive** don't
show a marker; open the toolkit to see them.

On a toolkit's page, **Toolkit version** shows what's running, as
`running v… (image v…)` after an update. The **Updates** card shows the
policy and where it comes from, the next version, the last result with a
short explanation, and **Recent updates**, which include the host's own
detail message for each result.

| Result | Meaning |
| - | - |
| **Updated** | The new version passed its health check and is running. |
| **Rolled back** | The new version failed its health checks, so the host went back to the previous one. It won't be offered that version again. |
| **Verification failed** | The download didn't match AO's signature, so it wasn't installed. |
| **Download failed** | The host couldn't download the update. It needs HTTPS access to `ghcr.io` and `pkg-containers.githubusercontent.com`. It will retry. |
| **Deferred** | Work was running on the host. It will try again shortly. This result also covers a temporary disk or state-file error on the host; the detail message says which. |
| **Inconclusive** | The host couldn't confirm the update because of a network or registry problem. It will retry. |
| **Refused** | The host declined this update. The detail message says why. |
| **Release withdrawn** | AO withdrew the version this host was running, so it moved to another release. |

## Troubleshooting

### A host isn't updating

Check, in order:

1. **The policy.** The **Update policy** line must be **Automatic**, or
   **Pinned** to a version above the one running. **Notify only** never
   installs anything. Under **Automatic**, a new release reaches hosts
   gradually, so some hosts get it later than others.
2. **The local switch.** If the host shows **updates off**, look for
   `AO_AUTO_UPDATE` in its environment, and check its image version:
   images 0.0.1-rc49 and earlier set it off.
3. **The entrypoint.** A container started with `--entrypoint` or a
   Kubernetes `command:` never updates and shows **updates off**. Remove
   the override.
4. **The install type.** Install-script and binary installs, and the
   cluster controller, don't update themselves.
5. **The last update result.** **Download failed** means the host can't
   reach the registry (see [Requirements](#requirements)). **Refused**
   carries a detail message, for example a `noexec` update directory or
   an image too old to verify the current release, which needs a
   recreate from a current image.
6. **Versions that failed here.** A version listed under **Not offered
   again on this host** won't be retried on that host. A later release
   will be.

A persistent volume doesn't affect whether an update installs, but
without one a recreated container goes back to its image's version.

### Read the logs

The Toolkit logs one JSON line per event. On Docker:

```bash theme={null}
docker logs ao-toolkit 2>&1 | grep -E 'upgrade|health.gate|started toolkit|rolling back|rollback|already failed|is available|yanked'
```

On Kubernetes, use `kubectl logs` on the DaemonSet pod. The messages to
look for:

| Message | Meaning |
| - | - |
| `upgrade scheduled` | The host will start downloading the new version after a short delay, up to 10 minutes, so that hosts don't all download at once. |
| `upgrade staged and selected; restarting onto it` | The new version is downloaded and verified, and the Toolkit is switching to it. |
| `started toolkit` with `"health_gate":true` | The new version is starting and must pass its health check. |
| `toolkit passed the health gate; committed` | The update succeeded. |
| `rolling back` | The new version failed. The line names the version, the version it went back to, and the reason. |
| `not retrying a version that already failed on this host` | The host skips a version that failed here earlier. |
| `toolkit … is available; updates are notify-only for this toolkit or its org, so it stays where it is` | The policy is **Notify only**. |
| `toolkit … is available; AO_AUTO_UPDATE is off on this host, so the toolkit stays where it is` | The local switch is off. |
| `classified the supervisor's rollback` | After a rollback, the previous version reconnected and recorded the result (**Rolled back** or **Inconclusive**). |
| `the running version is yanked; restarting onto the previous binary` | AO withdrew the running version, so the host is moving back. |
| `upgrade outcome` | A download, verification, deferral, or refusal result was recorded. Its `result` and `detail` match what AO shows. **Updated** and **Rolled back** are logged by the lines above instead. |

`docker exec ao-toolkit ao-toolkit version` reports the image's
version, not an installed update. The **Toolkit version** in AO shows
what is actually running.

### Go back to an earlier version

Apart from moving off a release AO has withdrawn, a host never goes
back to an older version on its own. A pin below the running version
holds the host where it is rather than downgrading it. To run an
earlier version on a host:

1. Stop further updates: set the toolkit's policy to **Notify only** or
   **Pinned**.
2. Recreate the container from the image tag you want, with
   `AO_AUTO_UPDATE=0` set, so the container runs exactly that image's
   version.

Recreating from an older image **without** turning the switch off isn't
enough: an update installed on the volume is newer than the image, so it
keeps running. Likewise, if you later remove `AO_AUTO_UPDATE=0`, the
host immediately runs the newer update still on the volume, whatever
its policy.

To discard installed updates for good, also delete the update directory,
`/var/lib/ao-toolkit/bin` (or the directory `AO_UPDATE_DIR` names, if you
set it), on the volume or, on Kubernetes, on each node. Then restart the
container. It starts from the image's version and follows its policy
from there.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.