writingmaclulich.com

We still deploy JARs over SSH. An AI made it reasonable.

I made an old-school Java deployment reasonable by putting an AI operator on top of the system Audience Republic already had: 17 long-lived production hosts, with no Terraform plan or disposable infrastructure between us and them.

We connected over SSH, applied PostgreSQL migrations, pulled JARs from S3, changed symlinks, restarted systemd services and watched the result. Rather than asking the AI to pretend this was a modern platform, I taught it how the platform actually worked, including the parts that could hurt us.

The deployment was simple, but operating it was not

Old deployment systems are awkward to talk about because the expected ending is a replatforming. Mention SSH and someone will ask why it is not Kubernetes. Mention JARs and the proposed fix quickly becomes a new delivery platform, even when the immediate problem is the risk around each release.

That was not the useful move here because the mechanics were simple and well understood. The expensive part was the operational judgement wrapped around them.

The basic operation was still this:

ssh "$host" "~/bin/deploy.sh 16.4"

On the host, that script pulled the production branch, stopped RabbitMQ consumption, waited two minutes for in-flight work, downloaded the selected JAR from S3, changed the service symlink and restarted it with systemd.

None of that is fashionable, but it is wonderfully legible. You can SSH to a box, inspect the symlink, ask systemd whether the service is active and curl its health endpoint without passing through seven control planes on the way to the process.

Rebuilding all of it would have been a large project, while making the next deployment safer was a problem we had today.

Kubernetes would have shifted the work

I knew from experience what the alternative cost. With a small team, managing a Kubernetes cluster is not an abstraction you buy once because it quickly becomes another production system in its own right.

Someone still owns cluster upgrades, node pools, networking, ingress, secrets, access control, logging, metrics, storage and the new failure modes introduced by all of it. That work lands on the same people who are meant to be building the product.

Kubernetes solves real problems, but it was not the bottleneck here. Our deployment mechanics were obvious, while the risk sat in their sequence and the operational memory around them. An AI safety layer addressed that risk without giving the team another platform to babysit.

The AI became the control plane

The agent did not replace the deployment. It made explicit a sequence that had previously lived across scripts, documentation, production incidents and the operator's memory.

  1. 01Migrateschema first
  2. 02Gateproduction clear
  3. 03Canaryone host
  4. 04Verifyhealth + logs
  5. 05Expandone wave
  6. 06Report17 hosts
The mechanics stayed simple while the agent made the order explicit and refused to skip the proof between stages.

One early version exposed a general remote-command method that accepted a host and an arbitrary command string. That is convenient until a version tag becomes part of the command and the remote shell interprets it.

# Too much authority in one interface
transport.run_command(host, command)

The replacement exposed typed actions instead, validating hosts, services and version tags at the boundary, quoting remote values and running local subprocesses without a shell.

runner.run_deploy_script(
    host=host,
    version_tag=version,
)

That boundary matters more when an AI is driving because the model can choose a deployment action, but it cannot invent a new remote shell action that the interface does not provide.

The agent had to know when not to deploy

A production restart could interrupt a CSV import or an outbound message send by removing the JVM while work was in flight. A message that reached RabbitMQ's dead-letter queue could be replayed manually, but often it simply disappeared with no delivery callback and no automatic retry.

The deployment agent checks database health and counts active imports and sends before a rollout. It repeats those gates between hosts because production state can change while a 17-host deployment is running.

The gate code preserves an important distinction between blocked and broken:

if error_results:
    can_deploy = False
elif blocked_results and not force:
    can_deploy = False

An active import is a known condition that an operator can investigate. A database connection error means the safety check never ran, so --force cannot turn missing evidence into a safe deployment.

The SMS and email senders are stricter again. They deploy alone, never bypass an active send, and must prove they have reattached to RabbitMQ before the next host starts.

A rollout advances one verified host at a time

The deployment tool had no production --single-host flag. The canary mechanism was more direct: temporarily narrow the YAML inventory to one host.

production:
  campaign:
    - worker-1  # canary

The agent deploys that host, waits for the result and verifies it before restoring a larger wave. Customer-facing hosts, HA pairs and the message senders still move one at a time.

A campaign host takes roughly 133 seconds, most of it intentionally. The remote script stops queue consumption and gives in-flight messages 120 seconds to drain before it restarts the JVM.

After the restart, the agent checks the JAR symlink for the expected commit, confirms systemd reports the service as active, calls /system/status and expects{"ok":true}, then reads the application's file log for new errors. On queue workers it also looks for consumers returning, and the next host does not become eligible until that evidence is present. A completed command on its own was never enough.

The old deployment got a live screen

The agent also needed a readable operational surface, so it kept one rolling message in the deployment channel and edited that message as the fleet changed.

A reconstructed deployment channel showing completed, active and waiting hosts with deployment times and a 17-host progress bar
A public-safe reconstruction of the live rollout message. Each host moved from waiting, to deploying, to verified without creating a new message for every step.

Every host had a line, and completed hosts gained a Sydney timestamp. The current host had a spinner that changed every 15 seconds, which mattered when one restart sat there for more than two minutes, while canary, HA, load-balanced and message-sender hosts carried their own annotations.

The progress bar was only 20 characters wide, but it showed enough to tell what had deployed, what was moving and what had not been touched.

That small interface changed the character of the system. Although the deployment was still SSH and JARs underneath, it now behaved like an observable rollout rather than a sequence of private terminal commands.

A JAR rollback is not a database rollback

Rolling a JAR back is straightforward once the operator chooses the exact previous version. The agent deploys it through the same one-host sequence and verifies the symlink and health checks again.

Rolling a database backwards carries a different class of risk because a down migration can drop data, change constraints or leave stored procedures in a state the old application does not understand.

The agent therefore rolls back JARs only, leaving the database forward-migrated while it checks whether the older application is compatible with the newer schema. Additive changes are normally fine, whereas a dropped column, a new required field or a changed stored procedure puts the decision back in front of a human.

The same gates, canaries and per-host verification apply during a rollback because urgency does not change what the system needs to prove.

The agent keeps the operational history visible

Sylvain Kalache recently wrote about AI incident tools creating comprehension debt. When software handles every routine incident, engineers lose the repetitions that taught them how their systems behave. They return only for the rare failure the AI could not solve. I think that risk is real, and this agent only helps if it keeps its reasoning visible: why a gate blocked, which host is active, what evidence passed and where the rollout will stop. The operator still chooses the rollout shape and confirms an exact rollback target.

Explanation still is not practice. Reading that a message sender must drain does not create the same understanding as diagnosing one that failed to resume, and I have not solved that part yet.

The next useful layer is a simulation mode: a deployment that injects a blocked gate, an unhealthy canary or an incompatible rollback and makes the human respond without touching production. The agent can preserve the incident memory, but we still need to exercise our own.

What changed was the way we operated it

We kept a deployment architecture that everybody could inspect rather than adding Terraform or moving the fleet to Kubernetes, then put a more disciplined operator over it.

That is the retro appeal for me: the machinery remains boring and direct, while the AI adds the memory, refusal and visibility that made operating it reasonable. You do not have to replace a system before AI can improve it, but you do have to teach the AI how the existing system really works and exactly when it should stop.