Runs and logs
Listing the runs of an Aetherfy agent
A run on Aetherfy is one ephemeral execution of a type: job agent. List
recent runs newest-first:
afy runs nightly-report
afy runs nightly-report --limit 50The default limit is 20 and the maximum is 100.
The listing shows six columns:
| Column | Contains |
|---|---|
| When | When the run was created, in UTC and labelled UTC — the instant -o json gives as created_at |
| Trigger | Why it ran — printed as the raw value, literally cron or manual |
| State | How it ended, or where it is now |
| Release | The release the run executed, as afy deployments numbers it (v3), or a dash when none was recorded |
| Duration | How long the machine ran |
| Run ID | The identifier to pass to afy logs --run |
Aetherfy lists only runs with a trigger source of cron or manual. Runs
started by a parent agent are excluded from this listing — they belong to the
parent’s history instead.
What the Aetherfy runs API returns
The equivalent Aetherfy REST call is:
GET /agents/{id}/runs| Query parameter | Accepted values | On anything else |
|---|---|---|
trigger_source | cron or manual only | 422 |
limit | 1–100 | 422 |
before | An ISO-8601 keyset cursor on creation time | 422 |
Each row Aetherfy returns carries these fields:
| Field | Contains |
|---|---|
id | The run’s identifier |
trigger_source | cron, manual, or spawn |
state | See the state table below |
created_at | When the run was created |
error_message | The failure reason, when the run failed |
release_version | The release the run executed — see below |
machine_started_at | When the machine began executing |
machine_stopped_at | When the machine stopped |
duration_seconds | Execution duration |
The two machine fields are null on a run of a type: service agent, which is a
request to a machine the service already has rather than a machine of its own. Its
duration_seconds runs from the moment that machine accepted the run to the moment
the run’s outcome was reported.
release_version is the version of the release a run executed — the number
afy deployments and afy rollback use, not the run’s own place in the
agent’s deployment history. A type: job run executes the release whose image it
was started from; when that image had been lost and the run rebuilt it, the
release whose stored code it was rebuilt from. A type: service run executes the
release the machine that took it was serving. It is null for a service run no
machine accepted, and for runs recorded before Aetherfy kept it.
The agent record carries its most recent entry of this list as last_run, which
is what afy status prints as Last run.
before is a keyset cursor rather than an offset, so paging backwards through a
long history stays correct even as new runs arrive.
Run and deployment states on Aetherfy
Aetherfy uses one state vocabulary for both deployments and runs. Four of the states are terminal — once a record reaches them it will not change again.
| State | Meaning | Terminal |
|---|---|---|
queued | Waiting in the build queue | No |
building | The image is building | No |
deploying | Machines are launching and being health-checked, or, for a run, the run is being handed to a machine | No |
active | Serving, or executing in the case of a run | No |
completed | A run’s machine ended on its own with exit code 0 | Yes |
failed | The build, the deploy, or the run failed | Yes |
superseded | Replaced by a newer successful deployment | Yes |
rolled_back | Replaced by a rollback | Yes |
completed applies to runs, which end by design. A long-lived service
deployment that is working sits at active and stays there until a newer
deployment supersedes it.
A run is active only once a machine has accepted it. Until then it is
deploying, and a run that no machine accepts within 2 minutes of Aetherfy
starting to hand it over is failed with “The run was never accepted by a
machine, so your code was not executed.” Your code did not start, so there is
nothing of yours to debug.
A run that stalls before that point, still queued, building or deploying 7
minutes after it was created, is also recorded as failed, with a reason naming the
step that never finished. The 7 minutes run from when the run was created, so time
the run spends queued behind other deployments counts toward them (see
/agents/task-contract). Aetherfy checks for stalled runs
every 5 minutes, so such a run can show its pending state for up to 5 minutes past
those 7 minutes before it is failed.
What decides completed versus failed for a run is the process exit code —
see /agents/task-contract.
Agent status values on Aetherfy
An agent has its own status, separate from the state of any individual deployment or run.
| Status | Meaning |
|---|---|
pending | Created, not yet built |
building | An image is building |
deploying | A version is launching and being health-checked; a version already deployed keeps serving meanwhile |
running | Deployed and operating |
failed | The most recent deployment failed and no earlier version is serving. A redeploy that fails leaves the agent running on the version it was replacing |
paused | Paused with afy stop — every machine is stopped and Aetherfy will not re-wake it. Resume with afy start |
stopped | Stopped by Aetherfy because the account was suspended for billing. You do not set this and cannot clear it directly — settling the account returns the agent to running. See Billing and spend caps |
usage_paused | Paused by Aetherfy at a spend limit — shown as paused for usage |
archived | Archived, freeing the plan quota slot |
deleting | Deletion in progress |
deleted | Deleted |
usage_paused is the one to recognise on sight: the agent is not broken and
nothing needs repairing in your code. Raise the limit or settle the payment at
https://app.aetherfy.com/dashboard/settings/billing
and the agent resumes.
Retrieving logs from Aetherfy
Everything an agent writes to stdout and stderr becomes its logs on Aetherfy.
afy logs nightly-reportScope the output to a single run, using a Run ID from afy runs:
afy logs nightly-report --run 9f3a1c72-5b8e-4d61-a0f4-7c2e9b5d3a18| Flag | Short | Default | Effect |
|---|---|---|---|
--tail | -n | 50 | How many lines to return. The server maximum is 1000. |
--follow | -f | off | Keep streaming new lines as they arrive |
--since | — | — | Only lines newer than a duration, e.g. 1h, 30m |
--level | — | — | Comma-separated severities, e.g. ERROR,WARN |
--stream | — | — | Comma-separated streams: stdout, stderr, system |
--run | — | — | Restrict to a single run id. On a type: service agent this returns Aetherfy’s lines about that run — sent, answered, how long — not your service’s own output |
Combined:
afy logs nightly-report --tail 200 --since 1h --level ERROR,WARN --stream stderrOne caveat worth knowing before you debug against it: on Aetherfy --follow
polls, and it ignores both --tail and --since. It also does not honour
JSON output. Use --follow to watch what happens next; use --tail and
--since to look at what already happened.
Log record fields on Aetherfy
Every log record Aetherfy stores carries:
| Field | Contains |
|---|---|
| agent | Which agent emitted it |
| deployment | Which deployment was running |
| timestamp | When the line was emitted |
| stream | stdout, stderr, or system |
| level | The severity |
| message | The line itself |
The system stream holds lines Aetherfy itself emits about the machine, as
distinct from your program’s own output on stdout and stderr.
How Aetherfy reads a line’s level
Your output carries no level field, so Aetherfy reads one from the start of each
line, and only from there. A level is recognised when the line begins — after an
optional timestamp such as 2026-09-17 12:40:36,123 or [2026-09-17T12:40:36Z] —
with one of DEBUG, INFO, WARN, WARNING, ERROR, CRITICAL, FATAL or
TRACE, written in one of three ways:
| Written as | Example |
|---|---|
| In square brackets, any case | [error] upstream closed |
| Followed by a colon, any case | error: config file missing |
| In capitals, followed by a space | ERROR disk full |
WARNING is stored as WARN, CRITICAL and FATAL as ERROR, and TRACE as
DEBUG. A line with no level at its start takes its stream’s: INFO on
stdout, ERROR on stderr. The word anywhere else is part of the message — No Error found on stdout is INFO — and so is a capitalised word opening a
sentence, such as Error handling is on. This is best-effort parsing: to be sure
of a line’s level, print it in one of the three forms above.
What Aetherfy does not record for you
Aetherfy logging is plain stdout/stderr capture, and nothing more. It does
not auto-instrument calls your agent makes to model providers — there is no
automatic capture of prompts, completions, token counts, latency or tool calls
from an OpenAI, Anthropic or other SDK running inside your agent.
If you want that, log it yourself. A print() or logger.info() around the call
lands in exactly the same place as the rest of your output and is searchable with
afy logs --level:
import json, time
started = time.monotonic()
response = call_your_model(prompt)
print(json.dumps({
"event": "llm_call",
"model": "your-model-id",
"elapsed_ms": round((time.monotonic() - started) * 1000),
"prompt_chars": len(prompt),
}))Keep the per-line and per-deployment limits below in mind if you log full prompts — they are easy to exceed.
Log limits and retention on Aetherfy
Aetherfy caps log volume in several dimensions. Logs are an operational aid, not a durable data store — write anything you need to keep to a real destination.
| Limit | Value |
|---|---|
| Retention | 7 days |
| Per deployment | 5 MB |
| Per line | 4 KB |
| Chunks per minute, per agent | 60 |
| Per upload body | 256 KB |
When an agent exceeds the per-line or volume caps, the Aetherfy log forwarder backs off and emits a marker in the stream so the gap is visible rather than silent:
[SYSTEM] N log line(s) droppedIf you see that marker, the agent is logging faster than Aetherfy will accept. Reduce per-line size or log volume — a very large payload dumped per iteration is the usual cause.
A run’s output is read to its last line before the run’s outcome is reported. If your code starts a background process that keeps the run’s output open after your main process exits, Aetherfy stops reading 2 seconds later, ends the run, and says so in the run’s log:
[SYSTEM] output from a process run <run id> left running was cut off 2s after the run endedAnything that background process prints after that point is not captured.
While a run is in flight, Aetherfy keeps its machine from being suspended. If
that cannot be established after several attempts, the run’s log says so once,
at ERROR, and Aetherfy keeps trying:
[SYSTEM] Aetherfy run <run id> could not hold its machine awake: 5 attempts in a row did not connect (last: <reason>). The provider may suspend the machine mid-run; still retrying.A long run that stops partway with this line in its log was most likely suspended by the platform, not ended by your code.